ABSTRACT. This paper presents a domain-specific language
(DSL) with limited input/output (I/O) by construction. By limiting
risky operations at the language level, this provides safety com
parable to the limited system access provided by standard tool
calling approaches. Using the DSL, input token costs of agentic
touchstone analysis were reduced by nearly 70% on well-crafted
prompts while retaining comparable or higher accuracy across
models of different capabilities. We evaluate both a standard
tool-calling approach and the DSL harness on 10 analysis tasks
used to gauge inference cost and model accuracy.
ABSTRACT. Product Design Guides (PDGs) contain dense visual information, including placement and routing rules, schematics, and plots that are time-consuming for design engineers to search for and interpret manually, while the existing off-the-shelf LLMs cannot comprehend the complex domain-specific data in these documents effectively. To this end, we introduce a multimodal fine-tuning dataset derived from PDGs to adapt LLMs for PDG image comprehension. To the best of our knowledge, this is the first publicly available dataset developed for PDG image comprehension task. We evaluate the dataset by fine-tuning four open-source LLMs across eight configurations and measuring
performance using semantic similarity and factual accuracy. The results show that fine-tuning with the proposed dataset improves accuracy by up to 19% relative to the corresponding base models. These findings demonstrate the potential of multimodal LLMs
to help design engineers extract information from PDGs figures quickly and accurately, thereby streamlining the design process.
Bayesian Optimization of Crossbar-Based Compute-In-Memory System Design for Efficient DNN Inference
ABSTRACT. Leveraging the high density and energy efficiency of Compute-In-Memory (CIM) crossbar–based Deep Neural Network (DNN) accelerators requires optimal Design Space Exploration (DSE), which becomes increasingly challenging as complex models for advanced AI workloads expand the highly non-convex design space. Among existing DSE approaches, multi-objective Bayesian Optimization (BO) is promising, as it explores high-quality design solutions while querying costly CIM simulators selectively. In this work, we propose a multi-objective BO framework that holistically co-optimizes hardware and algorithm parameters of a CIM crossbar–based hardware accelerator for various DNN inference tasks. Depending on NN model depth, our framework handles high-dimensional design spaces (with $26$ and $50$ dimensions) and extremely large search complexities on the order of $O(10^{12})$ and $O(10^{27})$ for VGG8/CIFAR-10 and VGG16/Tiny-ImageNet-200. Our method attains $91.72 \%$ and $57.2 \%$ accuracy, respectively, comparable to baseline designs, while improving chip area ($65.52 \%$ and $50.7 \%$), read latency ($9.52 \%$ and $13.27 \%$), read dynamic energy ($31.23 \%$ and $52.07 \%$) and increasing memory utilization ($13.41 \%$ and $2.67 \%$).
Cost-Aware Design Space Exploration of Yield-Compliant UCIe Interconnects
ABSTRACT. The design of UCIe die-to-die interconnects requires satisfying signal integrity requirements while minimizing fabrication cost under different temperatures, data rates, and process-variation levels. This paper presents a cost-aware design space exploration of UCIe-A die-to-die interconnects to identify cost-optimal interface geometries satisfying target signal integrity yield constraints. Four routing configurations are evaluated across UCIe 2.0 data rates, maximum interface temperatures, and process-variation levels. The resulting design space is used to quantify cost versus signal integrity yield tradeoffs and develop practical guidelines for selecting cost-effective UCIe interfaces.
Signal Integrity Design and Optimization of High-Bandwidth Memory Test System
ABSTRACT. The signal integrity design of high-bandwidth memory (HBM) test systems presents unique challenges due to significantly longer and more complex channels compared to end-use HBM platforms. This work presents an HBM test system architecture and a system-level signal integrity analysis of the complete tester channel and identifies the silicon interposer as the dominant factor to channel degradation. A silicon-substrate-isolated interposer routing is then proposed to reduce substrate-induced losses and improve signal integrity. Transient simulation results show the optimized silicon interposer improves the overall timing margin of the HBM test system.
Dual-Referenced Circuit Ports for Hybrid Modeling and Correlation of HPC Memory Interfaces
ABSTRACT. With the rapid escalation of data rates and routing densities in High-Performance Computing (HPC) memory interfaces, ensuring accurate simulation-to-measurement correlation is paramount for signal integrity (SI) design and verification. Standard full-wave 3D electromagnetic (EM) simulators demand excessive computational resources and are unsustainably slow when modeling complex and highly dense memory channels. In addition, for memory-dense designs, setting up numerous concurrent wave ports introduces significant engineering complexity and may even necessitate modifications to the layout. The hybrid EM partitioning has been widely used in industry to reduce simulation runtime. However, the usage of single-referenced ports in the conventional method often leads to inaccurate simulation results. To address these challenges, this paper proposes a robust, scalable, and efficient method which utilizes dual-referenced ports in hybrid 3D/2.5D EM simulation methodology for accurate correlation of HPC memory interfaces. The theoretical background of dual-referenced circuit ports is studied, and the associated EM formulation is given, which clearly shows the continuity of voltage and current is better maintained across the partition boundary with dual-referenced ports. The proposed approach is validated on a cutting-edge HPC hardware platform through a comprehensive simulation-to-measurement correlation study, where close agreement with vector network analyzer measurements in both frequency and time domains is reached.
Mitigating the Impact of Glass Weave on Bidirectional High-Speed Signaling
ABSTRACT. We analyze the impact of fiber-weave on high-speed bidirectional links and introduce fiber-weave-induced echo as an additional SI consideration. We then show a panel-rotation-based mitigation strategy to minimize fiber-weave impact on echo, skew, and loss resonance without performing time-consuming 3D EM simulations.
A Low-Power, Low-Cost USB4 Cable Using an Ultra-Wideband Signal Integrity Recovery Structure
ABSTRACT. This paper presents a low-power, low-cost USB4 active cable architecture based on an ultra-wideband signal integrity recovery structure (SIRS) employing planar balanced lines (BLs), with a linear redriver amplifier used solely for channel-loss compensation. To reduce the reliance on retimers, the passive SIRS functions as an ultra-wideband common-mode rejection filter and an autonomous phase and amplitude-balancing structure from DC to over 40 GHz, while the linear redriver efficiently compensates for insertion loss. Simulations based on component measurements under 40 Gb/s NRZ channel conditions demonstrate a significant improvement in the effective differential eye and a reduced common-mode voltage.
A Differential Tab Structure Design for Optimizing 200G+ PCB Via Transitions
ABSTRACT. Printed Circuit Board (PCB) via transitions increasingly limit the performance of 200G+ channels as signal frequencies extend beyond 50 GHz. Common optimization methods include anti-pad sizing, stub removal, non-functional pad (NFP) design, drill diameter, and ground return via placement, but these methods leave limited room for further improvement. This paper introduces a differential tab structure at the via breakout as an additional design parameter. The differential tabs add local capacitance to compensate for the inductive behavior of the via transition. Full-wave simulation results show that the optimized tabs extend the frequency range with Sdd11 below −20 dB from 60 GHz to about 90 GHz with about 4 dB lower PSNEXT at via break-out from 40GHz to 100 GHz, with minimum change in insertion loss and mode conversion. Differential tab parameter sweep identifies tab width and distance from the via as the main factors for high-frequency matching. The structure needs only small layout changes and is compatible with standard PCB fabrication.
Balanced Line-based Ultra-wideband Channel with Enhanced SNR for 224G/448G SerDes Links
ABSTRACT. Conventional differential lines on printed circuit boards (PCBs) face significant challenges for the ultra-high-speed digital channels for 224G/448G SerDes links due to excessive insertion loss, sub-picosecond skew requirements, crosstalk, and ultra-wideband impedance control. Through optimized designs and rigorous full-wave simulations, planar balanced-line structures are shown to resolve these critical bottlenecks. By mitigating signal loss and strongly suppressing common-mode noise, these structures significantly enhance the signal-to-noise ratio (SNR) from DC to over 100 GHz.
Common-Mode Sign-Based P/N Skew Calibration for 212.5-Gbps PAM-4 Wireline Receivers
ABSTRACT. This paper presents an RX-side P/N skew calibration technique for ADC/DSP-based PAM-4 wireline receivers. A system-level model is developed using measured S-parameters of a 1-m direct-attach copper (DAC) cable channel and package interconnects, which exhibit approximately 33 dB insertion loss at the Nyquist frequency. For a 212.5-Gbps PAM-4 link, the unit interval (UI) is only 9.41 ps, making the receiver highly sensitive to P/N skew. The proposed calibration exploits skew-induced common-mode modulation at the RX input and estimates the skew polarity using a sign-correlation metric between the sampled differential transition and common-mode polarity. The accumulated metric is then used to adaptively adjust the relative P/N sampling phase and reduce the residual skew. System-level simulations demonstrate that the proposed technique maintains the baseline 20.6-dB SNDR over a ±2.5 UI P/N skew range.
Voltage Supply Noise Analysis of Phase-Locked Loop (PLL) Structure in Custom Base Die of High Bandwidth Memory (HBM)
ABSTRACT. In this paper, we analyzed the voltage supply noise of phase-locked loop (PLL) circuit embedded in high bandwidth memory (HBM) custom base die, focusing on power integrity (PI). PLL can play a critical role in stabilizing the internal clock of next-generation HBM’s custom base die with active circuit integration; however, the long power delivery path of HBM injects supply noise into PLL and degrades its performance. For analysis, Hierarchical power distribution network (PDN) is modeled based on the HBM4 VDDC domain, and two types of PLL, each employing a different VCO—an LC-type and a ring-type—are designed to analyze the supply noise under different current profiles. Based on this analysis, the LC-type VCO was found to generate lower supply noise than the ring-type. Furthermore, a decoupling capacitor (decap) placement analysis is performed to mitigate the noise, reducing the peak-to-peak supply voltage noise by up to 46.40% for the LC-type and 77.19% for the ring-type, with an additional 1.5 percentage-point reduction on average depending on the decap placement.
Impact of Ground Bounce on Logic-HIGH in Differential Signalling: Case Study of a Current-Mode Driver
ABSTRACT. The work presented in this paper investigates the impact of ground bounce on logic levels in differential high-speed
signalling schemes. In the literature, Ground Bounce Noise (GBN) has been associated with deterioration in logic-LOW, whereas its impact on other logic levels in differential signalling has not been addressed yet. The paper provides a thorough investigation of the impact of GBN on both the logic levels, demonstrating this through both the simulations as well as steady-state analysis. As differential signalling inherently cancels common-mode noise, the investigation is focused on current-mode driver (CMD) circuits. This work provides practical insights for quantifying noise in high-speed systems. The close agreement between SPICE-based transient simulations and steady-state analysis across UMC 28
nm, 40 nm, and 65 nm process nodes and additionally the measurement results from the experimental setup validates the
observations/claims.
Simulation-driven Adaptive Estimation of Worst-case Power Noise
ABSTRACT. This paper proposes a fully automated algorithm to estimate the worst-case voltage droop of linearized voltage-regulated power distribution network models. The method is based on a multi-input adaptive Time-Domain Vector Fitting scheme, which iteratively produces an optimized PDN model, a set of worst-case current loads, and the associated droop.
Design of Integrated Voltage Regulators using a Generative Waveform Surrogate
ABSTRACT. This paper presents a machine-learning framework, Net2Pareto, for high-conversion-ratio integrated voltage regulators (IVRs) that replaces slow per-design SPICE simulation by pairing a transformer-based generative waveform surrogate with a reinforcement-learning (RL) optimizer, evaluating coupled circuit and package designs in milliseconds to autonomously return an efficiency-versus-footprint Pareto front.
Equal Source Voltages for Minimum DC Conduction Loss in High-Current Power-Delivery Interconnects
ABSTRACT. DC conduction loss in the resistive interconnect between voltage regulator module (VRM) outputs and load inputs is becoming a bottleneck as processor current demand exceeds 1000 A. Because this loss scales with the square of current, even a small interconnect resistance can cause substantial power dissipation. This paper derives a simple optimality condition for reducing interconnect conduction loss: for a given interconnect and fixed load currents, equal source voltages minimize the loss. In contrast, conventional equal-current control in multiphase VRMs can increase interconnect loss when the source paths are asymmetric. In a 360-A PDN example, equal-current control increases interconnect loss from 51.15 W to 54.46 W, a 6.5% penalty. This optimal condition provides practical guidance for the design and control of high-current power delivery networks (PDNs).
Non-Uniform Channel Allocation in Multi-Fin GaN Tri-Gate Transistors for Package-Integrated Power Delivery
ABSTRACT. Non-uniform allocation of vertically stacked channels across the fins of a multi-channel AlGaN/GaN tri-gate transistor is introduced as a footprint-neutral design variable for package-integrated power delivery. A four-fin device is considered as a fully enumerable case study, with one to four channels per fin yielding 256 candidate arrangements. Four Gaussian-process regression surrogates predict on-resistance, gate and gate-drain charges (QG and QGD), and the switching–conduction figure of merit (FoM). From 100 fully characterized TCAD arrangements, 80 are used for training and 20 are reserved as a fixed held-out test set. The corresponding mean absolute percentage errors are 10.3% for Ron, 3.8% for QG, 3.4% for QGD, and 5.6% for FoM. The best identified channel arrangement is verified by TCAD to achieve a FoM of 1.906 Ω·pC, which is 17.0% lower than that of the best uniform arrangement. A new effective filling factor is introduced to quantify the combined lateral fin density and vertical channel utilization. A TCAD comparison of two compositionally identical equal-effective filling factor arrangements confirms a 2.19× FoM difference associated with their spatial allocation.
Model-Driven Design Exploration of Hybrid Voltage Regulators for Vertical Power Delivery
ABSTRACT. Point-of-load power conversion for AI accelerators is no longer a single-converter problem—it is a coupled topology, device, passive, switching frequency, and packaging co-optimization that cannot be tractably solved by manual design. A model-driven framework for series-capacitor buck hybrid voltage regulators (SC-Buck HVRs) is proposed. Given user-specified conversion ratio, load current, current density target, and package model, the framework enumerates all feasible multi-phase architectures and co-optimizes devices, passives, and switching frequency over real component libraries via a closed-form loss model. Ripple, saturation, voltage-rating, and area constraints are enforced throughout the optimization. The framework is validated against Cadence Virtuoso/Spectre-designed SOTA topologies at 95.27 ± 1% loss prediction accuracy. Up to +4.6 pp efficiency improvement and 29.2× footprint reduction over SOTA baselines are demonstrated at identical topology. Packaging-aware design-space exploration demonstrates that the optimal converter topology and component selections vary significantly across package implementations, highlighting the necessity of package-aware optimization.