ABSTRACT. A novel algorithm is proposed for the computation of parameterized macromodels from sampled frequency responses. The main new contribution is a set of explicit constraints that allows placing the parameter-dependent model poles in arbitrary regions of the complex plane. We show that this capability leads to more robust, accurate, and noise-insensitive models.
An Efficient NN Architecture for Harmonic Balance-based Analysis for SI Applications
ABSTRACT. This paper introduces a two-stage NN framework that accelerates Harmonic Balance (HB) steady-state analysis for signal and power integrity applications such as analysis with nonlinear SerDes analog front-end circuits and power-delivery networks. Traditional HB simulations become computationally expensive as harmonic count increases, creating bottlenecks in high-speed receiver or power-supply induced jitter analysis where strong nonlinearity and square-wave type behavior require many harmonics. The proposed method combines an unsupervised Autoencoder trained on low order harmonic data with a supervised Regression Network that predicts the high-order harmonic spectrum using only a small number of expensive HB simulations. Validating examples demonstrate significant speed-up during training as well as several orders of speed-up during response computation.
Deep Equilibrium Recurrent–Transformer (DERT) Model for Large Coupled Interconnect Systems
ABSTRACT. This paper presents a deep equilibrium recurrent-transformer (DERT) model for efficiently analyzing large multiconductor transmission line systems. The proposed method is developed using input data generated from a small subset of coupled lines with distributed MTL model and a relatively lower-order lumped segment model, enabling faster training. The validation example demonstrates significant speedup with accuracy comparable to higher-order models.
Thermal-Aware Signal Integrity Analysis of GPU-HBM-HBF Interposer Channels Considering HBF Power Consumption
ABSTRACT. This paper presents a thermal-aware signal integrity (SI) analysis of interposer channels in a graphics processing unit-high bandwidth memory-high bandwidth flash (GPU-HBM-HBF) package for high-capacity artificial intelligence (AI) inference systems.
Since NAND flash-based HBF has operation-dependent power consumption, it can alter the package thermal distribution and degrade the SI of interposer channels.
To evaluate this effect, GPU-HBM and HBM-HBF interposer channels were analyzed by comparing a conventional GPU-HBM package with GPU-HBM-HBF packages under different HBF power consumption conditions.
The results showed that the GPU-HBM channel was weakly affected by HBF power consumption, with only 0.031 dB insertion-loss variation at the 3.2 GHz Nyquist frequency, whereas the HBM-HBF channel exhibited 0.409 dB insertion-loss degradation and eye-opening reduction associated with the HBF-side temperature rise.
These results indicate that HBF-induced thermal distribution should be considered in the SI design of HBM-HBF interposer channels.
Transformer-Based S-Parameter Prediction for Chiplet Interconnects with Sparse Global Context
ABSTRACT. We present a transmission-line-informed Transformer with sparse global context for chiplet-interconnect S-parameter prediction. It exceeds reimplemented GNN-SP in five of eight Chiplet-SI components by R2, and an ordered port-readout variant surpasses GNN-SP on NEXT crosstalk.
Physics-Informed Wideband S-Parameter Prediction for UCIe Interconnect using Resonance-Aware Gaussian Process Regression
ABSTRACT. Fast surrogate modeling of UCIe interconnects is critical for chiplet design. However,conventional wideband Sparameter models suffer from errors in the high-frequency range. We propose Resonance-Aware Gaussian Process Regression (RAGPR), a physics-informed machine learning framework. By incorporating temperature-dependent material profiles and extracting resonance components using Principal Component Analysis, RA-GPR accurately captures holistic S-parameter dynamics up to 20 GHz. Compared to conventional models, RA-GPR reduces the prediction error from 3.3 dB to 0.5 dB and drastically accelerates evaluation from over 30 minutes in 3D EM simulation to merely 0.024 seconds.
Design Optimization for 224Gbs Channel in High Density Interconnect Printed Circuit Boards
ABSTRACT. This work presents a scheme to remarkably optimize the insertion loss across the 224 Gb/s High Density Interconnect Printed Circuit Board channel by implementing an novel reference plane design and adjusting differential pair pitch in the main routing area. The effects of different pitch transition methods with the new reference plane layout are also analyzed.
Data-Efficient Training of Machine Learning Surrogates with Application to a Memory Interposer
ABSTRACT. This paper develops an active learning technique
for the surrogate modeling of multi-output system responses
based on Gaussian processes. In contrast to standard, datahungry
machine learning approaches, the proposed method
enables a more effective allocation of simulation samples and
achieves higher accuracy than conventional random sampling
strategies for the same computational budget. The methodology is
validated on two signal integrity case studies involving a memory
interposer, focusing on the prediction of the insertion loss across
all transmission paths and the complete scattering response.
Pattern- and Channel-Aware Physics-Assisted Deep Learning for Per-DQ Eye-Metric Estimation in RDL Interposer Channels
ABSTRACT. Pattern- and channel-aware physics-assisted deep learning estimates per-DQ eye metrics in RDL interposer channels. FiLM conditioning fuses pattern and SBR/VTF channel features, achieving 94–97% accuracy under unseen conditions while accelerating estimation 480× over transient simulation.
Machine-Learning-Based Multi-Channel Optical I/O Analysis for Co-Packaged Optics
ABSTRACT. This paper demonstrates a machine-learning-based methodology for large-scale multi-channel optical I/O analysis for co-packaged optics. A transformer model trained from multi-scale single-channel simulations is combined with geometric Monte Carlo analysis to evaluate channel-dependent performance across dense optical arrays. The model achieves 99% accuracy within a 0.05 dB loss threshold and completes more than 108 channel predictions in approximately 10s. This methodology enables scalable yield analysis and design evaluation for next-generation high-density optical interconnects.
Physics-Derived Machine Learning Surrogate Modeling for Rapid LPDDR Signal Integrity Triage
ABSTRACT. LPDDR signal-integrity closure requires repeated eye and jitter evaluation across package, PCB, DRAM, routing, and design-corner variations. High-fidelity statistical simulation provides signoff accuracy but can require several days per iteration. This paper presents a physics-derived machine-learning surrogate for rapid LPDDR signal-integrity triage using cascaded N-port S-parameter data. Loss, reflection, group-delay, and crosstalk features are extracted from each data-channel path to predict eye-area and jitter metrics at a bit-error rate of 10⁻¹⁵. XGBoost and a multilayer perceptron are compared, with XGBoost selected for its higher accuracy and faster training. Controlled two-dimensional interconnect variants increase the labeled dataset by approximately 3.4 times on one platform. Across two LPDDR platforms, XGBoost achieves R-squared values from 0.83 to 0.99 and average relative accuracy above 95 percent. Worst-channel recall reaches 93.3 percent, while entirely unseen datasets achieve average relative accuracy above 94 percent. The flow reduces the runtime of one evaluation iteration from approximately three days to two minutes, enabling targeted simulation of low-margin channels.
Active-Learning-Based TX FIR Tuning for High-Speed Links Under Dynamic Thermal Ramping
ABSTRACT. We present a thermal-aware active-learning framework for TX FIR tuning under dynamic temperature sweeps. Trajectory surrogates and pairwise error ranking guide a four-role candidate acquisition strategy, converging to thermally robust sweet spots within the top 4.1% of DTS-ranked configurations using 4.2% characterization cost.
Active Immersion Cooling Fixture for High Power Testing of Planar Microwave Transmission Lines
ABSTRACT. This paper introduces a test fixture for active liquid cooling of two-port planar microwave transmission lines under high power operating conditions. A rectangular waveguide (RWG) feed followed by a non-contacting transition to the planar line is used to avoid end-launch connectors and other potentially power-limited connections. The fixture is fabricated using fused deposition modeling (FDM), and its interior RWG walls are coated with silver paint. The cooling test fixture is demonstrated herein using a through coplanar waveguide (CPW) fully immersed in a liquid coolant and measured in the 3 − 6 GHz frequency range. Non-intrusive distributed temperature measurements using an optical fiber sensor are also presented for the CPW line operated at high power. The entire structure has a minimum insertion loss of 1.6 dB at 3.6 GHz and 2.2 dB at 4.4 GHz with and without the liquid coolant, respectively. The fixture itself has minimal loss and accounts for less than 1.4 dB of the total insertion loss without the coolant, which confirms its viability for high-power testing.
AI-Assisted Chiplet Design Space Exploration with Physics-Aware Reduced-Order Modeling
ABSTRACT. Chiplet-based design introduces a combinatorial explosion due to complex interactions among chiplet selection, layout, and multi-physics constraints. Conventional approaches are limited by heuristic pruning or costly simulations. This paper proposes an AI-assisted chiplet co-design framework integrating design space reduction with physics-aware evaluation. The method combines retrieval-augmented generation for knowledge-guided pruning and interface-aware reduced-order modeling for efficient multiphysics analysis. The proposed approach reduces the design space from O(106) to O(103) candidates while maintaining physically valid solutions, and enables systematic multi-objective evaluation of electrical, thermal, mechanical, and cost trade-offs. This enables scalable and physics-consistent design exploration.
Signal and Power Integrity Analysis of 32 Gbps Die-to-Die Interconnects for Chiplet
ABSTRACT. For 32 Gbps die-to-die(D2D) interconnects in chiplet systems, a high-speed channel design based on silicon interposer is presented. Comprehensive signal and power integrity (SIPI) evaluations are carried out. The individual impacts of channel loss and crosstalk on signal quality are investigated, and targeted routing design guidelines are proposed. The power supply induced jitter (PSIJ) introduced by voltage ripples of VDDIO and VDD_core power domains is investigated and the power regulation performance enhancement enabled by deep trench capacitors (DTC) is also verified.
SI/PI Trade-off Analysis of UCIe-A Interfaces Considering Power/Ground-to-Signal Bump Ratios
ABSTRACT. In this paper, the signal integrity (SI)/power integrity (PI) trade-off associated with the amount of allocated power/ground (P/G) bumps is investigated for a 32 Gb/s Universal Chiplet Interconnect Express-Advanced (UCIe-A) interface. This trade-off occurs because additional P/G bumps lower power distribution network (PDN) impedance and improve PI, but the resulting increase in channel length raises channel loss and degrades SI. To quantify this trade-off, four bump map candidates are comparatively evaluated by varying the power/ground-to-signal (P/G-to-S) bump ratio. Eye-diagram analysis is used for SI evaluation, while PDN impedance and simultaneous switching noise (SSN) analyses are used for PI evaluation. The results show opposite SI and PI trends as the P/G-to-S ratio varies and identify the lowest-ratio candidate as the most favorable design in this work because it satisfies the UCIe-A eye-mask and I/O supply-noise requirements with the smallest PHY footprint. This work provides practical insight into UCIe-A bump map design by showing that the amount of P/G bumps should be determined based on the SI/PI trade-off.
Design and Analysis of HBM-HBF-NMC Architecture for LLM Inference Considering Signal Integrity
ABSTRACT. In this paper, we propose a high bandwidth memory (HBM)-high bandwidth flash (HBF) near-memory computing (NMC) architecture that addresses the bandwidth and capacity demands of long-context large language model (LLM) decoding. The architecture places model weights and the frequently accessed hot key-value (KV) cache in HBM, while the larger cold KV cache is stored in HBF, supporting a batch size of up to 196 at a 128K context within the available memory capacity. By integrating NMC into the base dies of both tiers, data is processed locally through short TSVs, providing high internal bandwidth while limiting transfer to the GPU. Since decoding still requires repeated communication between HBM and HBF, a compact UCIe-A D2D interconnect is implemented as a 2 mm interposer channel with stacked signal layers and meshed ground shielding. The simulated eye diagrams satisfy the applied UCIe 3.0 masks, with openings of 166 mV/0.84 UI at 32 Gbps and 66 mV/0.76 UI at 64 Gbps, confirming the signal-integrity feasibility of the proposed compact HBM-HBF interconnect.
Package-level thermal optimization for high-power-density AI accelerators
ABSTRACT. The increasing power demands of AI applications necessitate advanced thermal management to maintain IC package temperatures within safe operating limits. This paper proposes modifications to conventional package architectures to eliminate thermal resistance bottlenecks. Simulation results indicate that the proposed design can achieve a thermal design power (TDP) density of 6–10 W/mm², representing at least a 3× improvement over existing state-of-the-art thermal management solutions.
Influence of interface roughness between Ag sintered bonding layer and Cu on thermal stress in SiC power chip system with AMB substrate
ABSTRACT. This paper clarifies the influence of interface roughness between Ag sintered bonding layer and Cu on the thermal stress profile in Ag sintered layer (Ag SL) using two-dimensional finite element analysis in a multiphysics solver for a SiC power chip system with an active-metal-brazed (AMB) substrate. We model this interface as a sinusoidal curve, varying the amplitude from 0 to 10 µm and the number of waves per 500 µm. It was found that von Mises stress in Ag SL is concentrated at the valley of the sinusoidal Ag SL–Cu interface. The normal stress at horizontal direction is tensile in Ag SL, and this tensile stress becomes the largest at the valley. Results also clarify that the thermal stress at the upper and lower sides of pores located near the valley shows the highest stress concentration, which is thought to become one of the crack origins in Ag SL.
8533 MT/s LPDDR5X Package Design with Memory-on-Board Topology
ABSTRACT. The growth of technologies requiring high bandwidth solutions has led to the development of increasingly faster memories. This has resulted in the current generation memory, LPDDR5X. However, increasing data rates lead to signal integrity-related challenges at the channel level. To mitigate these effects, memory devices are often integrated within the package due to the shorter trace lengths and reduced discontinuities. However, this increases the system cost and complexity. In this paper, we demonstrate a solution for achieving 8533 MT/s with the memory located on the printed circuit board (PCB). The associated package design challenges are discussed along with the design considerations in bump map, stackup and ball map.
Extension of a 2D CIM-Based Approach for Efficient and Automated Computation of Resistances for DC-Analysis of Multilayer PCBs
ABSTRACT. Efficient 2D CIM-based resistance calculation is extended for multilayer PCBs and used to automatically compute DC-solution for SI/PI simulations. Speed-up of 80x enables generation of large quantities of DC data for fast machine learning applications.
Physics-Aware Reinforcement Learning for Timing-Margin Optimization of High-Speed DDR5 RDIMM
ABSTRACT. We propose a physics-aware reinforcement-learning framework that automatically optimizes CA/CS/CK timing margins of high-speed DDR5 RDIMM. A reflection-wave phase-based surrogate guides a Double-DQN agent, achieving a worst-case gap of 0.0496, 48% better than manual design.