Multi-Die Integration and Advanced Packaging Trends
The semiconductor industry's favorite mantra for decades was simple: shrink the transistor. Moore's Law carried us from micron-scale features to nanometers, and each process node brought more transistors, better performance, and lower cost per function. That formula isn't dead, but it's getting exponentially more expensive. A leading-edge 3nm mask set costs $300-500 million. Not every design can justify that. Multi-die integration — splitting a large chip into smaller "chiplets" connected through advanced packaging — is the industry's answer to the economic and physical limits of monolithic scaling.
Why Chiplets Make Sense
The math is straightforward. Die yield (the percentage of working chips per wafer) drops exponentially with die area. A 100 mm2 die on TSMC's N5 process might yield 85-90%. A 400 mm2 die on the same process might yield only 50-60%. By splitting that 400 mm2 design into four 100 mm2 chiplets and connecting them in a package, you've dramatically improved your effective yield and reduced waste.
But yield is just the start. Chiplets also enable:
- Process mixing — put the compute cores on 3nm while the I/O and analog blocks stay on cheaper 12nm or 22nm. Not every function benefits from the latest node.
- Design reuse — the same memory chiplet or I/O chiplet can appear in multiple products, amortizing development cost across a product line.
- Heterogeneous integration — combine silicon with different materials (III-V for RF, SiC for power) in a single package.
- Disaggregation of supply chain risk — different chiplets can be manufactured at different fabs.
AMD was the company that really proved chiplets could work at scale in a commercial product. Their EPYC server processors, starting with the first-generation Rome in 2019, use multiple compute chiplets (CCDs) connected to a central I/O die (IOD). The Zen 4-based Genoa processor has up to 12 CCDs (each with 8 cores) plus an IOD, giving 96 cores in a single socket. Each CCD is manufactured on TSMC 5nm for compute density, while the IOD uses a more mature process (TSMC 6nm) since it's mostly I/O and memory controllers.
Die-to-Die Interconnect Technologies
The connection between chiplets is where the real engineering challenge lies. The interconnect needs to be fast, dense, energy-efficient, and — ideally — standardized so chiplets from different vendors can work together.
Organic substrate interconnects are the simplest: chiplets are placed on a standard organic package substrate and connected through traces in the substrate layers. AMD's first EPYC chips used Infinity Fabric links routed through the substrate. Bandwidth density is limited to maybe 2-5 GB/s/mm of edge bandwidth because organic substrate features are relatively coarse (5-10 um line/space).
We covered a related topic in Photonic Computing Explained: Silicon Photonics, Optical Int.
Silicon interposer (2.5D) places chiplets side by side on a thin piece of silicon that acts as an interconnect layer. The silicon can have very fine features (0.5-2 um line/space), enabling much higher bandwidth density. AMD's Instinct MI300X uses TSMC's CoWoS-S (Chip-on-Wafer-on-Substrate) with a silicon interposer to connect 8 HBM3 stacks to the compute dies. The bandwidth between the compute die and HBM reaches over 5 TB/s.
The downside of silicon interposers is cost — you're essentially making another large silicon die just for interconnect. TSMC's CoWoS capacity has been a major bottleneck, with demand from Nvidia, AMD, and Broadcom exceeding supply throughout 2023-2024.
Silicon bridge technology embeds small pieces of silicon into the organic substrate, only where high-bandwidth chip-to-chip connections are needed. Intel's EMIB (Embedded Multi-die Interconnect Bridge) is the prime example. It's cheaper than a full interposer because the silicon area is much smaller. Intel uses EMIB in their Ponte Vecchio GPU (with 47 active tiles and 5 bridge types) and in the upcoming Falcon Shores.
Hybrid bonding enables direct copper-to-copper and oxide-to-oxide bonding between two silicon surfaces at sub-micron pitch, without solder bumps. This is a 3D stacking technology where chiplets are literally stacked face-to-face or face-to-back. AMD's 3D V-Cache technology (used in the Ryzen 7 5800X3D and subsequent processors) bonds a 64MB SRAM cache die directly on top of the CCD using TSMC's SoIC hybrid bonding at roughly 9 um pitch.
Hybrid bonding can achieve extraordinary connection densities — over 10,000 connections per mm2, compared to about 400 for traditional micro-bumps. This enables bandwidth densities that rival on-die interconnects, blurring the line between multi-die and monolithic design.
For a related perspective, see Emerging Memory Technologies: MRAM, ReRAM, and Processing-in.
The UCIe Standard
Universal Chiplet Interconnect Express (UCIe) is an open standard that aims to make chiplets from different vendors interoperable. Version 1.0 was released in 2022, backed by Intel, AMD, ARM, TSMC, Samsung, ASE, and many others.
UCIe defines:
- Physical layer — specifying bump pitch (either 25 um for advanced packaging or 100+ um for standard packaging), signaling (single-ended or differential), and speed (up to 32 GT/s)
- Die-to-die adapter layer — handling link training, framing, and flow control
- Protocol layer — supporting CXL and PCIe protocols natively, with streaming and other protocols possible
The advanced packaging flavor of UCIe achieves 28 GB/s/mm in each direction, with an energy efficiency target of 0.5 pJ/bit. For comparison, an off-chip PCIe Gen5 link consumes about 5-10 pJ/bit — an order of magnitude more.
I think UCIe is important but I'm realistic about its near-term impact. Creating truly mix-and-match chiplet ecosystems requires solving not just the physical interconnect but also coherency protocols, power management, security, and testing across vendor boundaries. That's going to take years. In the near term, UCIe will mostly be used within a single company's chiplet portfolio, ensuring internal design consistency and future compatibility.
Advanced Packaging Technologies at the Foundries
TSMC, Intel, and Samsung are all investing heavily in advanced packaging:
For a related perspective, see Data Center Accelerators: GPUs, FPGAs, and Custom ASICs for .
TSMC offers CoWoS-S (silicon interposer, used by Nvidia's H100/H200/B100/B200), CoWoS-L (local silicon interconnect + organic redistribution, for larger interposer areas), InFO (integrated fan-out for mobile SoCs), and SoIC (3D stacking with hybrid bonding). They're building massive new packaging capacity in Taiwan and Japan.
Intel has EMIB (silicon bridge), Foveros (3D stacking with micro-bumps or hybrid bonding), and is combining both in what they call "co-EMIB + Foveros" designs. Ponte Vecchio was the first product to use this combination. Intel's packaging R&D in Chandler, Arizona is among the most advanced in the world.
Samsung offers I-Cube (silicon interposer, similar to CoWoS), X-Cube (3D stacking), and FOWLP (fan-out wafer-level packaging). Samsung has been less visible in the high-end AI accelerator packaging space but is working to catch up.
The Business Implications
Advanced packaging is reshaping the semiconductor value chain. OSATs (Outsourced Semiconductor Assembly and Test) like ASE, Amkor, and JCET are investing billions in advanced packaging lines. Meanwhile, TSMC is competing directly with OSATs by offering packaging as part of their turnkey service, which creates interesting competitive dynamics.
The packaging step is also becoming a larger fraction of total chip cost. For a high-end AI accelerator using CoWoS with HBM, the packaging and memory can easily cost more than the logic die itself. Some estimates put CoWoS packaging cost at $1,000-2,000 per chip for large configurations.
Looking ahead, I expect multi-die integration to become the default for any high-performance chip within the next 5 years. Monolithic designs above 200-300 mm2 will become the exception rather than the rule. The chiplet ecosystem will develop gradually — not the idealized vision of mix-and-match dies from a catalog, but a practical reality where companies design families of chiplets that they recombine across product lines. The real revolution isn't any single technology; it's the shift in design philosophy from "one big chip" to "a system of chips, integrated as one."