Huawei Kirin 9050 Pro Teardown: LogicFolding, Dual Dies, and a Much Bigger NPU

 

The Huawei Mate 90 Pro Max, featuring Huawei's latest chip, the Kirin 9050 Pro

Geekerwan has given us one of the closest looks yet at the Huawei Kirin 9050 Pro, the chip inside the Mate 90 Pro Max. The channel had a lab slice the SoC open and then rebuilt its internal structure in 3D, offering a rare view of how Huawei’s first LogicFolding chip actually works. If you want to see the full breakdown, Geekerwan’s teardown video is worth watching here.

Two dies, bonded together

A normal SoC typically has a substrate, metal layers, and a single logic layer on top. The Kirin 9050 Pro takes a different route. It uses two dies stacked and joined by copper-to-copper hybrid bonding, which directly links their metal layers.

According to the teardown, the compute die is made on SMIC N+3, which is roughly a notch better than 6 nm-class nodes from other foundries. The secondary die, which holds cache, I/O interfaces, and PLLs, is made on SMIC N+2, roughly comparable to TSMC N6. Neither Huawei nor SMIC has confirmed those node details directly, so treat them as industry estimates rather than official specifications.

Power and signals reach both dies through through-silicon vias, or TSVs. Geekerwan counts around 80,000 of them connecting the chip to the substrate. They pass through the lower die to reach the upper die, and their keep-out zones take up about 8% of the bottom die’s usable area.

The idea is not unlike AMD’s 3D V-Cache, where cache sits on a separate bonded die. But here, the hybrid bond connections are much denser. That should shorten the path between execution units and cache, which in turn should reduce latency.

LogicFolding is not just block stacking

Huawei’s LogicFolding approach does not simply put whole blocks on one die or the other. Each major block spans both dies, and that is visible in Kurnal’s die shot. In the prime core, most execution units sit on the top die, while the L1 and L2 caches sit on the bottom die directly beneath the load/store units. The same pattern applies to the GPU, ISP, NPU, and modem.

The shorter interconnects should cut latency and allow lower voltage at the same clock speed. Cache, I/O interfaces, and PLLs go on the bottom die because they run cooler and are less process-sensitive. Compute blocks go on the top die, where heat dissipation is better. You can see more detail in Kurnal’s die shot on X.

A revamped NPU

The NPU is the biggest upgrade in the Kirin 9050 Pro. Its combined area across both dies is 150% larger than the 9030 Pro’s. Geekerwan measured 68 TOPS of INT8 throughput and a threefold prefill improvement with a 3B-parameter model.

Phones with this chip support an on-device 30B-A2B mixture-of-experts model, meaning 30B parameters total with 2B active at a time. It handles offline photo organization and edit suggestions, and it powers Huawei’s AI assistant, paired with a 6B multimodal model for edits. For comparison, Apple uses a 20B-A4B on-device model, which the company has documented in its third-generation Apple Foundation Models research.

All Kirin 9050 Pro phones come with 16 GB of RAM, so Huawei can keep a larger model in fast memory. Even so, the NPU still trails the latest competitors. That is likely why Huawei only activates a smaller share of the parameters compared to Apple’s solution.

Impressive gaming results

In gaming tests, Geekerwan says the Huawei Mate 90 Pro Max’s performance in Genshin Impact beats that of some Snapdragon 8 Elite Gen 5 phones. It trails phones powered by the Snapdragon 8 Elite Extreme Gen 6 and Apple A19 Pro only slightly. In Wuthering Waves, the chip ranks ahead of Dimensity 9500 phones. In Neverness to Everness, it nearly matches Dimensity 9500 and Snapdragon 8 Elite Android phones, although the Android devices run at a slightly higher resolution.

Those results come with an important caveat. Performance depends heavily on how well games are ported for HarmonyOS. Titles that run through a translation layer, such as Arknights: Endfield, are less efficient and deliver lower frame rates.

What this means for Huawei

The Kirin 9050 Pro is not just another flagship SoC update. It shows Huawei experimenting with advanced packaging in a way that could reshape how future chips balance compute, cache, I/O, and heat. LogicFolding may be Huawei’s answer to the limits of conventional monolithic designs, especially as process node improvements become harder and more expensive to achieve.

The real test will be whether developers can take advantage of the design. If HarmonyOS games and AI apps are built natively, the chip’s shorter interconnects and larger NPU could deliver more consistent gains. If they rely on translation layers, some of that potential may stay locked away.

For now, the Kirin 9050 Pro looks like a bold step. It is not yet a clean win over the best Qualcomm and Apple silicon, but it is far closer than many expected. And with LogicFolding, Huawei has made it clear that it is willing to rethink the chip itself to get there.

Tags: