Hardware & Semiconductor

Mining and Cryptocurrency ASICs: Hash Rate Optimization and Power Efficiency

How Mining ASICs Work A cryptocurrency mining ASIC does exactly one thing: compute hash functions as fast as possible while consuming as little power as possibl

By Editorial Team · · 5 min read · 1220 words

How Mining ASICs Work

A cryptocurrency mining ASIC does exactly one thing: compute hash functions as fast as possible while consuming as little power as possible. For Bitcoin, that means SHA-256 double hashing. For Litecoin, it's Scrypt. For Ethereum Classic (since Ethereum moved to proof-of-stake), it's Ethash. The chip contains thousands of identical hash computation cores running in parallel, all doing the same operation on different input nonces.

The simplicity is the point. A general-purpose CPU can compute SHA-256 at maybe 25 MH/s. A GPU might hit 1-2 GH/s. A modern Bitcoin ASIC like Bitmain's Antminer S21 achieves 200 TH/s — that's 200 trillion hashes per second, roughly 100,000× faster than a high-end GPU. The specialization is total and the performance gap is enormous.

Architecture of a Bitcoin Mining Chip

Inside a Bitcoin ASIC, the basic building block is a SHA-256 engine. Each engine implements the 64 rounds of the SHA-256 compression function in hardware. Most designs use a pipelined architecture where a new hash computation enters the pipeline every clock cycle, and results emerge 64 cycles later. With hundreds or thousands of these engines on a single die, you get massive parallel throughput.

The pipeline can be unrolled or folded. A fully unrolled pipeline dedicates separate hardware to each of the 64 rounds, giving maximum throughput at the cost of large silicon area. A folded design reuses the same hardware for multiple rounds, saving area but reducing throughput per unit area. Most production ASICs use full unrolling because hash rate per chip is the primary metric that sells miners.

Bitmain's BM1397 (used in the Antminer S17 series) contained an estimated 144 hash engines. Its successor chips have pushed that number higher as process technology allows more cores to fit on the die. MicroBT's Whatsminer M50 series uses a chip with even higher core density, though exact numbers aren't publicly disclosed.

Nonce Distribution and Communication

Each mining ASIC receives a block header template from the mining controller (usually a small ARM or RISC-V SoC on the miner board). The ASIC then iterates through a portion of the 32-bit nonce space, trying each value to see if the resulting hash meets the difficulty target. When the nonce space is exhausted, the controller provides a new template with different extranonce values.

See also: DDR5 vs LPDDR5X: Memory Architecture, Bandwidth, and Power E.

Communication between the controller and ASIC is typically through a serial interface — SPI or a custom protocol. Some designs chain multiple ASICs together, where one passes work to the next in a daisy-chain configuration. Bitmain uses this approach, stringing dozens of ASICs on a single serial bus per hash board.

Power Efficiency: The Only Metric That Matters

In practice, mining profitability comes down to one number: joules per terahash (J/TH). A miner that does 200 TH/s but consumes 3,500W achieves 17.5 J/TH. If a competitor achieves 15 J/TH, they'll earn more profit per dollar of electricity consumed. Everything else — hash rate, upfront cost, form factor — is secondary.

The evolution of Bitcoin mining efficiency tells the story clearly:

  • 2013 (28nm): ~500 J/TH (Bitmain S1)
  • 2016 (16nm): ~100 J/TH (Antminer S9)
  • 2019 (7nm): ~30 J/TH (Antminer S17 Pro)
  • 2022 (5nm): ~21 J/TH (Antminer S19 XP)
  • 2024 (3nm): ~15 J/TH (Antminer S21)

Each process node jump gives roughly a 30-40% efficiency improvement, which tracks with the power-per-transistor scaling of the foundry process. There's almost no architectural innovation left — the SHA-256 algorithm is fixed and the implementation is already optimal. It's purely a process technology race.

Thermal Design and Board Layout

A single hash board in an Antminer S21 has about 108 ASIC chips dissipating a combined 1,000+ watts. The chips are mounted on both sides of the PCB in some designs, with aluminum or copper heatsinks attached via thermal interface material. Airflow from industrial fans — often running at 5,000+ RPM, generating 75+ dB of noise — provides the cooling.

For a related perspective, see CPU Microarchitecture Explained: Pipeline Stages, Branch Pre.

The thermal design is aggressive. Junction temperatures regularly operate at 85-105°C. This shortens chip lifetime compared to consumer electronics targets, but miners don't care about 10-year reliability — a mining ASIC is economically obsolete in 3-4 years when the next generation of chips arrives. Pushing the thermal limits extracts more hash rate per chip during its profitable window.

Power delivery is another critical board-level challenge. Each ASIC runs at a low voltage (typically 0.3-0.5V for the core logic) but draws significant current. The total board current can exceed 200A at these voltages. Power conversion uses multi-phase buck converters, often custom-designed for the specific current and transient requirements of the hash board.

The ASIC-Resistance Debate

Some cryptocurrency algorithms were specifically designed to be "ASIC-resistant" — making it hard to build specialized hardware that massively outperforms GPUs. Ethereum's original Ethash algorithm used a large memory-bound dataset (the DAG) that required gigabytes of memory bandwidth, which ASICs couldn't easily provide cheaply.

It didn't work forever. Companies like Bitmain and Innosilicon eventually built Ethash ASICs by integrating large amounts of on-chip or on-package SRAM. The performance advantage over GPUs was smaller (maybe 2-5× rather than 100,000×), but it was enough to shift the mining economics.

Honestly, I'd argue that ASIC resistance is a losing battle in the long run. Any fixed algorithm can be optimized in silicon given enough economic incentive. The only question is how long it takes and how large the ASIC advantage ends up being. Algorithms like RandomX (used by Monero) that emphasize random memory access patterns and general-purpose computation have been more successful at maintaining GPU competitiveness, but even they aren't immune to specialized hardware forever.

Related reading: ASML and the EUV Monopoly: Why One Company Controls Advanced.

Manufacturing and Supply Chain

Mining ASIC companies are among the largest customers at leading-edge foundries. Bitmain has historically been one of TSMC's top-10 customers by wafer volume. When Bitcoin prices spike, the demand for modern wafer capacity from mining companies can actually displace other customers — this happened in 2021 when TSMC's 5nm capacity was partially consumed by mining chip orders during a broader chip shortage.

The business model is peculiar. Mining ASIC companies like Bitmain, MicroBT, and Canaan design the chips, manufacture them at TSMC or Samsung, assemble them into mining rigs, and sell the complete machines. They also operate their own mining farms using a portion of their production. This vertical integration means they profit on both the hardware sale and the mining output.

The boom-bust cycle of cryptocurrency prices creates wild swings in demand. During bear markets, second-hand miners flood the market at below-manufacturing cost, and ASIC companies cut wafer orders dramatically. During bull markets, new machines sell at 2-3× their component cost with months-long waitlists. It's probably the most volatile end market in the semiconductor industry.

Beyond SHA-256: Other Mining Algorithms

While Bitcoin mining dominates the ASIC market, there are specialized chips for other algorithms too. Scrypt ASICs (for Litecoin/Dogecoin) integrate relatively large SRAM arrays because Scrypt is memory-hard. Blake2b ASICs exist for Siacoin. KHeavyHash ASICs appeared for Kaspa.

Each algorithm presents different optimization challenges. Memory-hard algorithms require careful memory subsystem design. Algorithms with complex control flow are harder to parallelize efficiently. But the fundamental approach is the same: identify the computational bottleneck, throw silicon at it, and optimize for energy efficiency above all else.

E

Editorial Team

Technical Writer

Expert analysis at Universal Aide.

Editorial Transparency

Our Standards

  • Expert-written technical analysis
  • Fact-checked by domain specialists
  • No sponsored content without disclosure

Content Transparency

  • 100% written by human experts
  • No AI-generated content
  • Advertising content clearly labeled (if any)

Universal Aide is committed to Google Search Essentials, Spam Update 08/2026 compliance, and E-E-A-T principles. Contact: [email protected]