Heat per millimetre
Checked on three cards (26 September 2026). This page's claims were re-measured under a pre-registered plan, in six new passes on each of aifoundry2, aifoundry3 and aifoundry1 card 1. Of 17 claims tested here, this page counts 8 held, 1 corrected and 8 differ by card; the hub’s scoreboard, 7 “proven on the cards”, 1 “a test behind it failed”, 8 “differs by card” and 1 “within noise”. The correction is all ones' excess over random data on board power, which no card resolved from zero; where the cards differ, the text says so. Every figure those passes cover now comes from them; the record is in docs/reports/data/2026-09-25-claims-v3.
Bill Dally's rule of thumb for on-chip communication is "~100 fJ/b-mm": the heat of carrying one bit one millimetre across a chip. We measured it on the ET-SoC-1's mesh (the 6 × 6 grid of shires plus a column of memory shires down each side: 8 × 6 stops) by reading data from a scratchpad 0 to 6 hops away, with the bits on the wires chosen, in three runs (the last with six passes on each of three cards), and two meters: board power and the mesh's own supply rail. One hop is of silicon. At the mesh's 0.485 V, with no other traffic on its links, a random bit costs per millimetre on the mesh rail (), of it depending on the data; board power, which also carries the regulator's loss, says (). On a busy mesh, where flows share links, the same bit costs : .
What costs energy is surprising: not only bits that differ between consecutive flits, but the ones carried. A stream of all ones, which never changes from one flit to the next, costs as much per hop as random data (), though at one hop it costs .
Dally states no voltage and no process: the same figure has served a 28 nm slide and a 14 nm cost model (§7). Scaled to 0.9 V (the voltage of the 40 nm table he published with Keckler in 2011, its likeliest source; §1), the mesh rail's data-dependent cost is , so the rule of thumb holds, to within its own vagueness. The gap of about two below it at the chip's own 0.485 V is expected, and mostly voltage: for a network-on-chip on a 5 nm chip Dally himself gives "~50fJ/bit-mm", his group's scaling takes a 28 nm wire to 7 nm at of its energy, most of it through the lower supply, and this mesh runs below even 7 nm's nominal one. Read instead at the "~0.5V" his 2023 talk sets beside it, the figure is what this mesh spends on the data.
Terms used on this page
The ET-SoC-1's cores are minions (two hardware threads, harts, each), 32 to a shire; part of each shire's 4 MB of SRAM is a 2.5 MB scratchpad that any shire can address, and a tensor load is a minion's bulk load of up to 1 KB. Shires talk over a mesh network-on-chip (NoC), one mesh stop per shire; a hop is one step between neighbouring stops, and a flit is the unit the mesh moves as a whole, here at least one 64-byte line (§9). Both meters are in the board's power-management controller (PMIC): board power is its reading of the card's 12 V input, the mesh rail its reading of the regulator that feeds the mesh, and the service processor (SP), the chip's management core, reports them. a2, a3 and a1c1 are the three cards: those in the lab machines aifoundry2 and aifoundry3, and card 1 of aifoundry1; more in the hub's glossary.
1. The number, and what it hides
Dally's keynote to the Stanford AHA retreat (2023) splits the energy of computing, E = ½CV², into three kinds of
capacitance: "Communication (~100fJ/b-mm on-chip) · Memory (~50fJ/b for small RAM) · Operations (~1fJ/b for add)"; at
Hot Chips the same week he put communication at "about ... 100 femtojoules per bit millimeter". The same 100 fJ/bit-mm
is in his CACM 2020 paper with Turakhia and Han. None of them gives a process, a voltage or a data activity.
CACM 2020 states it twice, in a cost model whose arithmetic and local memory are "in 14 nm", and the 2023 talk's
conclusion slide sets it beside "Reduce V until it gets too slow (~0.5V)" without saying it applies there. A wire's energy
goes as CV², so the voltage alone moves the number by 3.4× between 0.9 V and 0.485 V, and "per bit" can mean per bit
sent (random data flips half its bits), per transition, or a full charge of the wire for every bit. Its likeliest
ancestors — an inference; no source states it — are Dally's own earlier figures: 110 fJ/bit-mm at 32 nm and 0.6 V in the
2008 DARPA exascale study, whose authors include him; a keynote slide labelled 28 nm that he showed from 2010 to 2017;
and 310 pJ for 256 random bits over 10 mm at 40 nm and 0.9 V (Keckler, Dally and colleagues, IEEE Micro 2011), 121 fJ
per bit·mm. His later numbers are 68 fJ at a projected 10 nm and 0.7 V, 20–40 fJ/bit-mm "in present day chips" (VLSI
Symposium 2018, a paper set in 16 nm), 30 in 2022, and for the wires of a network-on-chip on "a typical chip say five
nanometer chip today", "~50fJ/bit-mm" (his NOCS keynote, 2022). Which process the figure is for, and whether a gap of two
below it is expected: §7. The sources and quotes are in
docs/reports/data/2026-09-24-wire-energy/research/DALLY-NODES.md and SYNTHESIS.md beside it.
2. How long is a hop?
What the die diagram shows
Esperanto publishes the die area (570 mm²) and a die plot, but not the dimensions. We measured the plot's tile period in pixels — the shire outlines and an autocorrelation of the image agree — and scaled it to the published area: the minion-shire tiles are 3.73 mm east–west and 3.70 mm north–south, square to within 1%, so a hop in either direction is about 3.72 mm ( depending on whether the 570 mm² includes the drawn frame). The shire grid spans 86% of the die's width; the rest is the memory shires and their DRAM PHYs down each side. A pair of shires d hops apart on the logical map (marty1885's shire coordinates, which the on-chip communication report confirmed from latency on all three cards, aifoundry2, aifoundry3 and aifoundry1 card 1) crosses d links each way: on each card round-trip latency is linear in Manhattan distance, to within 1.1–1.4 cycles, over all 496 pairs, in every one of three passes. "Per mm" below means per mm of mesh travel, one router and one link per 3.72 mm — per mm of displacement; the metal actually routed can only be longer, so per mm of wire the cost would be lower.
3. The experiment
Method in full
- Traffic. Hart 0 of every minion streams 1 KB tensor loads, two in flight, from the scratchpad of a shire exactly d hops away (d = 1, 2, 3, 4, 6), at most two readers per target; d = 0 reads the shire's own scratchpad and never touches the mesh. The energy per payload byte is fitted against d; the slope is the cost of one hop, and everything that happens once per byte — the SRAM read, the L1 write, the request, leaving and entering a shire — stays in the intercept.
- The bits on the wires are chosen. Before each configuration every scratchpad is filled with a known image (the store kernel's output checked byte for byte on a DRAM slice, on all three cards; the tool cannot read a scratchpad back): bits independently 1 with probability P (so two consecutive data flits differ in a bit with probability 2P(1−P): a difference rate between flits, not necessarily the transition rate on the wires), blocks of N bytes alternately all-zero and all-one, and one random 64 B line frozen everywhere (half ones, no two flits differ). The second run makes every line on the chip unique and P and 1−P exact bitwise complements, and adds reader–target pairs chosen so that no two flows share a link.
- Two meters. Board power (10 mW steps; under the 10 Hz sampler used here its value changes about every 156 ms on aifoundry2 and 158 ms on aifoundry1 card 1, and every 263 ms on aifoundry3, where a light poller sees it change every 223–224 ms: the three-card check resolves the sampler's lengthening of that refresh at 99% only there, while the service processor's own pass, timed in its trace, lengthens under the sampler on every card; the observability report has the meter chain) and the mesh rail, which feeds the routers and links at 0.485 V and 400 MHz. For board power each 3 s burst is measured against idle on both sides and corrected for the extra leakage of a burst that ran warmer. The mesh rail is read over the burst's last 0.6 s against the idle before it, without a leakage correction (about 2%, roughly cancelling a 2.5% under-correction of the meter's 1 s running average). A burst is dropped if the clock left 600 MHz or the meter itself was starved (§10).
- Replication. Three runs, every configuration in shuffled order, the bars the range over every pass of every card. The first two (24 September) ran each configuration three times on each of two cards (aifoundry2, aifoundry3): first run 82 configurations, 492 bursts; second run 44 configurations, 264 bursts; dropped. The third (the three-card check, 25–26 September) ran 28 of the second run's configurations — all pairs at P = 0, ½ and 1 over 0–6 hops, link-disjoint pairs at P = 0 and ½ over 1–5 hops — six times on each of aifoundry2, aifoundry3 and aifoundry1 card 1: . Every figure those configurations give comes from the third run; the split into ones and differences (§5), the block patterns (§9) and the x/y pairs come from the first two, which alone ran them. A figure given without a card holds on every card that measured it; where the cards differ, the text names them, and what these runs leave open is listed in §10.
4. Energy grows with every hop
How to read the chart
Density of ones P runs from the faintest line (all zeros) to the strongest (all ones); bars are the range over every pass of every card. Hover, tap or tab to a point.
How to read the chart
5. Ones cost energy, not just flips
How to read the chart
What a, b and s₀ mean, and the fitted coefficients as a table
Here a is the energy per bit that differs from the same bit of the previous flit, b the energy per one-bit carried (a 1, whether it changed or not), and s0 the part that does not depend on the data (clocking, headers, requests).
What any bit pattern would cost
Try it: the interactive plane
The fitted model spans a plane of patterns: the density of ones P across, the fraction of bits that differ from the previous flit t up. Independent random bits sit on the dashed curve; the frozen line and the block patterns of §9 leave it. Drag the cursor, use the sliders or pick a pattern.
6. Sharing a link costs energy
How to read the chart
What the link-sharing map shows
7. Against the rule of thumb
How the scaling was done
Which process is Dally's figure for, and is a 2× gap expected?
Each figure is moved to 0.485 V as CV² at constant capacitance: through the capacitance the
source gives where it gives one (the 2008 study, Gebhart 2011, VLSI 2018), from its own voltage where it states one, and
otherwise from each of the three voltages Dally's own sources name (0.9 V, his 40 nm table; 0.7 V, his 10 nm projection;
"~0.5V today", 2023). A random bit switches its wire half the time, at ½CV² a switch, so it costs a quarter of a
figure that counts a full charge, CV², for every bit (the 2007, 2008 and SatIn rows). The quotes, with their pages and
slides, are in docs/reports/data/2026-09-24-wire-energy/research/DALLY-NODES.md, and
research/lit/dally_gap.py there prints the arithmetic.
8. What it means in practice
Try it: price your own transfer
The same comparison for any route: pick where the data sits and who reads it, and the readout sets the hops against reading the same bytes from DRAM or from the reader's own scratchpad.
9. Lanes and flits
These patterns test the model of §5 with data whose ones and differences do not move together. Every block pattern has half its bits ones, so the ones term is the same for all of them; any difference is transitions.
Blocks of 16 to 128 bytes cost the same, and 256-byte blocks cost more. That is what the shire's four mesh lanes predict: a line goes to lane PA[7:6] (bits 7–6 of the physical address), so on one lane consecutive lines are i and i+4 — 256 bytes apart. With 16–128-byte blocks lines i and i+4 hold the same value and nothing flips; with 256-byte blocks they are opposite and every bit flips. It also says a flit carries at least a 64-byte line at a time: with 32-byte flits the 32-byte pattern would flip on every flit, and with 16-byte flits the 16- and 32-byte patterns would too, on every or every other flit; neither costs more.
More on the 256-byte pattern
10. What this cannot tell
Caveats in full
- Router against wire. Every hop is one router plus one pitch of link, and the x and y pitches are equal to within 1% on the die plot, so the experiment cannot split a hop's energy between the router's flops and crossbar and the wire itself.
- Board against rail. Board power is 12 V input power: it includes the regulator's loss and anything else that moves with the traffic. The regression in the observability report (§4.2) finds part of what the minion rail delivers lost between the 12 V input and the core, consistent with a regulator's loss, as far as the rails' own meters can be trusted. The mesh rail is the die's own supply but its meter's gain has not been checked independently. The truth is between the two.
- Fixed per-hop cost. The part that does not depend on the data (clocking, headers, the request travelling the other way) is measured, but it is confounded with time: bandwidth per reader falls with distance, so part of it may be cost per second rather than per hop. The data-dependent part is immune to per-second costs that do not depend on the data (P = ½ and P = 0 run at the same bandwidth); a per-cycle cost of ones held in the mesh would still look per-hop.
- The meter can be starved. Other workloads that starve the meter, a sampler crash that stopped it between the two runs, and a stray write the card never reported (no measurement here comes from those launches) are in the observability report (§4.1; the stray write in §3).
- Voltage scaling assumes full-swing CMOS at constant capacitance; whether the links are low-swing is not known.
Gate capacitance is somewhat lower near threshold, so constant-C scaling slightly understates the 0.9 V figure for
full-swing links; low-swing links would make it overstate. Experiment E54 (NV) is registered to measure it: the mesh rail
at 485, 540 and 600 mV, with its predictions frozen on 28 September (
tools/claims-v3/nv/; the table in §7). A read-only probe of aifoundry3 at 22:54 PDT that day found the rail set to 485 mV and reading 484 mV on the die, with the mesh at 400 MHz. The first voltage write waits for the owner's go-ahead: it changes a shared card, and the service-processor bootloader these cards run retries a failed regulator write without end (fixed upstream in et-platform7c6049087, 10 October 2024), so a failed write would end in the 10 s watchdog resetting the card. Development is to run on aifoundry3, and validation on aifoundry2 after its DVFS validation ends.
11. Data and tools
Reproduce this
- Runs and analysis:
workloads/enercat/run_wire.py(the first two runs),tools/claims-v3/wire/(the third:block.sh, andreduce.pyfor the check's verdicts),workloads/enercat/analyze_wire.py,workloads/enercat/analyze_wire_v3.py(the same reduction over the third run's passes) andtools/ettelem/build_wire_report.py; the fill modeststore_rawandtstore_uniqand the patternsbern:P,alt:N,uq:Pandfrzare inworkloads/enercat/host/main.cpp. The four commands that rebuild the two analyses, this page's data and this page from the raw data are in the docstring oftools/ettelem/build_wire_report.py. - Raw data: the first run in
docs/reports/data/2026-09-24-wire-aifoundry2/anddocs/reports/data/2026-09-24-wire-aifoundry3/, the second indocs/reports/data/2026-09-24-wire2-aifoundry2/anddocs/reports/data/2026-09-24-wire2-aifoundry3/, the third indocs/reports/data/2026-09-25-claims-v3/raw/<card>/wire/p1–p6(its verdicts inresults/wire.jsonthere). The analyses aredocs/reports/data/2026-09-24-wire-energy/wire.json(the first two runs) andwire3.json(the third), and the page draws its charts and computed figures fromdocs/reports/data/2026-09-24-wire-energy/report.json. - Research notes behind the conversion to millimetres (die geometry, the NoC, the literature), gathered by AI agents
run by the author:
docs/reports/data/2026-09-24-wire-energy/research/SYNTHESIS.md; which process Dally's figure is for, with every statement of it we found and its page or slide:research/DALLY-NODES.md(the arithmetic inresearch/lit/dally_gap.py). An independent re-analysis and adversarial review of this page, also by AI agents:docs/reports/data/2026-09-24-wire-energy/review/VERDICT.md; the measurements reproduced, and its corrections are applied here.
Version history
Versions. 24 September 2026: first published, and corrected the same day after a review.
25 September: after a second review, the per-second bound on the fixed per-hop cost given per meter and the block
patterns placed below the two-term model; then version 3, every claim checked against both cards' passes (unresolved
differences no longer findings; the 16–128 B blocks' shortfall dropped; the regulator's loss no longer said to grow with
load). 26 September (version 4): the three-card check's third run is the basis of every figure it covers; all ones'
excess over random data on board power is not resolved from zero on any card. 27 September: the lede's contention figure
compares like with like over one to four hops (it had mixed two bases); the regulator's loss given per card. 28 September:
the review's cuts, and a chart of what each further hop adds beside the link-sharing share (§4); later that day, at the
owner's request, which process Dally's figure is for and whether a gap of two below it is expected (§7), with his
figures from 2008 to 2023. 29 September: the voltage test of §7's first explanation registered as E54, with its frozen
predictions (§7, §10); the route map (§6) and the transfer calculator (§8) draw read data y first, the order E56 measured
on two cards, with the analysis's x first a choice on the map. Every earlier
wording is in the repository's history of docs/reports/sources/heat-per-mm.*.
12. Related reports
- The energy manual, §4.3 — the first wire measurement, from the energy catalogue; its fit over more hops reads lower.
- The energy manual, §5 — what messages between cores and shires cost, from a pair of minions to rings across the mesh.
- Hand it to the next shire — a computation that hands its intermediate to another shire's scratchpad instead of DRAM: 12× faster at a thirteenth of the energy, with bandwidth within a quarter over the ring offsets tried.
- On-chip communication — the shire map behind these hop counts, and the first per-hop estimate, from messages around rings.
- Anatomy of a memory access — each level’s energy split between the rails on three cards, with the energy manual’s per-hop cost beside it.
- Spatial temperature — the 35 temperature sensors on the same grid of 3.72 mm tiles.
- Limits of observability — the meters used here, and what they cannot see.