openTPU Runs Ten Small Models on a Kintex-7 Card, and Its DDR3 Bandwidth Sets Every Token Rate
Hardware / analysis
openTPU Runs Ten Small Models on a Kintex-7 Card, and Its DDR3 Bandwidth Sets Every Token Rate
The AI-built accelerator posted 209 points on Hacker News. Its README numbers, all author-reported, work out to one pass over the weights per token.

The openTPU repository's headline claim, that an AI accelerator was "developed by AI", is the least interesting number on its page. The one worth reading is 17.1 GB/s, the peak bandwidth of the two DDR3-1066 channels on the FPGA card it runs on, because that figure sets every token rate in the project's results table.
openTPU posted 209 points on Hacker News on Oct. 6. Its README describes an Apache 2.0 monorepo holding SystemVerilog RTL, an instruction set, a bit-exact Python simulator, a kernel language and compiler, and host tools. It runs on an Inspur YPCB-00338 PCIe card carrying a Xilinx Kintex-7 xc7k480t, and the card produces the same tokens as the simulator, bit for bit. All performance figures below are the author's own measurements; no independent party has reproduced them in anything we found.
What the card actually delivers
The machine is deliberately plain. A sequencer issues one instruction per cycle to a DMA unit, a four-column systolic matrix unit working on int8 weights, an fp32 vector unit and a quantizer. There is no cache. Every data movement is an instruction, so a trace shows where each cycle went. The design closes timing at 133.33 MHz with a worst negative slack of +0.032 ns, which the README itself calls "only just".

The picture above is Google's TPU 3.0, not openTPU; the open project shares only a name and a systolic matrix unit with that lineage. The lead photograph is likewise a stand-in: a Xilinx Kintex part (an XCKU025) on a PCIe card, a different chip from the xc7k480t openTPU uses. openTPU's own board is a decommissioned data-centre FPGA card, which the submitter described on Hacker News as popular with hobbyists.
| Model (weights) | Decode, tok/s on device | DRAM read while decoding |
|---|---|---|
| LFM2.5-230M (int8) | 59.0 | 14.5 GB/s (85% of peak) |
| Qwen3-0.6B (int8) | 21.6 | 14.4 GB/s (84%) |
| Qwen3.5-2B (int8) | 8.02 | 16.0 GB/s (94%) |
| Phi-4-mini, 3.8B (int8) | 3.99 | 16.0 GB/s (94%) |
- LFM2.5-230M59 tok/s
- Qwen3-0.6B21.6 tok/s
- Qwen3.5-0.8B17.6 tok/s
- Qwen3.5-2B8.02 tok/s
- Phi-4-mini (3.8B)3.99 tok/s
Source: openTPU README, decode on device, 64 greedy tokens after a 512-token prompt, accessed 2026-10-06
The number that matters is bytes per token, not tokens per second
Divide the DRAM bandwidth the card reads by the decode rate and the model's size falls out. For LFM2.5-230M, 14.5 GB/s over 59.0 tok/s is about 0.25 GB per token, which is a 230-million-parameter model at one byte per weight. For Phi-4-mini, 16.0 GB/s over 3.99 tok/s is about 4.0 GB, which is 3.8 billion parameters at one byte each. The match is close enough that the card is doing what the README says: streaming every weight once per token and doing almost nothing else.
That makes the accelerator memory-bound by construction, and it is why 4-bit weights, stored at 4.25 bits including block scales, lift decode speed by 40% on Qwen3.5 and 45% on Qwen3 and LFM2. It also explains why a faster clock would help prefill but not decode, a point the README concedes. The author said the same on Hacker News: for inference "the name of the game is memory bandwidth".
What "developed by AI" does and does not show
The README says the project applies lessons from auto-arch-tournament, an earlier repository by Felipe Sens Bonetto in which an agent proposes changes to a RISC-V core and only designs that beat the current champion merge. In its documented run, 10 of 73 hypotheses were accepted over 9 hours 51 minutes, lifting CoreMark/MHz from 2.23 to 2.91. That repository supports Codex CLI and Claude Code as agents. The openTPU README names no model, and we did not find which one produced its RTL.
Readers on Hacker News split on what that proves. One commenter, mbgerring, argued that "a human prompted an LLM to build a software simulation environment", and that the claim that AI developed its own inference hardware overreaches. Another, scarmig, answered that a prompt can come from a harness or another model, so the human kickoff is beside the point. Both are arguing about attribution; neither disputes the bit-exact match against the simulator.
A third commenter, chris_money202, supplied the scale check: this is the smallest unit of a production AI chip. The card is an FPGA at 133.33 MHz. A shipped accelerator adds dozens of compute units, plus PCIe and Ethernet subsystems that move data between chips, and that periphery is where much of the difficulty sits. The 35B-parameter Qwen3.5-35B-A3B mixture-of-experts model runs at 3.95 tok/s with experts streamed from the host at 1.41 GB/s over PCIe, which is a demonstration of reach, not of throughput.
What would change this read
Two things would. An independent rebuild of the bitstream that reproduces the 59.0 tok/s figure would move the numbers from claim to result. A synthesis run of the RTL on an ASIC flow, with area and power, would show whether the design survives outside a Kintex-7. The README's own to-do list is smaller: close the last few percent of DRAM efficiency, widen timing margin and speed up prefill.
Agent-driven engineering is spreading in software first; REA, which hands coding agents a disassembler, gained 2,963 stars in a day. Chip design is the harder test, and the commercial side of it looks different: we covered Qualcomm's multiyear cross-licence of Huawei's LogicFolding patents as a patent deal, not a tooling one.
Sources
More in Hardware
- 01Tesla's Model 3 and Model Y Can Feed 11.52 kW to a Home, but Only With a Powerwall 3 and a New OrderPowershare Home Backup reaches Tesla's two best sellers on Oct. 6, with a 4.8-fold jump over their old 2.4 kW output and fine print on eligibility.
- 02Replacing 89,000 Satellites Every Five Years Would Mean 17,800 a Year Against 3,400 US Launches in 2025SpaceNews' constellation counts and the AIA-PwC supply-chain study point at nine components with capacity gaps and, for some, three or fewer qualified US suppliers.
- 03Boston Dynamics Hires Rohit Prasad as CEO While Atlas Is Still Listed as 'In Development'Hyundai's reported 30,000-a-year Atlas goal and 300,000-actuator target do not obviously fit together, and the CEO announcement gives no unit counts.
- 04Teradyne Takes an Undisclosed Stake in Bright Machines to Put Robots on AI Server LinesThe Oct. 5 announcement gives no dollar figure and no ownership share, and says the companies will only "evaluate" integrating Teradyne's robots and board testers.