A Hobbyist FPGA Decodes Qwen3.5-9B at 8.18 Tokens a Second, Against a 17.3 Ceiling Set by DDR4
Hardware / explainer
A Hobbyist FPGA Decodes Qwen3.5-9B at 8.18 Tokens a Second, Against a 17.3 Ceiling Set by DDR4
Two MIT-licensed projects run the 9-billion-parameter model on bitcoin-mining FPGA cards. The numbers show where the time goes: memory bandwidth, not logic.

A developer named Corey Hahn has Qwen3.5-9B decoding at 8.18 tokens per second on a single SQRL BCU-1525 FPGA card, a board built for cryptocurrency mining, using custom SystemVerilog and a custom instruction set. A second project, Nero7991's llm.vhdl, runs the same model across two SQRL FK33 cards at about 2.5 tokens per second. Both are MIT-licensed on GitHub, and both are proofs of concept rather than products.
The interesting figure is not 8.18 but the 17.3 next to it. Hahn's repository, fable5_llm, lists a theoretical ceiling of about 17.3 tokens per second, and that ceiling is set by memory, not by the FPGA's logic.

Why a 9B model on this card is capped near 17 tokens a second
Decoding one token means reading every weight once. The BCU-1525 carries a Xilinx xcvu9p FPGA and four DDR4 DIMMs of 4 GiB each, 16 GiB in total. Hahn measured 70.70 GB/s of aggregate bandwidth across the four channels.
The INT4 weight pack for Qwen3.5-9B is 3,902 MiB, or 4.125 bits per weight once the per-128-element scales are counted. Divide 70.70 GB/s by roughly 4.09 GB of weights and you get about 17.3 reads per second. That is the ceiling.
The 8.18 tokens per second the repository reports from the card works out to about 33.5 GB/s of weight traffic, or 47% of the measured bandwidth. The project is leaving more than half the memory system idle on this arithmetic, which is where the next speedups would come from if timing allowed.
- llm.vhdl, two FK33 cards, 75 MHz2.5 tok/s
- fable5_llm, first silicon run, Sept. 97.29 tok/s
- fable5_llm, best measured8.18 tok/s
- fable5_llm, DDR4 bandwidth ceiling17.3 tok/s
Source: fable5_llm and llm.vhdl GitHub repositories, accessed 2026-10-05
What the fable5_llm repository admits
The README is blunt about limits. The context limit is 511 positions, so this is a demonstration of decoding, not a usable chat assistant. The resident bitstream has zero effective timing margin, listed as 0.000 ns. The faster R3 build exists as predictions and has not been loaded on hardware.
The repository gives no power figure and no board price. It needs the specific BCU-1525 and a Vivado licence, and the Qwen weights are downloaded separately from Hugging Face under their own licence.
On correctness, Hahn describes a verification chain from a bf16 checkpoint through an integer reference model, a Python executor and a Verilator testbench to silicon, and reports that all 24 tokens across four prompts matched at the gate test. The project's own rule is "Hardware and reference disagree, the bug is real." Hahn's earlier milestones: 16.1 tokens per second on a 2B model on August 25, then 7.29 on the 9B model on September 9.
The second project: HBM on the FK33
Nero7991's llm.vhdl targets the SQRL FK33, which carries a Xilinx XCVU33P with 8 GiB of HBM, and validates only at 75 MHz on silicon. Running the 9B model as a two-card pipeline "roughly doubles decode throughput" to about 2.5 tokens per second. The 27B target on a single VU35P module is at 46% LUT utilisation but blocked by a missing GTY reference clock and slow weight-load paths, and the tensor-parallel collective subsystem is still skeletal.
Startup Fortune, which wrote up both projects, puts the FK33's HBM2 bandwidth at around 400 GB/s and says the cards sell for roughly $300 secondhand on eBay, against $700 to over $1,000 new. That bandwidth figure comes from the article rather than either repository. On the same arithmetic as above, 400 GB/s over a 4 GB weight pack would put a ceiling near 100 tokens per second, far above the 2.5 measured, which suggests the llm.vhdl bottleneck at 75 MHz is the clock and the pipeline, not memory.
Why anyone is doing this now
Startup Fortune ties the work to scarcity, noting Micron's record quarter and its refusal to commit to a date when memory supply catches up. The practical read is narrower. These cards were cheap because mining ended, and an open model that fits in 4 GB of INT4 weights turns them into a teaching platform for inference hardware, much as Valve's Timur Kristóf has done for old AMD GPUs on Linux. Neither project beats a consumer GPU on speed or cost per token, and neither claims to.
The fable5_llm repository names no model beyond Qwen3.5-9B, but Startup Fortune names Qwen3.8-27B across dual modules as a goal for the llm.vhdl effort; for how that model family is licensed, see the Qwen3.8-27B licence breakdown. The number to watch is the R3 build: if it reaches silicon with timing closed, the gap between 8.18 and 17.3 is the measure of what is left.
Sources
More in Hardware
- 01Teradyne Takes an Undisclosed Stake in Bright Machines to Put Robots on AI Server LinesThe Oct. 5 announcement gives no dollar figure and no ownership share, and says the companies will only "evaluate" integrating Teradyne's robots and board testers.
- 02Starship Reaches Orbit on Flight 14 and Deploys 26 Starlink V3 Satellites, Then Splashes Down EarlyAn engine failure after stage separation cut the planned nine-hour mission to about three, and ship reuse remains undemonstrated.
- 03Norway Wants a Temporary Ban on Smart Glasses in Parks, Gyms and SchoolsDigitalisation minister Torgeir Micaelsen has not yet named the devices or the bill's text, and his Labour government needs other parties' votes to pass it.
- 04Qualcomm Licenses Huawei's LogicFolding Patents in Multiyear Cross-LicenseHuawei says its patent licensing contracts will exceed $6.9 billion after the deal, but neither company has disclosed what Qualcomm is paying.