Doom on a RISC-V core written in TxHDL
Doom now runs on Vreteno, the RISC-V core written in TxHDL, on an Artix-7 FPGA board. The game draws into a framebuffer in DDR3 memory, the board shows it on HDMI, and the keys come over the serial line. This post shows what runs, how fast, and how the way TxHDL designs are built lets changes like this one go from a plan to the board quickly.
What runs on the board
The board is an Alinx AX7A200B with a Xilinx Artix-7 XC7A200T. The design on it is the TxHDL flagship system:
- Vreteno, an RV32IMAC core with Sv32 paging, at 100 MHz.
- DDR3 memory, with a small write-through data cache in front of it.
- Razboj, a small GPU that draws triangles into a framebuffer.
- A scanout engine that reads the framebuffer from DDR3 and sends it to HDMI at 640 by 480.
- An Ethernet port with receive and send engines that move frames to and from DDR3.
Doom is doomgeneric, a portable version of the game whose platform layer is five functions. The board’s platform layer does three things:
- It scales each 320 by 200 frame by two and writes it into the
framebuffer at
0x4200_0000. The scanout shows it from there, with no register writes per pixel. - It reads keys from the serial port. A terminal on the next machine sends them.
- It measures time with the core’s cycle counter, and prints the frame rate every 64 frames.
The game data is Freedoom 0.13.0, which may be passed on.
The 28.8 MB WAD file goes to the board with fastboot over Ethernet,
next to the program: fastboot boot doom.bin freedoom1.wad. The
transfer takes about 32 seconds. The program finds the WAD in memory
right after itself, so the board needs no storage.
How fast
These are the board’s own cycle counts, at 100 MHz:
| What | Cycles per frame | Frames per second |
|---|---|---|
| The demo loop | 16.5 to 21.3 M | 4.6 to 6.0 |
| Playing the first level | 13.2 to 15.7 M | 6.3 to 7.5 |
That is slow, and the reason is known. The core does all of Doom’s work in software: the game logic, the renderer, and the copy into the framebuffer. The core runs one instruction per cycle at best, at 100 MHz, without floating point. Doom needs none, since it uses fixed-point arithmetic.
The next steps are measured already in the cycle-timed model of the machine:
- Compiler settings. Every program for the core is built for size today. Built for speed instead, the GL test’s work per frame falls by a third or more in the model. A board run will show whether the larger code still fits the 16 KiB instruction cache.
- A second core. The plan is to add a second Vreteno that first works as a coprocessor, and later runs Linux in SMP mode.
The changes that led to it
Doom came at the end of a run of changes, each checked on the board:
- Texture mapping in hardware. Razboj gained perspective-correct texture coordinates, a texture cache over DDR3, and texture environments. Its netlist matches the software model byte for byte. A textured, lit icosahedron then ran on the board at 12 frames per second, and at 15 after the program overlapped its work with the GPU’s.
- Network loading. Loading a Linux image over Ethernet went from 53.5 seconds to about 4. The receiver had held one frame, so it lost the second of every pair that came back to back. A copy routine that worked a byte at a time was the second cause.
- A data cache. The core gained a write-through cache for DDR3. Two coherence bugs were found and fixed before it went to the board, one by an outside run of the RISC-V architectural tests. With the cache, Linux goes from the start of the load to a shell prompt in 16.3 seconds, down from about 100.
Why it is fast to build
Four things make this pace possible. None of them is new on its own. The combination is what matters.
One source for the model, the netlist and the tests
A TxHDL design is a Rust library. A unit is a struct, and its behaviour
is an async function that waits for clock edges. The same source
does three jobs:
- It runs as a fast simulation in plain Rust tests.
- It lowers to Verilog and VHDL for synthesis.
- It records its own run, so the build can check the lowered netlist against that run.
The build co-simulates every lowered netlist under nvc and Verilator against the Rust run. When the texture unit went in, the check was that its netlist drew the same bytes as the model, for ten rounds of random triangles. A mismatch fails the build before anyone builds a bitstream.
A hermetic Bazel build
Everything is built by Bazel, and every tool comes from a pinned archive fetched by checksum:
- LLVM, for the Linux kernel and its user space.
- The RISC-V GCC toolchain with newlib, for the bare-metal C programs like Doom.
- nvc and Verilator, for the co-simulations.
- Vivado, for place and route. Its runs share one lock, so that only one runs at a time on the build machine.
Doom itself follows the same rule. The doomgeneric source and the Freedoom data are fetched at a pinned commit and checksum. They are never copied into the TxHDL tree. doomgeneric is GPL-2.0 and TxHDL is Apache-2.0, so a Doom image is GPL, and the source tree stays Apache-2.0.
The effect is that any machine with bazelisk builds the same bits.
bazel build //demo/doom:doom gives the board image, and
bazel run //demo/doom:host -- <wad> <directory> <frames> plays a
scripted game on the host and writes its frames as PNG files, from the
same source. The Doom port was debugged on the host and
in the machine model before it touched the board.
A model that predicts the board
TxHDL’s machine model runs the whole system in software, and has a mode that counts cycles. Most changes are measured there before any board time. For example, the model predicted the gain from a word-at-a-time copy in the network receive path. On the board the copy then fell from 82,100 to 28,300 cycles per frame, as predicted.
The model still has gaps. For the GL code it reports about half the board’s cycles. That gap is now a tracked issue, and board runs check every prediction that matters.
Agents, with checks at each step
Several Claude agents did this work under my direction, as every commit in the repository says, and each commit includes its prompts word for word. One agent runs the board sessions and owns the timing. Others own the memory system, the bus and Ethernet, the graphics, and the coordination. Another agent runs on the machine next to the board and records the HDMI output.
Nothing merges on an agent’s word. Each change passes the full test suite and a place and route of the whole flagship. Hardware changes then get a board session, with the HDMI capture checked frame by frame. I merge the pull requests.
Read more
- TxHDL’s pages, with the paper and the documents.
- The icosahedron demo, the first thing Vreteno drew on HDMI.
- Razboj, a minimal GPU in TxHDL.
- Vreteno, a RISC-V core in TxHDL.