Razboj: a minimal GPU in TxHDL

#GPU#FPGA#Rust#TxHDL#auto

Razboj is a minimal graphics rasteriser implemented in approximately one hundred lines of TxHDL. It reads a display list from memory and writes rendered pixels into a framebuffer over an AXI bus. TxHDL lowers the design to synthesizable Verilog and VHDL, and the build verifies both netlists against the software simulation trace. This post describes the rasteriser architecture and hardware design tradeoffs.

What it draws

The reference demonstration scene measures 64 by 64 pixels (4,096 pixels total). The scene consists of eight display list entries: a background clear, three rectangles, and four triangles. The rasteriser renders the entire scene in approximately 13,000 clock cycles. A verification harness reads the completed framebuffer from memory and saves the output image.

A house with a red roof, a hill
    and a yellow sun on a dark sky, 64 by 64 pixels
Figure 1: Demonstration scene rendered by Razboj from eight display list primitives.

The name Razboj comes from the Serbian word for a weaver’s loom. It pairs with our RISC-V CPU core, Vreteno, which is the Serbian word for a spindle.

The display list

The display list acts as a compact graphics instruction set. Razboj supports three primitive commands:

  • Clear: accepts a 24-bit RGB colour and fills the entire screen.
  • Rectangle: accepts a colour and bounding box coordinates.
  • Triangle: accepts a colour and three vertex coordinate pairs.

Software defines these primitives as a Rust enum. When written into memory, each command occupies an eight-word entry with fixed field layouts. Fields do not cross 32-bit word boundaries. Software writes fields using bit shifts and masks, while hardware extracts attributes via fixed slice offsets. Because entry sizes are powers of two, hardware calculates consecutive entry addresses with simple shifts rather than integer multipliers.

A dedicated control word stores the primitive count at the head of the list. The rasteriser polls this word until it receives a non-zero count. Software writes all primitive entries before updating the count word, signaling that the display list is ready for rendering.

A lightweight software helper performs per-primitive setup. It clips bounding boxes to screen limits and winds vertices clockwise so that interior pixels produce non-negative edge functions. The helper discards degenerate primitives that produce zero visible area.

The rasteriser

The rasteriser processes primitives sequentially, visiting bounding box pixels one pixel per cycle. For triangles, three linear edge functions evaluate whether a given pixel lies inside the geometry. Because edge functions vary linearly with pixel coordinates, stepping one pixel horizontally or one row vertically adds a precalculated constant. For each edge, the unit tracks the current accumulator value, the row-start value, and the two coordinate increments.

The rasteriser only executes multiplications once per triangle during setup. Per-pixel evaluation requires only three additions and three sign checks.

Early iterations used 16-bit fixed-point arithmetic. Randomized fuzz tests with vertices outside the visible viewport quickly triggered arithmetic overflow. Widening the datapath to 32 bits resolved the issue, demonstrating the value of property-based randomized testing over manual test vectors.

Writing over AXI

Razboj functions as an AXI host. Each rendered pixel is issued as a single-beat write burst. The unit asserts address and data beats in the same cycle, ensuring compliance with AXI4 transaction ordering requirements. When the memory interconnect signals backpressure, the rasteriser pipeline stalls without dropping pixel writes.

The host tracks up to four concurrent in-flight transactions. This limit corresponds to four allocated AXI transaction identifiers, which the host reclaims upon receiving write acknowledgments.

The framebuffer occupies a dedicated dual-port memory holding one 32-bit word per pixel, with display lists stored in upper addresses. The current reference memory processes writes sequentially, taking approximately three cycles per pixel. A pipelined framebuffer supporting single-cycle write beats represents a natural performance optimization.

Netlist generation and cross-verification

Both the rasteriser and the framebuffer use the synthesizable subset of Rust supported by TxHDL. The Bazel build compiles the design to Verilog and VHDL, builds a 16-by-16 pixel testbench, and simulates both netlists under nvc and Verilator. The test harness verifies cycle-accurate equivalence across all port signals.

Automated lowering uncovered three identifier collisions with HDL reserved keywords. A register named next is a keyword in VHDL, while tri and inside are reserved keywords in Verilog. Adding automated keyword checks to the lowering compiler resolved these conflicts before code generation.

Integration with a CPU core

Display lists can originate from software executing on an integrated processor. In our network-on-chip demonstration, Vreteno computes the 3D projection of a rotating icosahedron, removes back-facing facets, calculates surface lighting, and writes the resulting triangle display list to shared memory. Razboj reads the display list across two router hops and renders the image.

Design constraints and trade-offs

Razboj prioritizes architectural simplicity over peak performance. Its current microarchitecture includes the following characteristics:

  • It renders at most one pixel per clock cycle.
  • It omits a depth buffer, drawing primitives strictly in display-list order.
  • It performs bounding-box clipping rather than polygon geometry clipping.
  • It fetches display list commands sequentially via single-word reads.

Each of these subsystems can be extended modularly without redesigning the core edge-walking datapath or the AXI bus interface.

Complete implementation listings are available in the Razboj chapter of the TxHDL documentation. The full source code is mirrored in the repository.