Vreteno: a RISC-V core written in TxHDL
Vreteno is a 32-bit RISC-V processor core designed in TxHDL. TxHDL is a hardware description language implemented as a native Rust library. The core implements the RV32IMC instruction set architecture: the base integer set, standard multiply/divide extensions, and compressed instructions. It supports traps, hardware interrupts, and system memory access via an AXI bus. Lockstep verification checks core state against a reference model on every cycle, while gate-level netlists match cycle-accurate Rust simulations. This article details the core microarchitecture, verification strategy, and FPGA implementation.
Design motivation
Validating a hardware description language requires nontrivial designs that exercise realistic microarchitectural structures. A processor core serves as an ideal testbench. It incorporates multi-stage execution pipelines, memory interfaces, bus logic, exception handling, and strict compliance requirements defined by the RISC-V specification.
The name derives from the Serbian word for a spindle, the rotating component of a spinning wheel. A companion GPU core developed subsequently was named Razboj, meaning loom.
Pipeline microarchitecture
Vreteno uses a three-stage execution pipeline:
- Fetch: Retrieves the instruction at the current program counter into the instruction register. Compressed 16-bit instructions expand into standard 32-bit representations, providing a uniform decoding interface for subsequent stages.
- Execute: Decodes instructions, evaluates arithmetic and logical results, resolves branches, and dispatches memory transactions across the bus.
- Writeback: Formats incoming load data and commits results to the integer register file.
The processor implementation comprises a single TxHDL hardware unit.
A Rust struct declares register and memory state, while an async fn run loop
drives execution.
Each loop iteration synchronizes to a rising clock edge and computes pipeline
transitions concurrently.
Taken branches and jumps invalidate the instruction in the fetch stage, incurring a single-cycle bubble. Data forwarding bypasses values from writeback to execute stages when instructions reference pending results. Load instructions require a one-cycle stall when dependent instructions immediately consume read data. This constraint prevents critical path degradation through address computation adders.
Multiply and divide operations utilize a dedicated multi-cycle sequencer. Multiplication completes through FPGA DSP blocks, while division uses a 32-cycle shift-and-subtract state machine. The execute stage holds dependent instructions until the sequencer completes.
The core implements RISC-V machine mode with eight control and status registers.
System calls (ecall) and illegal instructions vector to an exception handler,
resuming execution via mret.
External interrupts and timer interrupts preempt normal instruction execution.
Memory and bus interfaces
Memory transactions outside the primary internal page route over an AXI interconnect. On hardware, an AXI crossbar connects four distinct peripherals to the core: internal data memory, a system timer, an AXI-Lite serial UART bridge, and one gigabyte of external DDR3 memory. Instruction memory resides within internal block RAM. Because the core cannot read constants directly from instruction memory, linker scripts allocate read-only constants into data memory regions.
Verification methodology
Comprehensive verification relies on lockstep co-simulation. A software reference model written in Rust executes alongside the hardware simulation. Every time Vreteno retires an instruction, the reference model steps forward. The test framework compares the program counter, all 31 general-purpose registers, control status registers, and halt flags on each clock cycle. At test completion, the framework validates data memory, timer counters, and transmitted UART byte streams.
The verification suite evaluates standard benchmark routines and 64 randomized program sequences. Randomized programs interleave supported instructions, dual instruction lengths, branch targets, and asynchronous interrupts. The harness tallies instruction coverage to ensure test generators exercise corner cases consistently.
Fault-injection tests confirmed that verification checkers fail predictably:
- Substituting logical shifts for arithmetic shifts failed immediately.
- Omitting branch pipeline flushes triggered sequencing mismatches.
- Altering status flags during
mretexecution violated state checks. - Modifying return address offsets in compressed calls halted test execution.
Co-simulation uncovered subtle timing bugs during development. In one instance, a division sequencer continued stepping during interrupt servicing, producing corrupted results. In another, a multiply operation read uncommitted register state when following an incomplete load transaction. Automated tests detected these bugs before hardware deployment.
Dual simulation and synthesis flows
A single Rust source file defines the core architecture.
The #[lower] procedural attribute compiles the execution loop into
synthesizable Verilog and VHDL.
Bazel targets simulate generated netlists under NVC and Verilator, verifying
cycle-by-cycle parity with Rust simulations.
Vivado synthesizes Vreteno for an AMD Artix-7 xc7a200t in approximately two
minutes.
The design utilizes 2,151 LUTs, 489 flip-flops, two block RAM tiles, and four
DSP blocks, meeting a 10 ns clock constraint (100 MHz).
Targeting the Nangate45 open cell library via Yosys and OpenROAD achieves 3 ns
timing closure (333 MHz).
Software toolchain integration
Executing compiled software requires standard toolchain support. Bazel rules provision bare-metal RISC-V targets for Rust and GCC toolchains for C and C++. Build targets compile source code into ELF binaries, while a formatting utility generates initial memory images for simulation and hardware bootloaders. The formatting utility enforces memory boundary limits, rejecting programs that exceed the 4 KiB instruction memory limit.
Using the compressed instruction extension reduced executable size significantly: a standard Rust initialization binary shrank from 33 words to 25, while a C++ equivalent dropped from 89 words to 63.
FPGA hardware bring-up
An early build of the core was deployed to an Alinx AX7A200 FPGA development
board, incorporating internal memory, a system timer, and a serial interface.
The hardware successfully transmitted OK over UART at 115,200 baud, echoed
received input bytes, and updated status LEDs, matching simulation traces
identically.
Subsequent board targets integrate DDR3 memory controllers. These configurations have undergone complete place-and-route and board-level simulation with memory models in preparation for hardware flashing.
Current limitations and roadmap
Vreteno currently implements machine mode without physical memory protection.
Debugger breakpoints (ebreak) halt the processor rather than trapping to an
external debug unit.
Instruction memory contents remain static at bitstream generation time pending
serial bootloader integration.
These enhancements are tracked in the project repository.
Complete architectural details and source files appear in the Vreteno section of the TxHDL documentation. Source code is available in the TxHDL repository.