openxc7 or Vivado: Build Times and Results Measured

#bazel#FPGA#openxc7#Vivado#yosys#nextpnr#RISC-V#auto#EDA

I built the same three designs with the open toolchain and with Vivado, through the same Bazel rules, and timed every build from cold and warm caches. Small designs build two to three times faster with the open tools. Sixteen RISC-V cores build faster with Vivado, and Vivado’s circuits use about half the logic. Setting up the measurements turned up five problems, three of them in my own rules.

What was compared

rules_openxc7 and rules_vivado share one target interface: vivado_project, vivado_synthesis and vivado_place_and_route. The first runs Yosys, nextpnr-xilinx and Project X-Ray. The second runs Vivado 2025.2, which Bazel installs from the installer archive on first use. Switching a design between them changes one load line.

That makes a fair comparison possible. Each design has a target for each flow, Bazel runs both the same way, and the time measured is what a user of the rules waits for.

The three designs, all for an Artix-7 200T with a 100 MHz clock:

  • blinky: an 8-bit counter on a LED. It measures tool start-up more than work.
  • PicoRV32, one core: the PicoRV32 RV32I core, 1 KiB of block RAM and a LED register.
  • PicoRV32, sixteen cores: sixteen of those, side by side.

How it was measured

All measurements ran on one Google Compute Engine VM: 20 vCPUs of an AMD EPYC 7B12, 47 GiB of memory, and a pd-balanced persistent disk for the builds. They ran inside the same container image that my CI uses for Vivado jobs, with the CI runner stopped.

A script ran one bazel build per measurement, with a fresh Bazel output base for each cold and warm run. It timed five cases:

  • cold: a new output base and an empty cache. The open tools are unpacked again.
  • warm: a new output base, with the disk cache that the cold run filled. This is what a CI job with a shared cache sees.
  • no-op: the same build again, with nothing changed.
  • edit at the end: a comment line appended to the design’s source.
  • edit at the top: a comment line put first in the source.

Synthesis and place and route were timed separately. Each case ran three times. The full report lists every flag and every mount.

Vivado’s installation was not timed with any build. rules_vivado keeps it in a cache outside Bazel’s output bases. The first install on this machine took 100 minutes, most of it copying the 103 GB installer. Every later build reused it, and the check before each pass took about 9 s.

Build times

Median seconds from source to bitstream, synthesis and place and route together:

Design openxc7 cold Vivado cold openxc7 warm Vivado warm
blinky 79.5 194.4 10.1 7.8
PicoRV32, one core 91.7 221.6 9.8 8.4
PicoRV32, sixteen cores 587.1 334.2 10.2 9.0

Vivado costs about three minutes per build at any size. blinky has eight flip-flops, and Vivado still takes 102 s to synthesise it and 94 s to place and route it. Most of that time goes to starting Vivado, creating a project and writing reports.

The open tools scale with the design. Their place and route takes 37 s for blinky, 56 s for one core and 483 s for sixteen. Between one core and sixteen, Vivado becomes the faster flow.

With a warm disk cache, both flows rebuild in 8 to 10 s, and a no-op build takes under a second. Shared caches remove most of the difference.

Edits

After an edit, the two flows behave differently:

Design openxc7, edit at end Vivado, edit at end openxc7, edit at top Vivado, edit at top
blinky 6.7 149.5 22.4 150.0
PicoRV32, one core 13.5 178.8 52.3 172.7
PicoRV32, sixteen cores 76.3 289.5 513.2 280.6

A comment at the end of the file leaves Yosys’s netlist the same. Bazel sees an unchanged input and reuses the place and route from before, so the open flow pays only for synthesis. A comment at the top moves the source line numbers that Yosys records in the netlist, and then both steps run.

In the Vivado flow, place and route depends on a directory that holds copies of the sources. It runs again after any change to them.

A real edit changes the logic, so the “edit at the top” columns are the fair ones. They show the same picture as the cold builds: the open flow is faster for small designs and slower for large ones.

What each flow builds

Time is half of the comparison. The reports from both flows show what came out:

openxc7, 1 core Vivado, 1 core openxc7, 16 cores Vivado, 16 cores
LUTs 1,671 890 26,668 14,216
Registers 553 555 8,852 8,880
Speed 147 MHz about 165 MHz 131 MHz about 144 MHz
Hold time one violation met seven violations met

Vivado’s netlists use about half the LUTs of the open flow’s. The register and block RAM counts are the same, so the difference is in logic optimisation and LUT mapping.

The speeds come from different timing models, so they compare only roughly. Hold time is a clearer difference. Vivado’s router adds delay to short paths until hold time is met. nextpnr-xilinx does not, and leaves violations of a few hundredths of a nanosecond, here on paths into block RAM. I have not tried these bitstreams on a board, so I do not know whether those paths fail in practice.

What broke on the way

The benchmark found five problems before it produced a single number.

Yosys left a cell that nextpnr cannot place. PicoRV32 declares a wire that only renames another. Yosys kept a $buf cell for it, and nextpnr stopped with “no BELs remaining to implement cell type ‘$buf’”. Blinky never had such a wire. rules_openxc7 now maps leftover $buf cells to plain connections after synthesis.

nextpnr failed the build on a timing violation. One hold violation of 0.02 ns, and no bitstream. Vivado reports violations and finishes the build. rules_openxc7 now does the same: the violation stays in the log.

rules_vivado modified my source files. Its place and route copied files into its working directory with cp -a, onto sandbox symlinks that point at the sources. The copy wrote through the symlinks and made the sources read-only, so every second build failed with “Permission denied”. My own tests kept their sources in subpackages, where the paths never collided. Version 3.16.2 fixes it.

Vivado deleted my benchmark. In the first version, all sixteen cores ran the same program, and the LED showed the exclusive or of one bit from each. Sixteen equal bits give 0 every time. Vivado proved that and removed the entire design: 0 LUTs, 0 registers. Yosys did not find it and built all sixteen cores. Vivado’s fast “sixteen-core” times measured an empty chip. Now each core counts in a different step, and Vivado keeps all of them. The lesson for anyone writing a benchmark: check the utilisation report before you trust a time.

Bazel’s caches made a cold build warm. Bazel 9 keeps unpacked repositories next to its repository cache and reuses them in every output base. In a trial run, a second “cold” synthesis took 8.8 s, because it found the open toolchain already unpacked. The cold case now turns that cache off, and the same build takes 28.5 s.

Which one to use

For small and medium designs that change often, the open flow is faster. It is also free, and Bazel downloads it on first use, with no installer. For large designs, and for anything that must meet timing with margin, Vivado builds faster and makes a smaller, faster circuit.

The rules make the choice cheap to change. With both flows behind one interface, a project can use the open tools day to day and run Vivado in CI before a release.

The report, the analysis notes and the benchmark script are in the rules_openxc7 repository, under integration/bench.