API Reference¶
Generated from the package's docstrings, including the compiled
extension's, so it describes the installed version. The examples call a
running simulator cpu.
Configuration¶
Config
¶
Full simulator configuration with flat parameter access.
Example::
from rvsim import Config, Cache, BranchPredictor, Prefetcher
cfg = Config(
width=4,
branch_predictor=BranchPredictor.TAGE(),
l1i=Cache("128KB", ways=8, prefetcher=Prefetcher.NextLine(degree=2)),
l1d=Cache("128KB", ways=8, prefetcher=Prefetcher.Stride(degree=2, table_size=128)),
l2=Cache("4MB", ways=16, latency=12),
)
replace
¶
Return a new Config with the given fields overridden.
Example::
base = Config(width=4, branch_predictor=BranchPredictor.TAGE())
wide = base.replace(width=8)
ooo = base.replace(backend=Backend.OutOfOrder(rob_size=128))
See Configuration for every parameter.
Cache
¶
BranchPredictor
¶
Namespace for branch predictor configurations.
TageBanking
¶
TAGE-SC-L's banked tables: banks before first_long_bank
share an array of short_factor slices of table_size
entries, the rest one of long_factor; each pair of banks forms
a 2-way table, and enabled (one flag per bank) says which exist.
ScGehl
¶
One statistical corrector GEHL component: a counter table per
history length (longest first), each 2**log_entries entries.
No lengths turns the component off.
ScLocalGehl
¶
A statistical corrector GEHL over per-branch local histories:
histories of them (a power of two), a branch's at
(pc ^ (pc >> index_shift)) % histories; mix_pc XORs the
branch's pc & 15 into each update.
ScLTage
¶
SC-L-TAGE + ITTAGE composed predictor.
Combines TAGE (direction), Loop Predictor, Statistical Corrector, and Indirect Target TAGE into a single high-accuracy predictor.
The TAGE parameters are shared with the standalone TAGE config;
the defaults add TAGE-SC-L's own TAGE rules (use_alt_counters=16,
CBP-5 allocation and update with 1-bit useful counters, two
allocations, a reset_interval of 1024 allocation penalties,
pc_bits history with a 27-bit path, tage_sc_l hashing, and
the 64KB TAGE-SC-L's 36 banked 1024-entry tables over history
lengths 6 to 3000).
The loop predictor, SC and ITTAGE have their own sub-configs; the
loop predictor and SC defaults are Seznec's 64KB TAGE-SC-L (CBP-5).
The SC's GEHL components are BranchPredictor.ScGehl and
BranchPredictor.ScLocalGehl values; None takes the default.
MemDepPredictor
¶
Namespace for memory dependence predictor configurations.
Backend
¶
Namespace for pipeline backend configurations.
Fu
¶
Functional unit pool configuration for the O3 backend.
Instantiate Fu with a list of unit descriptors (the inner classes).
Any FU type omitted will be absent from the pool, so include every type
your workload exercises::
Fu([
Fu.IntAlu(count=4, latency=1),
Fu.IntMul(count=1, latency=3),
Fu.IntDiv(count=1, latency=35),
Fu.FpAdd(count=2, latency=4),
Fu.FpMul(count=2, latency=5),
Fu.FpFma(count=2, latency=5),
Fu.FpDivSqrt(count=1, latency=21),
Fu.Branch(count=2, latency=1),
Fu.Mem(count=2, latency=1),
])
IntAlu
¶
Integer ALU: add, sub, logic, shift, compare, set-less-than.
IntMul
¶
Integer multiplier: mul, mulh, mulhsu, mulhu.
IntDiv
¶
Integer divider: div, divu, rem, remu. Non-pipelined.
FpAdd
¶
FP adder: fadd, fsub, fmin, fmax, fcmp, fcvt.
FpMul
¶
FP multiplier: fmul.
FpFma
¶
FP fused multiply-add: fmadd, fmsub, fnmadd, fnmsub.
FpDivSqrt
¶
FP divider/sqrt: fdiv, fsqrt. Non-pipelined.
Branch
¶
Branch/jump unit: all conditional branches, jal, jalr.
Mem
¶
Memory address calculation for loads and stores.
MemoryController
¶
Namespace for memory controller configurations.
Simple
¶
Fixed-latency controller: answers latency core cycles after it
starts a request, and starts requests serialised on bandwidth_gib_s.
DDR5
¶
Command-level DDR5 controller with JEDEC timing.
Timing comes from speed_bin ("4800B" or "5600B") and may
be overridden per field with timing={"t_rcd": 40, ...} in DRAM
command clocks. The controller runs at the DRAM clock; the core clock
is Config(cpu_clock_mhz=...).
Policies: scheduler is "FrFcfs" (default) or "Fcfs";
refresh is "AllBank" (default) or "SameBank";
address_mapping is "RoRaBaChCo" (default), "RoRaBaCoCh"
or "RoCoRaBaCh"; ecc is "None", "SecDed" or
"ChipKill". power_down_idle_ns enables rank power-down after
that many idle nanoseconds; patrol_scrub_ns enables ECC patrol
scrubbing at that interval.
Prefetcher
¶
Namespace for prefetcher configurations.
ReplacementPolicy
¶
Namespace for cache replacement policies.
Coherence
¶
Coherence fabric between the private caches, used when hart_count > 1.
HomeAgent
¶
Namespace for coherence home-agent policies (who must be snooped).
Interconnect
¶
Namespace for coherence interconnect topologies.
Every kind takes the cycles a message spends per hop and the bytes a port or link moves per cycle.
Crossbar
¶
Bases: _Kind
Any port to any port in one hop (default).
Ring
¶
Bases: _Kind
Bidirectional ring, routed the shorter way; the home is one stop.
Mesh
¶
Bases: _Kind
Square 2-D mesh with XY routing.
Torus
¶
Bases: _Kind
Square 2-D torus (mesh with wraparound) with XY routing.
Hypercube
¶
Bases: _Kind
Hypercube with dimension-order routing.
presets
¶
Built-in configuration presets.
basic— modest 4-wide OoO core, small caches, good for quick runs.fast— Apple M4 P-core class: 8-wide OoO at 4.4 GHz, 630-entry ROB, 192KB L1I, 128KB L1D, 4MB L2, 36MB L3, the 64KB TAGE-SC-L with ITTAGE, 4 unified FP/SIMD pipes, and DRAM controller.linux— a multi-corefastsystem with the memory map the bundled Linux image expects, kept coherent over an interconnect, with DDR5 memory.cortex_a72,m1,p550— models of the Arm Cortex-A72, an Apple M1-class core and the SiFive P550, from their published microarchitecture.
Usage from the CLI::
rvsim mandelbrot.elf --preset fast
Usage from Python::
from rvsim import presets, Simulator
cfg = presets.fast()
Simulator(cfg, binary="mandelbrot.elf").run()
basic
¶
Modest 4-wide out-of-order core with small caches.
This is identical to Config() with no arguments.
fast
¶
Apple M4 P-core class configuration.
Based on publicly known M4 Everest P-core microarchitecture:
- 8-wide rename/dispatch (Apple decodes up to ~10 but dispatches 8)
- 630-entry ROB, 108-entry store buffer, ~160-entry issue queues
- 6 integer pipes (4 simple ALU + 2 complex with mul/div)
- 4 unified FP/SIMD pipes (each handles add, mul, FMA)
- 2 branch units, 4 load/store AGUs (3 load + 2 store capable)
- 192KB 6-way L1I, 128KB 8-way L1D (3-cycle hit), 4MB L2, 36MB L3
- 4.4 GHz P-core clock
- Seznec's 64KB TAGE-SC-L (CBP-5) with ITTAGE (Apple's predictor is proprietary but believed to be TAGE-class)
- Non-inclusive cache hierarchy
- LPDDR5-class DRAM controller
Vector units are modeled as a RISC-V V equivalent of Apple's 4 NEON/AMX pipes — VLEN=256 with 4 lanes and chaining.
Note: the sim models FU types independently, so having count=4 for FpAdd/FpMul/FpFma slightly overstates mixed-FP throughput vs the real M4 (which has 4 unified pipes). The 8-wide dispatch width naturally limits total throughput to realistic levels.
linux
¶
linux(harts: int = 8, *, memory: str = 'ddr5', speed_bin: str = '5600B', interconnect: str = 'mesh', real_time: bool = True, core: Config | None = None) -> Config
core (the fast preset by default) in a system that boots the
bundled Linux image: its memory map, harts, coherence and memory
replace the core config's.
harts harts boot through OpenSBI's HSM into an SMP kernel; with
more than one, the private caches are kept coherent by a snoop-filter
home agent over interconnect (crossbar, ring, mesh,
torus or hypercube). memory is ddr5 (JEDEC
command-level timing at speed_bin, four channels) or dram (the
core config's own memory controller).
With real_time the CLINT ticks at the device tree's 10 MHz
timebase, so the guest's clock keeps time with the modelled one.
Without it the CLINT ticks every cycle: guest time runs
cpu_clock_mhz / 10 times fast, which shortens a boot's sleeps and
timeouts but floods measurements with timer interrupts.
cortex_a72
¶
Cortex-A72: 3-wide O3, 48KB I\(, 32KB D\) (8 MSHRs), 1MB L2.
ARM Cortex-A72 machine config. https://en.wikipedia.org/wiki/ARM_Cortex-A72
Microarchitecture (publicly documented): - 3-wide fetch/decode/rename/dispatch/issue - Out-of-order execution, 128-entry ROB - Eight issue queues of 8 entries (10 for branches), 66 in all; modelled as one 66-entry queue until per-pipe queues exist - 16-entry store queue, 32-entry load queue - One load AGU, one store AGU - PRF: 128 integer + 128 FP physical registers - Execution units (Chips and Cheese, Graviton at 2.3 GHz): - 2x integer ALU (latency 1), 1x multi-cycle pipe (multiply 3, divide non-pipelined) - 2x FP/NEON pipeline (modeled as FpAdd + FpMul/FpFma per pipe) - 1x FP div/sqrt (latency ~17-38, non-pipelined) - 1x branch unit (latency 1) - 48KB L1-I (3-way), 32KB L1-D (2-way, 4-cycle load-to-use), 8 MSHRs - 1MB L2 (16-way) on the Raspberry Pi 4's BCM2711, 21 cycles - 48-entry L1 ITLB, 32-entry L1 DTLB, 1024-entry 4-way L2 TLB - 4096-entry BTB, 31-entry return stack; mispredict penalty ~15 cycles - Load/store prefetcher (TRM 6.4.9): loads prefetch into the L1D and 22 requests ahead into the L2 (CPUECTLR_EL1 reset), crossing pages through the TLB (CPUACTLR_EL1[43] reset); store misses prefetch into the L2 only. The L1D distance, table sizes and store run length are not published. - Clocked at 1.5 GHz as on the Raspberry Pi 4
m1
¶
M1-style: 4-wide, 128KB L1-I/D, 4MB L2.
p550
¶
SiFive Performance P550 — 3-wide, 13-stage, out-of-order.
Microarchitecture notes (from Chips and Cheese reverse engineering): - 3-wide fetch/decode/rename/retire - ROB ~72 entries (comparable to Core 2 / Goldmont Plus class) - Modest issue queue, ~32 entries estimated - Load queue ~24 entries, store buffer ~16 entries (described as "thin") - PRF sized with "plenty of capacity compared to ROB size" - 9.1 KiB branch history table with good pattern recognition - 32-entry BTB handles taken branches with zero bubbles - 32KB 4-way L1i (3-cycle), 32KB 4-way L1d (3-cycle load-to-use), 64B lines - 256KB 8-way private L2 at 13 cycles; 4 MB L3 at ~38 cycles on the EIC7700X - DRAM at 194 ns on the HiFive Premier P550 (272 cycles at 1.4 GHz) - FP add, multiply and FMA at 4 cycles; one load AGU and one store AGU - 32-entry fully associative L1 TLBs, 512-entry L2 TLB - 13-stage pipeline → ~11-13 cycle mispredict penalty - No hardware misaligned access support (trap-based emulation) - Prefetchers unpublished: a load stride prefetcher that keeps to the page, and no store prefetcher, as the cautious reading
SiFive Performance P550 machine config.
Based on published microarchitecture analysis:
- Chips and Cheese: "Inside SiFive's P550 Microarchitecture" (Jan 2025)
- SiFive official specs: 13-stage, triple-issue, out-of-order, RV64GC
- Measured on Eswin EIC7700X SoC @ 1.4 GHz, 32KB+32KB L1, private L2, 4MB shared L3
- Published SPECInt2006: 8.65/GHz
- Observed IPC: approaching 3.0 on favorable workloads
Running a binary¶
Environment
dataclass
¶
Immutable description of a simulation run for reproducibility.
config
class-attribute
instance-attribute
¶
Config or dict. If None, uses Config() defaults.
load_addr
class-attribute
instance-attribute
¶
Load address for the binary.
run
¶
Run the simulation and return a Result.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
quiet
|
bool
|
Suppress exceptions and return error Result instead. |
True
|
limit
|
int | None
|
Max cycles to simulate. |
None
|
Example::
env = Environment(binary="software/bin/benchmarks/qsort.bin")
result = env.run()
print(result.stats["ipc"], result.stats["cycles"])
Result
dataclass
¶
Structured result of a single run.
stats
class-attribute
instance-attribute
¶
All stats as a Stats object.
wall_time_sec
class-attribute
instance-attribute
¶
Wall-clock time of the run in seconds.
compare
staticmethod
¶
compare(results: dict[str, Any], *, metrics: list[str] | None = None, baseline: str | None = None, col_header: str = '') -> None
Print a comparison table for experiment results.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
results
|
dict[str, Any]
|
Either |
required |
metrics
|
list[str] | None
|
Specific metric names to show. If None, shows a default set. |
None
|
baseline
|
str | None
|
Config name to normalize against (shows speedup ratios). |
None
|
col_header
|
str
|
Label for the config-name column (e.g. "size", "width"). |
''
|
The simulator¶
Simulator
¶
Bases: Simulator
Native simulator.
Example::
sim = Simulator(Config(width=4), binary="qsort.elf")
exit_code = sim.run(limit=10_000_000)
sim = Simulator(kernel_config(), kernel="linux.img", disk="rootfs.img")
sim.run()
The sim-control device¶
A simulator-only MMIO device at sim_control_base (default 0x0010_2000),
not in the device tree, with two 64-bit registers:
| Offset | Register | Access |
|---|---|---|
0x00 |
COMMAND |
write 1 to reset the stats, 2 to dump them labelled with ARG, 3 to end the simulation with ARG as the exit code, 4 to stop the host's run_to here with ARG as the label |
0x08 |
ARG |
read/write: the argument of the next command |
A guest writes ARG, then COMMAND. Under Linux, map the page through
/dev/mem, or use the image's rvsim tool (rvsim dump-stats LABEL,
rvsim break LABEL, rvsim run START END CMD); bare-metal programs
include software/libc/rvsim.h (rvsim_dump_stats, rvsim_break,
rvsim_reset_stats, rvsim_exit). The command takes effect at the end of
the cycle the write reaches the device.
Instruction
¶
One retired instruction, returned by Simulator.step.
PipelineSnapshot
¶
Point-in-time contents of every inter-stage latch; each stage is a
list of slot dicts (length at most width).
Sessions¶
Session
¶
A workload run in phases on one or more configurations.
Example::
s = Session.linux(harts=8)
s.fast_forward(until=Session.LOGIN_SHELL) # cached after the first time
s.switch(presets.linux(harts=8, core=my_core))
s.warm_up(command="coremark 0x0 0x0 0x66 20 7 1 2000")
r = s.measure("coremark 0x0 0x0 0x66 20 7 1 2000")
print(r.ipc, r.stats["core0.bp.committed.accuracy"], r.exit_code)
The session runs on config. Fast-forwards run on
fast_forward_config (config unless given) and return to
config at their stop, so what a session measures always runs on the
configuration it was given or last switched to. Direct changes made
through sim are not part of the history cache keys are made
of; fast-forward with cache=False after them.
linux
classmethod
¶
linux(config: ConfigLike | None = None, *, harts: int | None = None, image_dir: str | None = None, kernel: str | None = None, firmware: str | None = None, disk: str | None = None, **kwargs: Any) -> Session
A session on the bundled Linux image (built by make linux).
config defaults to presets.linux(harts). Fast-forwards run
on presets.linux with its timer ticking every cycle, which
compresses the boot's sleeps, and with config's RAM, VLEN and
ISA options, so the cached boot is shared by every core
configuration of the same system.
resume
classmethod
¶
A session continuing from a checkpoint save wrote, on
config (the configuration it was saved on by default).
Raises ValueError if the workload's files changed since.
run
¶
run(until: Stop | None = None, *, every: int | None = None, on_every: Callable[[Session], None] | None = None) -> Stopped
Runs until until holds (the workload's end by default),
calling on_every(session) every every cycles on the way.
fast_forward
¶
Gets to until quickly: restores it from the cache when this
history has reached it before, else runs to it on the fast-forward
configuration (and caches it). Either way the session continues
from the stop's checkpoint on its own configuration, with caches,
TLBs and predictors cold, so what follows does not depend on
whether the cache held the stop.
Raises WorkloadEnded if the workload ends first.
switch
¶
Continues on config, which must show the guest the same
system (harts, RAM, memory map, VLEN and ISA options). Caches,
TLBs and predictors start cold.
warm_up
¶
Runs unmeasured to until, or through a shell command, to
warm caches and predictors before measuring.
measure
¶
measure(command: str | None = None, *, until: Stop | None = None, name: str | None = None) -> Region
Measures a shell command (Linux) or the run to until.
A command runs under the guest's rvsim run, which snapshots the
stats just before it starts and just after it exits; the region is
the difference, so the session's whole-run stats stay intact.
expect
¶
Runs until the console prints pattern and returns the match.
Raises WorkloadEnded if the workload ends first.
save
¶
Saves the session's state to path (with its console and
history beside it in path.json) for resume.
Saving drains the pipelines first, as gem5 does, so this session
continues a few cycles later than it would have; a session resumed
from path starts with caches, TLBs and predictors cold.
fork
¶
Continues from this point once per configuration: yields
(name, session) pairs, each an independent session on its
config. This session is left as it was.
Region
dataclass
¶
Stopped
dataclass
¶
Where a run stopped, and which stop ended it.
instructions
instance-attribute
¶
Instructions retired by every hart since the system started.
exit_code
class-attribute
instance-attribute
¶
The workload's exit code, when it ended.
label
class-attribute
instance-attribute
¶
The guest's label, for a Marker.
match
class-attribute
instance-attribute
¶
The console match, for a Console or LoginShell.
WorkloadEnded
¶
Bases: RuntimeError
The workload ended before the run reached its stop.
Stop points¶
A run ends the moment one of its stops holds; a | b stops at whichever
comes first. Counts are relative to the run's start, and every run also
ends if the workload does.
Stop
¶
Where a run stops.
key
¶
A stable description of the stop for checkpoint-cache keys, or
None when it has none (a predicate, say).
Marker
dataclass
¶
Bases: Stop
When guest software asks the host to stop: rvsim break LABEL in
Linux, rvsim_break(label) from rvsim.h on bare metal. Any
label stops the run when label is None.
Console
dataclass
¶
Bases: Stop
When console output the session has not yet matched matches
pattern, a regular expression; the match consumes the output up to
its end. The run stops on the cycle the matching output is written.
When
dataclass
¶
Bases: Stop
When predicate(session) returns true, checked every every
cycles. Give it a name to let a fast-forward to it be cached; the
name stands for the predicate, so it must change when the predicate
does.
LoginShell
dataclass
¶
Bases: Stop
At a Linux shell: waits for the login prompt, logs in as
user (with password if the image asks for one), and stops once
the shell prompt appears.
LOGIN_SHELL is LoginShell() with the image's defaults.
rvsim bench¶
The benchmark suite inside Linux from the command line:
rvsim bench # every benchmark, fast core, 8 harts
rvsim bench coremark stream --warm # warm each benchmark before measuring it
rvsim bench --config my_core.py --harts 4 --json results.json
rvsim bench --list
It boots once per system (cached), places the core in the Linux system
with presets.linux(core=...), and prints each benchmark's cycles,
instructions, IPC and branch, L1D, L2 and LLC misses per thousand
instructions.
Sweeps¶
Sweep
¶
Run multiple configs x binaries in parallel across CPU cores.
Example::
from rvsim import Sweep, Config
results = Sweep(
binaries=["qsort.elf", "dhrystone.elf"],
configs={
"baseline": Config(width=2),
"wide": Config(width=4),
},
).run(parallel=True, limit=100_000_000)
results.compare()
run
¶
run(*, parallel: bool = True, limit: int | None = None, max_workers: int | None = None) -> SweepResults
Execute all (binary, config) combinations.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
parallel
|
bool
|
Use multiple processes. |
True
|
limit
|
int | None
|
Maximum cycles per run. |
None
|
max_workers
|
int | None
|
Max parallel workers. |
None
|
Returns:
| Type | Description |
|---|---|
SweepResults
|
|
SweepResults
dataclass
¶
Structured results from a sweep run.
Organised as results[binary_name][config_name] = Result.
compare
¶
compare(*, metrics: list[str] | None = None, baseline: str | None = None, col_header: str = '') -> None
Print a comparison table across all configs and binaries.
Statistics¶
Stats
¶
Bases: dict
Dict-like simulation statistics with querying and comparison.
Keys are the simulator's stat paths (core0.cache.l1d.misses,
core0.bp.committed.accuracy, hart0.retired_insts), plus the
run-level cycles, instructions_retired and ipc.
Example::
result.stats["ipc"]
result.stats["core0.pipeline.stalls.data"]
result.stats.query("miss_rate")
from_core
classmethod
¶
Flatten the native stats object into path-keyed entries.
Every registered path (core0.cache.l1d.hits, system.retired_insts)
becomes a key, and the run-level cycles, instructions_retired
and ipc are added under their short names.
query
¶
Search for statistics matching pattern (case-insensitive regex or substring).
tabulate
staticmethod
¶
Build a comparison table from labeled Stats objects.
Each Stats is typically a .query() result, so all share similar
keys. Columns are the sorted union of all keys across the provided
Stats objects.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
rows
|
dict[str, Stats]
|
|
required |
title
|
str
|
Optional table title rendered above the header. |
''
|
Returns:
| Type | Description |
|---|---|
Table
|
|
compare
¶
Print a two-column comparison table (self vs other) to stdout.
Table
¶
Rendered comparison table. Created by tabulate, displayed via
print() or REPL auto-repr.
ISA utilities¶
reg and csr¶
Register and CSR lookups: constants, name lookup, and the reverse.
from rvsim import reg, csr
reg.A0 # 10
reg("a0") # 10
reg.name(10) # "a0"
csr.MSTATUS # 0x300
csr("mstatus") # 0x300
csr.name(0x300) # "mstatus"
Disassemble
¶
Fluent disassembler for RISC-V binaries and raw bytes.
Usage::
Disassemble().binary("software/bin/programs/qsort.bin").print()
Disassemble().binary("qsort.bin").at(0x80000024, count=10).print()
Disassemble().bytes(data).print()
Disassemble().inst(0x00a00513) # single instruction -> str