Configuration¶
Every aspect of the simulated machine is set through the Config class.
Its parameters are flat keywords; caches, predictors, backends, memory
controllers and the coherence fabric are small builder classes passed to
them. A Config is serialised to the Rust core when a simulator is built,
where it is validated as a whole: an unknown key, or a combination no
machine can have (for example a BTB whose set count is not a power of two,
or an exclusive L1/L2 with more than one core), raises ValueError before
anything runs.
Basic Usage¶
from rvsim import Config, Cache, Backend, BranchPredictor, MemDepPredictor
config = Config(
width=4,
backend=Backend.OutOfOrder(rob_size=128),
branch_predictor=BranchPredictor.TAGE(),
l1d=Cache("32KB", ways=8, latency=1, mshr_count=8),
l2=Cache("256KB", ways=8, latency=10),
)
Use replace() to derive new configs from a base; every Config keeps
its own copies of its components, so changing one never changes another:
base = Config(width=4, branch_predictor=BranchPredictor.TAGE())
narrow = base.replace(width=2)
wide = base.replace(width=8)
rvsim.presets holds complete machines: basic(), fast(), p550(),
cortex_a72(), m1(), and linux() for a system that boots the bundled
Linux image. See Benchmark Configs.
Pipeline¶
| Parameter | Type | Default | Description |
|---|---|---|---|
width |
int |
4 |
Instructions per cycle for every stage that has no width of its own |
fetch_width, decode_width, rename_width, issue_width, commit_width |
int |
width |
Per-stage widths |
writeback_width |
int |
width |
Results the out-of-order backend writes back per cycle; the rest wait for later cycles |
trap_latency |
int |
13 |
Cycles from commit detecting a trap or interrupt to the squash into its handler; an interrupt first lets everything already fetched retire |
redirect_latency |
int |
2 (O3), 1 (in-order) |
Cycles from execute resolving a misprediction, CSR write, fault or ordering violation to the squash into the redirect; commit retires nothing the pending squash will remove |
store_forward_latency |
int |
L1D hit latency | Cycles from a load matching a store in the store buffer to its data reaching writeback, where a load the L1D answers takes the L1D hit latency; 0 writes the load back in the cycle it matches (a forwarded vector span takes at least one cycle) |
backend |
Backend.* |
OutOfOrder() |
Backend.OutOfOrder(...) or Backend.InOrder() |
branch_predictor |
BranchPredictor.* |
TAGE() |
Direction predictor (see below) |
btb_size |
int |
4096 |
Branch target buffer entries |
btb_ways |
int |
4 |
BTB associativity; btb_size / btb_ways must be a power of two |
ras_size |
int |
32 |
Return address stack depth |
mem_dep_predictor |
MemDepPredictor.* |
StoreSet() |
Memory dependence predictor (see below) |
Backend: Out-of-Order¶
Backend.OutOfOrder(
rob_size=128, # Reorder buffer entries
issue_queue_size=32, # Unified issue queue entries (CAM wakeup/select)
store_buffer_size=32, # Store buffer entries
load_queue_size=32, # Load queue entries
load_ports=2, # Loads issued per cycle
store_ports=1, # Stores issued per cycle
prf_gpr_size=256, # Physical integer registers
prf_fpr_size=128, # Physical floating-point registers
fu_config=Fu([...]), # Functional unit pool (see below); Fu() by default
checkpoint_count=0, # Rename-map checkpoints for branch recovery (0: rebuild from the ROB)
squash_width=8, # ROB entries commit squashes per cycle after a squash
prf_vpr_size=64, # Physical vector registers
vec_chaining=True, # Let a dependent vector op start on the first element group
vec_store_buffer_size=8, # In-flight vector stores
vec_store_forwarding="byte_mask", # Vector store-to-load forwarding: "byte_mask", "stall" or "off"
)
- Register files.
prf_gpr_sizeandprf_fpr_sizemust each hold the 32 architectural registers plus one for every instruction that can be in flight with a destination; 32 +rob_sizealways suffices. - Squash recovery. After a misprediction, trap or ordering violation,
commit squashes the flushed ROB entries
squash_widthper cycle and rename waits until it has finished, one cycle more; the rename map itself is restored at once, from a checkpoint when the squashing branch has one. - Stores issue in two halves when their data is not ready: the address as soon as the base register is, the data when the value is (see Pipeline).
- Vector store forwarding.
byte_maskforwards any bytes a vector store holds, as most out-of-order cores do;stallmakes an overlapping load wait for the store to be written, as Saturn does;offtreats every overlap as a stall.
Backend: In-Order¶
The in-order backend takes no parameters of its own. It issues in program
order, up to issue_width per cycle, with a scoreboard tracking operands;
its functional units are the default Fu() pool, and it has a 64-entry
ROB and a 16-entry store buffer.
Functional Units¶
from rvsim import Fu
fu = Fu([
Fu.IntAlu(count=4, latency=1), # add, sub, logic, shift, compare
Fu.IntMul(count=1, latency=3), # multiply (pipelined)
Fu.IntDiv(count=1, latency=35), # divide and remainder (not pipelined)
Fu.FpAdd(count=2, latency=4), # FP add, subtract, compare, convert
Fu.FpMul(count=2, latency=5), # FP multiply
Fu.FpFma(count=2, latency=5), # FP fused multiply-add
Fu.FpDivSqrt(count=1, latency=21), # FP divide and square root (not pipelined)
Fu.Branch(count=2, latency=1), # branch and jump resolution
Fu.Mem(count=2, latency=1), # load and store address generation
Fu.VecIntAlu(count=1, latency=1), # vector integer arithmetic and logic
Fu.VecIntMul(count=1, latency=3), # vector integer multiply
Fu.VecIntDiv(count=1, latency=20), # vector integer divide (not pipelined)
Fu.VecFpAlu(count=1, latency=4), # vector FP add, compare, convert
Fu.VecFpFma(count=1, latency=5), # vector FP multiply and fused multiply-add
Fu.VecFpDivSqrt(count=1, latency=20), # vector FP divide and square root (not pipelined)
Fu.VecMem(count=1, latency=1), # vector load and store address generation
Fu.VecPermute(count=1, latency=1), # slides, gathers, compress, moves
])
The list above is Fu(), the default. A scalar unit type left out of a
Fu list has no units, so an instruction that needs one never issues:
include every scalar type your workload uses. A vector unit type left out
gets one unit with the default latency. A vector instruction's time on its
unit also scales with vl over the number of lanes (num_vec_lanes,
below).
Branch Prediction¶
BranchPredictor.Static() # Always predicts not-taken
BranchPredictor.GShare() # PC XOR global history, 2-bit counters
BranchPredictor.Tournament( # gem5's TournamentBP (Alpha 21264): local, global, choice
global_size_bits=12,
local_hist_bits=10,
local_pred_bits=10,
)
BranchPredictor.Perceptron( # Perceptron predictor
history_length=32,
table_bits=10,
)
BranchPredictor.TAGE( # TAGEBase-style TAGE (defaults shown)
num_banks=8,
table_size=2048,
reset_interval=256_000,
history_lengths=[5, 11, 22, 44, 89, 178, 356, 712],
tag_widths=[8, 8, 9, 9, 10, 10, 11, 11],
)
BranchPredictor.ScLTage() # 64KB TAGE-SC-L with ITTAGE (Seznec's CBP-5 configuration)
TAGE and ScLTage take further keywords for every structure of the
predictor (allocation and update rules, history kind, hashing, banking,
the bimodal table, USE_ALT_ON_NA counters); ScLTage adds the loop
predictor (loop_*), the statistical corrector (sc_*, with
BranchPredictor.ScGehl and BranchPredictor.ScLocalGehl components) and
ITTAGE (ittage_*). Their defaults reproduce Seznec's 64KB TAGE-SC-L.
Branch Prediction describes each
parameter.
Memory Dependence Prediction¶
Controls how a load decides whether it may issue ahead of older stores whose addresses are not known yet.
MemDepPredictor.Blind() # Loads wait for every older store's address
MemDepPredictor.StoreSet( # Store-set predictor (Chrysos & Emer 1998), the default
ssit_size=1024, # Store Set ID Table entries (PC -> store set)
lfst_size=1024, # Last Fetched Store Table entries (store set -> last store)
)
The store-set predictor learns from ordering violations and wipes both tables every 250,000 memory instructions.
Caches¶
Each level is configured independently. The builder's own defaults (a
4 KiB direct-mapped cache) are shown; Config's default levels are in the
table below.
Cache(
size="4KB", # "4KB", "32KB", "1MB", or bytes
line="64B", # Line size; at least the 64-byte block CBOs act on
ways=1, # Associativity
latency=1, # Tag and data access latency in cycles
response_latency=1, # Cycles from a fill arriving to answering its requests
mshr_count=0, # Lines fetched at once (0 = the default, 8)
write_buffers=0, # Evicted lines in flight to the next level (0 = the default, 8)
targets_per_mshr=0, # Requests one MSHR can hold (0 = the default, 20)
policy=None, # Eviction policy; ReplacementPolicy.LRU() when None
prefetcher=None, # Hardware prefetcher; Prefetcher.Off() when None
)
| Parameter | Type | Default | Description |
|---|---|---|---|
l1i |
Cache or None |
32 KiB, 4-way, 1 cycle, next-line prefetch | L1 instruction cache |
l1d |
Cache or None |
32 KiB, 4-way, 1 cycle, stride prefetch | L1 data cache |
l2 |
Cache or None |
256 KiB, 8-way, 10 cycles | Private L2 |
l3 |
Cache or None |
None |
Shared last-level cache |
inclusion_policy |
Cache.* |
Cache.NINE() |
Relationship between L1 and L2 |
wcb_entries |
int |
0 |
Write-combining buffer entries between the store buffer and the L1D (0 = none) |
load_prefetcher |
LoadPrefetcher.Stride or None |
None |
The load/store unit's load prefetcher (see below) |
store_prefetcher |
StorePrefetcher.Stream or None |
None |
The L1D's store-miss prefetcher, which fills the L2 (see below) |
A load that hits takes one cycle of address generation plus the L1D's
latency to reach its dependents, so latency=3 models a 4-cycle
load-to-use. None disables a level.
MSHRs and writeback buffers
Every level fetches at most mshr_count lines at a time and keeps at
most write_buffers evicted lines in flight to the next level; while
either is exhausted, or one MSHR holds targets_per_mshr requests, the
cache blocks and later requests queue. Passing 0 leaves the simulator
default in place (8, 8 and 20); mshr_count=1 gives a blocking cache
that serialises its misses.
Replacement Policies¶
ReplacementPolicy.LRU() # Least recently used (default)
ReplacementPolicy.PLRU() # Tree pseudo-LRU
ReplacementPolicy.FIFO() # First in, first out
ReplacementPolicy.Random() # Random eviction
ReplacementPolicy.MRU() # Most recently used
Prefetchers¶
Prefetcher.Off() # None (the default for Cache())
Prefetcher.NextLine(degree=1) # The next `degree` lines on every access
Prefetcher.Stride(degree=1, table_size=64) # PC-indexed constant-stride detection
Prefetcher.Stream(degree=1) # Ascending or descending streams
Prefetcher.Tagged(degree=1) # Next lines on a miss or a first use of a prefetched line
A prefetch is a real fetch: it takes an MSHR (never the last free one) and travels down the hierarchy like a demand miss. A cache sees physical addresses only, so its prefetcher never crosses the 4 KiB page of the access that triggered it.
Load and store prefetchers¶
The L1D's prefetching on a real core lives in the load/store unit, where each load's PC, virtual address and translation are known. These follow the Cortex-A72's documented prefetcher; the memory hierarchy page gives the design and its sources.
LoadPrefetcher.Stride(
table_size=64, # PC-indexed entries, a power of two
l1_lines=4, # Lines kept ahead in the L1D
l2_lines=0, # Lines kept ahead in the L2 alone (the A72 keeps 22)
page_boundary=None, # PageBoundary.Stop() (when None) or PageBoundary.CrossWithTlb()
)
StorePrefetcher.Stream(
streams=4, # Runs of store misses tracked at once
l2_lines=8, # Lines kept ahead in the L2, with write permission
)
PageBoundary.Stop() keeps a stream inside the page of the load that
trained it, at that page's size; PageBoundary.CrossWithTlb() continues
into the next page when the data TLB holds its translation and drops the
prefetch when it does not. Both are passed to Config:
Config(
load_prefetcher=LoadPrefetcher.Stride(l1_lines=1, l2_lines=22,
page_boundary=PageBoundary.CrossWithTlb()),
store_prefetcher=StorePrefetcher.Stream(),
)
Inclusion Policies¶
Cache.NINE() # Neither inclusive nor exclusive (default)
Cache.Inclusive() # An L2 eviction back-invalidates the L1 copies
Cache.Exclusive() # L1 victims go to the L2; an L1 fill takes the L2's copy
Memory and Translation¶
| Parameter | Type | Default | Description |
|---|---|---|---|
ram_size |
str or int |
"256MB" |
Main memory size |
memory_controller |
MemoryController.* |
Simple() |
Memory controller (see below) |
tlb_size |
int |
64 |
Entries in each of the instruction and data L1 TLBs |
tlb_ways |
int |
0 |
L1 TLB associativity; 0 is fully associative |
l2_tlb_size |
int |
0 |
Entries in the shared L2 TLB; 0 disables it |
l2_tlb_ways |
int |
4 |
L2 TLB associativity |
l2_tlb_latency |
int |
4 |
L2 TLB hit latency in cycles |
paging_mode_max |
str |
"sv57" |
Strongest paging mode satp accepts ("bare", "sv39", "sv48", "sv57"); a stronger mode written to satp reads back as Bare, which makes a kernel fall back |
misaligned_access_trap |
bool |
False |
Raise address-misaligned exceptions instead of performing misaligned accesses in hardware |
svadu |
bool |
False |
Implement Svadu: with menvcfg.ADUE set the page-table walker sets A and D bits itself; otherwise a missing A or D bit faults (Svade) |
Memory Controllers¶
MemoryController.Simple( # Fixed latency (default), serialised on a bandwidth:
latency=120, # Core cycles from the controller starting a request to its data
bandwidth_gib_s=12.8, # each request busies the controller for its bytes' time
)
MemoryController.DRAM( # Row-buffer DRAM: per-bank open rows and refresh
t_cas=14, # Column access, cycles
t_ras=14, # Row activate, cycles
t_pre=14, # Precharge, cycles
)
MemoryController.DDR5( # Command-level JEDEC DDR5 (see Memory Hierarchy)
speed_bin="4800B", # "4800B" or "5600B"
channels=2, # Channels; each has two 32-bit sub-channels
subchannels_per_channel=2,
ranks_per_channel=2,
bank_groups_per_rank=8,
banks_per_group=4,
row_bits=16,
column_bits=6, # Rows of 64 << column_bits bytes
read_queue_entries=64,
write_queue_entries=64,
write_high_watermark=54, # Start draining writes at this depth
write_low_watermark=32, # Return to reads at this depth
min_writes_per_switch=16,
frontend_latency_ns=10, # Controller pipeline
backend_latency_ns=10,
scheduler="FrFcfs", # or "Fcfs"
refresh="AllBank", # or "SameBank"
address_mapping="RoRaBaChCo", # or "RoRaBaCoCh", "RoCoRaBaCh"
power_down_idle_ns=None, # e.g. 200 to enable rank power-down
ecc="None", # "SecDed" or "ChipKill"
patrol_scrub_ns=None, # e.g. 100_000 to enable patrol scrubbing
timing=None, # Per-field overrides in DRAM command clocks, e.g. {"t_rcd": 40}
)
On the DRAM controller a row hit costs t_cas and a row miss
t_pre + t_ras + t_cas. The DDR5 controller runs at the DRAM command clock
(half the data rate) and converts to and from the core clock through cpu_clock_mhz; its statistics appear under
memctrl0.ch<C>.sc<S>.*. Every controller sits behind the system bus, so a
miss to memory also pays bus_latency each way.
Vector Extension¶
| Parameter | Type | Default | Description |
|---|---|---|---|
vlen |
int |
128 |
Vector register length in bits, a power of two from 128 to 2048 |
num_vec_lanes |
int |
vlen / 64, at least 1 |
64-bit lanes the vector units process per cycle |
vector_mem_width |
int |
vlen / 8, at most 64 |
Bytes one unit-stride vector memory access moves (the vector load-store datapath), a power of two from 8 to 64 |
ELEN is 64 and Zvfh is implemented.
System¶
These parameters set the SoC's memory map, clocks and devices. You
normally need to change only cpu_clock_mhz, hart_count and the
console.
| Parameter | Type | Default | Description |
|---|---|---|---|
cpu_clock_mhz |
int |
2400 |
Core clock: converts between cycles and nanoseconds for the DDR5 controller, device latencies and the RTC |
hart_count |
int |
1 |
Harts in the system, one per core (see Multi-core) |
bus_width |
int |
8 |
System bus width in bytes |
bus_latency |
int |
4 |
System bus latency in cycles, each way |
device_latency_ns |
int |
100 |
Time every device takes to answer a register access |
device_latency_ns_overrides |
dict |
None |
Per-device access latency by name (UART0, CLINT, PLIC, VirtIO-Blk, SysCon, GoldfishRTC, HTIF) |
clint_divider |
int |
10 |
CPU cycles per mtime tick |
rtc_epoch_seconds |
int |
1767225600 |
Wall-clock time the RTC reports at cycle zero (2026-01-01), advanced by simulated time so runs are reproducible |
ram_base |
int |
0x8000_0000 |
RAM base address |
uart_base |
int |
0x1000_0000 |
UART base address |
disk_base |
int |
0x9000_0000 |
VirtIO disk base address |
clint_base |
int |
0x0200_0000 |
CLINT base address |
syscon_base |
int |
0x0010_0000 |
SYSCON base address |
sim_control_base |
int |
0x0010_2000 |
Sim-control device base address (guest statistics reset, dump and exit) |
kernel_offset |
int |
0x0020_0000 |
Kernel load offset from ram_base |
Multi-core¶
hart_count=N builds N single-threaded cores, each with its own
pipeline, branch predictor, TLBs and private L1 and L2, sharing the LLC,
memory and devices. Every hart has its own CLINT timer and
software-interrupt registers and its own PLIC contexts, and the generated
device tree enumerates them. Bare-metal programs start every hart at the
entry point with a0 holding the hart id and a1 the hart count.
from rvsim import Config, Coherence, HomeAgent, Interconnect
config = Config(
width=4,
hart_count=4,
coherence=Coherence(
home_agent=HomeAgent.SnoopFilter(capacity_factor=1.5, ways=8),
interconnect=Interconnect.Mesh(hop_latency=2, bytes_per_cycle=32),
),
)
With more than one core the private L2s become requesting agents on a
coherence fabric: MESI states in every private cache, a home agent at the
LLC that serialises requests per line and decides who is snooped, and an
interconnect that carries request, snoop, response and data messages on
separate virtual channels. The L2 is made inclusive of its L1s so snoops
are answered from its tags; Cache.Exclusive() is therefore rejected
with hart_count > 1. A single core builds no fabric.
| Parameter | Type | Default | Description |
|---|---|---|---|
coherence.home_agent |
HomeAgent.* |
HomeAgent.SnoopFilter() |
Who must be snooped for a request |
coherence.interconnect |
Interconnect.* |
Interconnect.Crossbar() |
Message transport between the L2s and the home |
coherence.txn_entries |
int |
32 |
Transactions the home can have live at once |
Home agents¶
HomeAgent.SnoopFilter(capacity_factor=1.5, ways=8) # exact sharers and owner per tracked line (default)
HomeAgent.Broadcast() # track nothing; snoop every other core
The snoop filter tracks capacity_factor times the aggregate private L2
lines in a ways-way set-associative array. When a set is full, its least
recently used line is recalled (every holder invalidated) before a new
line is tracked, as Arm's snoop filter and AMD's probe filter do.
Interconnects¶
Interconnect.Crossbar(hop_latency=2, bytes_per_cycle=32) # any port to any port (default)
Interconnect.Ring(hop_latency=2, bytes_per_cycle=32) # bidirectional ring, shorter direction
Interconnect.Mesh(hop_latency=2, bytes_per_cycle=32) # square 2-D mesh, XY routing
Interconnect.Torus(hop_latency=2, bytes_per_cycle=32) # mesh with wraparound
Interconnect.Hypercube(hop_latency=2, bytes_per_cycle=32) # dimension-order routing
hop_latency is the cycles a message spends per hop and bytes_per_cycle
the width of a port or link; a 64-byte data message on a 32-byte link
occupies it for two cycles. The crossbar is one hop; the routed networks
place the cores and the home on their nodes and charge every hop.
The fabric reports under coherence.ha.* (requests by kind, snoops,
cache-to-cache transfers, recalls, transaction latency) and
coherence.interconnect.* (messages, bytes, busy and blocked cycles);
each private cache counts its snoops under
core<N>.cache.<level>.coherence.*. See Multi-core.
General¶
| Parameter | Type | Default | Description |
|---|---|---|---|
trace |
bool |
False |
Emit a trace event at every pipeline stage an instruction passes. RUST_LOG selects which are printed: rvsim=trace for all, or targets such as rvsim::commit=trace, rvsim::mem=trace and rvsim::fwd=trace |
initial_sp |
int or None |
None |
Stack pointer a bare-metal program starts with; ram_base + 16 MiB when unset |
console |
str or None |
None |
Where the UART connects: "stdout", "stderr", "quiet", or "captured" (kept in memory for read_console(), input given with write_console()); overrides the two shorthands below |
uart_quiet |
bool |
False |
Shorthand for console="quiet" |
uart_to_stderr |
bool |
False |
Shorthand for console="stderr" |
Example Configurations¶
Minimal embedded core¶
Config(
width=1,
backend=Backend.InOrder(),
branch_predictor=BranchPredictor.Static(),
l1d=Cache("4KB", ways=1, latency=1),
l1i=Cache("4KB", ways=1, latency=1),
l2=None,
)
High-performance out-of-order core¶
Config(
width=4,
backend=Backend.OutOfOrder(
rob_size=128,
issue_queue_size=48,
load_queue_size=32,
store_buffer_size=32,
prf_gpr_size=256,
prf_fpr_size=128,
checkpoint_count=32,
fu_config=Fu([
Fu.IntAlu(count=4, latency=1),
Fu.IntMul(count=1, latency=3),
Fu.IntDiv(count=1, latency=35),
Fu.FpAdd(count=2, latency=4),
Fu.FpMul(count=2, latency=5),
Fu.FpFma(count=2, latency=5),
Fu.FpDivSqrt(count=1, latency=21),
Fu.Branch(count=2, latency=1),
Fu.Mem(count=2, latency=1),
]),
),
branch_predictor=BranchPredictor.ScLTage(),
mem_dep_predictor=MemDepPredictor.StoreSet(),
l1d=Cache("32KB", ways=8, latency=3, mshr_count=8,
prefetcher=Prefetcher.Stride(degree=2, table_size=128)),
l1i=Cache("32KB", ways=8, latency=1,
prefetcher=Prefetcher.NextLine(degree=2)),
l2=Cache("256KB", ways=8, latency=12, mshr_count=16),
l3=Cache("4MB", ways=16, latency=30, mshr_count=32),
memory_controller=MemoryController.DDR5(speed_bin="5600B"),
l2_tlb_size=1024,
)
Linux-capable system¶
presets.linux() places a core in a system that boots the bundled Linux
image; see Linux Boot.