Memory Hierarchy¶
rvsim models a complete memory hierarchy from TLBs through L3 cache to DRAM, with configurable parameters at every level. Every cache, the bus and the memory controller is an event-driven component exchanging request and response packets, so each access pays the latency and waits for the occupancy of every component it passes. Caches hold tags and coherence state only; data lives in one memory image and an access takes effect where it is served (decision 11).
Overview¶
flowchart TD
CPU["CPU Pipeline"] --> ITLB["I-TLB\n64 entries · fully assoc."] & DTLB["D-TLB\n64 entries · fully assoc."]
ITLB --> L1I["L1-I Cache"]
DTLB --> L1D["L1-D Cache"]
ITLB & DTLB -->|miss| L2TLB["L2 TLB\noptional · off by default"]
L2TLB -->|miss| PTW["Hardware PTW\nSv39 / Sv48 / Sv57"]
PTW -->|PTE reads| L1D
L1I & L1D -->|miss| MSHR["MSHRs\ncoalescing"]
MSHR --> L2["L2 Cache"]
L2 -->|miss| L3["L3 Cache"]
L3 -->|miss| MC["Memory Controller"]
MC --> DRAM["DRAM\nrow-buffer timing"]
L1D <--> STB["Store Buffer\nforwarding · WCB"]
Virtual Memory¶
The MMU implements the RISC-V Sv39, Sv48 and Sv57 paging modes, with
three, four and five levels of page table. paging_mode_max caps the
modes satp accepts: writing a stronger mode leaves satp reading back
as Bare, which is how a kernel probes for the deepest mode and falls back.
- Pages. 4 KiB base pages and every superpage size the mode allows (2 MiB, 1 GiB, 512 GiB, 256 TiB). A TLB entry maps a whole page of the size its leaf was found at, so a 2 MiB kernel mapping is one entry.
- L1 TLBs. Separate instruction and data TLBs of
tlb_sizeentries (default 64), fully associative whentlb_waysis 0 (the default, as gem5's RISC-V TLB) and LRU within a set otherwise. - L2 TLB. An optional TLB shared by the core's instruction and data
sides (
l2_tlb_size,l2_tlb_ways), hitting afterl2_tlb_latencycycles. It is off by default, as gem5 has none. - Page-table walker. A miss in every TLB starts a hardware walk. Each level's PTE is an 8-byte read sent to the L1D, so page-table entries are cached like data and a walk's cost depends on where they are found. Superpage alignment, reserved bits, the U, SUM and MXR rules and the PMP check on each PTE are all applied.
- Accessed and dirty bits. By default a page whose A bit, or D bit on
a store, is clear raises a page fault and the kernel sets the bit
(Svade). With
svadu=Trueandmenvcfg.ADUEset, the walker sets the bit itself and writes the updated PTE back (Svadu). - Flushes.
SFENCE.VMAflushes the TLBs at commit, by address and ASID when it names them.
Translation is skipped when satp.MODE is Bare or the hart runs in M-mode
(with mstatus.MPRV clear for loads and stores).
Cache Hierarchy¶
L1 Instruction Cache¶
Accessed by the Fetch1 stage. Configurable size, associativity, latency, and replacement policy. Supports hardware prefetching (typically next-line).
Invalidated by:
FENCE.Iinstruction (deferred to commit, drains store buffer first)- Inclusive L2 eviction back-invalidation (if inclusion policy is Inclusive)
L1 Data Cache¶
Accessed by the Memory1 stage. The critical path for load-to-use latency. An atomic (AMO) is one access: the cache takes the line writable, the read-modify-write is performed on it and it is left modified, so commit has no separate store to send.
Every level: MSHRs, writeback buffer, blocking¶
Each cache level is one event-driven component, modelled after gem5's classic cache:
- A miss allocates an MSHR and sends one line-sized request to the
next level after the tag-lookup latency. A second miss to a line already
in flight joins that MSHR instead of fetching again; when the fill
arrives every joined request is answered at once.
mshr_countbounds the fetches in flight (default 8;mshr_count=1gives a blocking cache, and0from Python leaves the default). An MSHR holds at mosttargets_per_mshrrequests (gem5'stgts_per_mshr, default 20): the request that fills it blocks the cache until that line's fill returns. - A write miss allocates: the line is fetched, then installed dirty. A line written back from above merges into a held line or is forwarded without allocating.
- A hart's access takes effect where it is served, as gem5's cache satisfies a request. The caches hold tags only; the data lives in one memory image, and a load reads it, a store writes it, when the first cache holding the line with the permission the access needs serves the request: a hit as it arrives, a miss when its fill does. A request no cache serves takes effect at the memory controller. Line fills and writebacks between levels move permission and timing only.
- A fill is forwarded to every request the MSHR gathered as it is
written into the array,
response_latencycycles after it arrives (gem5'sresponse_latency, default 1), rather than after a second array access. - A fill that evicts a dirty victim puts it in the writeback buffer and sends it to the next level; the entry is freed when that level acknowledges. A fill whose dirty victim finds the buffer full waits, holding its MSHR, until a writeback is acknowledged, as a real core's linefill does. A dirty line a probe or back-invalidation demands goes back on the snoop-response path and takes no buffer slot. Dirty lines leaving the last cache reach the memory controller as writes, so DRAM sees the real write traffic.
- While every MSHR or every writeback buffer entry (
write_buffers, default 8) is busy, or one MSHR holds its target limit, the cache is blocked: new requests queue in arrival order and are retried as entries free up, which is what a blocked port does to its requester. - Prefetches are real fetches: a candidate line the prefetcher wants takes an MSHR (never the last free one) and travels down the hierarchy like a demand miss. A request that joins it in flight makes it late; the first request to find its line once installed makes it useful; its line dropped before any request found it makes it unused.
Per-level counters live under core<N>.cache.{l1i,l1d,l2} and llc:
hits, misses, mshr_hits, blocked_requests, fills, evictions,
writebacks, back_invalidations, prefetches.issued,
prefetches.late, prefetches.useful, prefetches.unused, and the
derived prefetches.used, prefetches.accuracy and miss_rate.
A load's dependents wake when its data returns: one cycle of address
generation plus the L1D's latency on a hit, or whenever the fill
arrives on a miss. Neither backend issues dependents speculatively on a
predicted hit.
L2 / L3 Caches¶
Unified caches accessed on L1 miss. Each level has independent size, associativity, latency, replacement policy, and prefetcher configuration.
Inclusion Policies¶
The relationship between adjacent levels is configurable:
| Policy | Behavior | Trade-off |
|---|---|---|
| NINE (default) | No inclusion enforcement | Simple, no back-invalidation traffic |
| Inclusive | An eviction back-invalidates the same line in the caches above; a dirty copy above is written back first | Guarantees each level is a superset of the levels above it, which a snooping lower level needs |
| Exclusive | L1 victims (clean or dirty) are handed to the L2; the L2 gives up its copy when it fills an L1 | Maximizes effective L1+L2 capacity; the LLC stays non-inclusive |
An exclusive L2 keeps a shadow tag for each line it handed up without keeping, so it neither prefetches a line held above nor loses track of it: the line comes back as a victim, or the tag is dropped when the line is flushed or invalidated from above.
Invariant audit¶
Simulator::set_audit_caches(true) (Python: sim.audit_caches = True)
checks every cache invariant after every event: no duplicate tags, MSHRs
or writebacks, no MSHR over its target limit, no more MSHRs or eviction
writebacks than the cache has entries, inclusion as each level's policy
sets it, every copy above a level recorded by it, and, with several cores,
the coherence invariants. Lines with a request, fill, probe or writeback
in flight are left out of the cross-level checks. The first broken
invariant ends tick() with SimError::CacheInvariant;
Simulator::cache_violations lists them all. The audit walks every cache
after each event, so it is off by default and costs nothing then.
Store Buffer¶
The store buffer sits between the pipeline and L1D, holding stores that have executed but not yet committed.
- Store-to-load forwarding — when a load address matches a pending store in the buffer, the data is forwarded directly without accessing L1D. Supports full and partial overlap detection.
- Commit-time draining — a store is written to L1D only after it commits, one per cycle from the head of the buffer, and keeps its slot (and keeps forwarding) until the L1D has performed it and acknowledged
- Write-combining buffer (WCB) — optional merging write buffer (
wcb_entries) between the store buffer and the L1D. A committed store merges into the entry for its line, which the hart's loads read from; the entry is written to the L1D as one masked line write when a new line needs its slot, when its line is fully written, when a load needs bytes it holds only some of, or when the store buffers leave the write port idle. A sent line keeps forwarding until the L1D acknowledges its write, since the cache can still serve a load from its old copy of the line while it fetches write permission. Barriers wait for its lines like any other committed store
Cache-block operations¶
cbo.clean, cbo.flush and cbo.inval (Zicbom) and cbo.zero (Zicboz)
act on a 64-byte block, which every cache line must hold whole. A CBO
translates in memory1 and takes a store-buffer slot, so it drains to the
L1D in order with the stores around it after it commits, and barriers wait
for it like a store.
cbo.zerois the hart's write of a zeroed block, taking effect where the L1D serves it.- The management operations follow gem5's
CleanSharedReq,CleanInvalidReqandInvalidateReqto the point of coherence: each cache on the way applies the operation to its copy (a clean keeps the line clean, a flush or invalidate drops it and the inclusive copies above it) and passes it on, carrying any dirty data it found, which an invalidate discards. A coherent L2 sends the home agent a maintenance request; the home snoops the other harts (the owner is cleaned for a clean, every holder invalidated for a flush, dropped without its data for an invalidate), passes the operation to the LLC, and completes the requester once memory acknowledges it. The memory controller counts a line write when dirty data arrives with it. - Because caches hold tags only, an invalidate cannot lose data: memory
keeps the latest value, which the specification allows since
cbo.invalmay perform a flush. - Device regions do not support CBOs: every CBO,
cbo.zeroincluded, to a device address raises a store access fault.
Hardware Prefetching¶
Prefetchers sit where the hardware puts them and know only what it knows
there (decision 14).
The design follows the Cortex-A72's load/store hardware prefetcher, the
one core in the presets whose prefetchers are documented: Arm, Cortex-A72
MPCore Processor Technical Reference Manual r0p3, §6.4.9 (load/store
hardware prefetcher), §4.3.66 (CPUACTLR_EL1) and §4.3.67
(CPUECTLR_EL1); and on Intel's description of its prefetchers in the
64 and IA-32 Architectures Optimization Reference Manual Vol. 1
(248966-049), §9.5.2 and §4.1.7. What follows says, for each part, which
behaviour comes from those manuals and which is rvsim's own choice.
flowchart LR
LSU["Load/store unit<br/>load prefetcher<br/>(PC, VA, DTLB)"] -- "Prefetch into L1D" --> L1D
LSU -- "Prefetch into L2" --> L1D
L1D -- "passes it down" --> L2
L1D -- "store misses<br/>(ReadOwn)" --> SP["L1D store prefetcher<br/>(PA, 4 KiB page)"]
SP -- "Prefetch into L2, exclusive" --> L2
Each cache may also run a cache-side prefetcher of its own on the physical addresses it sees.
Prefetch requests¶
A prefetch travels as a MemReq with MemOp::Prefetch { into, exclusive }
from the load/store unit or the L1D. Caches above into pass it down
unchanged; the cache at into starts a prefetch fetch of the line unless
it already holds it (with write permission, for an exclusive prefetch),
is already fetching it or is writing it back. Nothing answers a prefetch.
A cache drops one rather than give up its last free MSHR
(prefetches.dropped), so prefetches never block demand misses, and a
disabled level passes on prefetches meant for the level below it and
drops its own.
Load prefetcher (load/store unit)¶
Configured with Config(load_prefetcher=LoadPrefetcher.Stride(...)).
- Where it trains. In memory1, on every load that goes to the memory system, whether the L1D or store-buffer forwarding answers it: scalar loads, vector spans and vector elements. It sees the load's PC, its virtual address and its physical address. A load that waits to retry does not train until it goes. (Source: the A72's prefetcher is part of the load/store unit, §6.4.9.)
- How it detects streams. A reference prediction table of
table_sizeentries, direct-mapped and tagged on the load's PC, holds each load's last virtual address, stride and a 2-bit saturating confidence. A repeated stride raises the confidence; a different one lowers it, and replaces the stride once it reaches zero. A stream is confident once the same nonzero stride has followed a saturated confidence, that is from a load's sixth access at one stride. A load whose PC maps to another load's entry takes it over. (rvsim's choice: neither manual describes the detection algorithm; this is Chen and Baer's reference prediction table, as gem5'sStridePrefetcheruses. The cache-side stride prefetcher shares the same rule.) - How far ahead. A confident stream keeps
l1_lineslines ahead of the demand access in the L1D and, beyond them,l2_lineslines ahead in the L2 alone. Each line is requested once: a stream remembers the furthest line it has requested at each level, and starts again when the load leaves that window (a second pass over the same array). (Source for the L2 distance:CPUECTLR_EL1[33:32], "the number of requests by which the prefetch request to the L2, on a load stream, is ahead of the demand request stream", 16 to 22, reset 22. The L1D distance is not published.) - Line granularity. Prefetches go out a line at a time, so a stride shorter than a line advances one line per prefetch rather than naming the line the load is already in, and a longer stride names the line it lands in.
- Page boundaries.
page_boundary=PageBoundary.Stop()keeps every prefetch in the page of the load that trained it. The page's size comes from the data TLB's entry for it (4 KiB, 2 MiB, 1 GiB...), and is 4 KiB when translation is off or the entry has gone.PageBoundary.CrossWithTlb()continues into the next page when the data TLB already holds its translation, looked up without disturbing the TLB's replacement state and without starting a walk; on a miss the prefetch is dropped. With translation off the address is physical and crossing needs no lookup. (Source:CPUACTLR_EL1[43]— reset 0, "Enables the Load/Store hardware prefetcher to use VA in generating prefetches that can cross page boundaries"; set, "prefetch is restricted to within the page boundary of the demand request". Intel's Gracemont prefetcher crosses pages in the linear address space and "start[s] translations for TLB misses"; rvsim drops instead, as the A72 manual does not say it walks.) - What it may touch. A prefetch whose page the load could not read
(permissions,
mstatus.SUM/MXR, PMP at the load's effective privilege) or whose line is not RAM is dropped, so a prefetch never reaches a device. - Where a level stops. A level stops at its first line it cannot place and picks up from that line on the load's next access.
Stats under core<N>.prefetch.loads: l1 and l2 (prefetches sent to
fill each level) and dropped.page_boundary, dropped.tlb_miss,
dropped.denied, dropped.not_ram.
Store prefetcher (L1D)¶
Configured with Config(store_prefetcher=StorePrefetcher.Stream(...)).
It watches the L1D's store misses that start a fetch for write permission
(ReadOwn, the ReadUnique of the coherence protocol), finds runs of
misses to adjacent lines inside one 4 KiB physical page, tracking
streams runs at once, and once a run has gone two lines in one direction
keeps it l2_lines lines ahead with exclusive prefetches into the L2,
each line once. prefetches.store_stream on the L1D counts them.
(Source: §6.4.9, "Prefetching on store accesses is managed by a PA based
prefetcher and only prefetches to the L2 cache", and CPUACTLR_EL1[42],
prefetch requests "generated by ReadUnique transactions". The run
detection and its length are rvsim's choice. Stores drain after commit as
merged lines with no translation attached, so the prefetcher keeps to the
smallest page.)
Cache-side prefetchers¶
Each cache level can also have a prefetcher of its own
(Cache(prefetcher=...)). A cache sees physical addresses only, and the
physical page after the one an access touches may belong to anything, so
like a hardware PA prefetcher it drops every candidate outside the 4 KiB
page of the access that produced it (prefetches.page_crossing). (Source:
Intel, "it will not prefetch across a 4-KByte page boundary"; the A72's
PA mode keeps to the page.)
| Prefetcher | How it works |
|---|---|
| NextLine | On any access, prefetch the next degree cache lines |
| Stride | The reference prediction table above, kept in the cache: demand loads and fetches carry their PC to the cache, and stores, page walks and writebacks do not, so they do not train it. A confident stream prefetches the next degree lines along its stride, a line at a time. gem5's StridePrefetcher is this design, and the gem5 comparison uses it |
| Stream | Detects ascending or descending runs of consecutive lines and prefetches degree lines ahead in that direction |
| Tagged | Prefetches the next line on a demand miss, and again when a demand access first uses a prefetched line, so a useful stream keeps extending |
The prefetcher sees every demand access to its cache. A candidate is dropped when its line is already present, already being fetched or being written back, so no level fetches a line twice, and when it would take the last free MSHR; each level's prefetcher works independently.
In the presets¶
| Preset | L1D prefetching |
|---|---|
cortex_a72() |
Load prefetcher, l2_lines=22 and CrossWithTlb (the reset values of CPUECTLR_EL1[33:32] and CPUACTLR_EL1[43]); store prefetcher into the L2. The L1D distance (1 line), table size (32) and store run length (8 lines, 4 runs) are not published |
p550() |
SiFive has not published the P550's prefetchers: a load prefetcher that keeps to the page (Stop, 1 line ahead, no L2 stream) and no store prefetcher, as the cautious reading |
fast(), m1(), basic() |
The cache-side stride prefetcher on the L1D, as before; Apple's prefetchers are not published either |
DRAM Controller¶
Three memory controllers are available; all sit behind the L3 (or the last enabled cache level) and the system bus.
Simple controller (the default) — every access takes latency cycles
(120 by default) once the controller is free: each request busies it for the time its bytes take
at bandwidth_gib_s (gem5's SimpleMemory), and later requests wait.
DRAM controller — models row-buffer aware timing over 8 banks of 2 KiB rows:
- Row hit:
t_cascycles (a column access to the open row) - Closed bank:
t_ras + t_cas(activate, then the column access) - Row conflict:
t_pre + t_ras + t_cas(precharge the open row, activate, access) - Bank interleaving: consecutive rows map to different banks, and accesses to different banks overlap; two activates are at least 4 cycles apart (tRRD)
- Refresh: every 7,800 cycles all banks close their rows and are busy for 350 cycles
DDR5 Controller¶
MemoryController.DDR5() is a command-level model of a DDR5 memory
subsystem in the style of gem5's MemCtrl / DRAMInterface. Every request
becomes a sequence of JEDEC commands scheduled against per-bank state, and
every command must clear the timing constraints of JESD79-5B.
Clock domain. The controller runs at the DRAM command clock (data rate
/ 2, so 2400 MHz for DDR5-4800). Requests arrive stamped with the core cycle
and are converted through cpu_clock_mhz; responses are converted back. Set
cpu_clock_mhz to the core you are modelling; the default 2400 MHz gives a
1:1 ratio with DDR5-4800.
Topology. channels × two sub-channels (each with its own command and
32-bit data bus) × ranks_per_channel × bank_groups_per_rank ×
banks_per_group. Physical addresses are split into these coordinates by
the address_mapping interleave; rows are 64 << column_bits bytes.
Timing. A speed bin (4800B, 5600B) carries each JEDEC parameter as
"the larger of N clocks and T ns" and resolves it for the bin's clock,
rounding up as the standard does. Enforced per command: tRCD, tRP, tRAS,
tRC, tRRD_S/L, tCCD_S/L, tCCD_L_WR, tFAW (four-activate window per rank),
tWTR_S/L, tWR, tRTP, tPPD, tRTRS (rank switch on the data bus), the
read-to-write bus turnaround, CL / CWL and the BL16 burst. ACT, RD and WR
occupy the command bus for two clocks, PRE and REF for one. Any field can be
overridden with timing={...}.
Queues and scheduling. Reads and writes have separate bounded queues
(64 entries each); requests wait for admission in arrival order. Writes are
posted: acknowledged when queued and drained later, once the write queue
crosses its high watermark or when there are no reads, for at least
min_writes_per_switch writes. A write to a line already queued merges into
it; a read to a line in the write queue is answered from the queue. Reads
pay the fixed front-end and back-end latencies (10 ns each) on top of the
DRAM access. The scheduler is FR-FCFS by default: an open-row hit that can
issue now wins, otherwise the request that becomes ready soonest; Fcfs
keeps arrival order.
Refresh. AllBank issues REFab every tREFI: the rank stops taking
commands, open rows close with a PRECHARGE-ALL, REFRESH issues after tRP,
and the rank is busy for tRFC1. SameBank issues REFsb every
tREFI / banks-per-group, rotating through the bank sets so only one bank
per bank group is busy (for tRFCsb) while the rest of the rank keeps
serving.
Power-down. With power_down_idle_ns, a rank with no command, no burst
in flight and no queued request for that long enters precharge or active
power-down, and pays tXP after the exit command before its next command.
ECC. SecDed and ChipKill do not change DRAM timing; with
patrol_scrub_ns they add a background scrubber that reads every line in
address order at that rate.
Statistics live under memctrl0.ch<C>.sc<S>: reads, writes,
writes_merged, reads_hit_write_queue, scrub_reads, activates,
precharges, precharge_alls, refreshes, row_hits, row_misses,
row_hit_rate, power_down_entries, power_down_exits, bus_busy_clocks,
clocks, data_bus_utilization, read_admission_stalls,
write_admission_stalls, and the histograms read_latency,
read_queue_depth, write_queue_depth. Per-bank counters sit under
rank<R>.bank<B> and are queryable but omitted from the summary.