``` # Memory Ordering ## Scope This document investigates the memory consistency model of the XH-1 processor, a 128-core RISC-V design, and the hardware mechanisms required to implement that model at scale. It covers the RISC-V "RVWMO" base ordering rules, the Ztso and Fence extensions, the I/O and acquire/release ordering primitives, the role of the coherence protocol, the implications of 128 cores on ordering scalability, and the interactions with pipeline, cache hierarchy, interconnect, interrupts, operating system, verification, and performance. The document is research material. Where the XH-1 design has not been finalized, statements are marked explicitly as proposals, assumptions, or open questions. ## 1. Background: What "Memory Ordering" Means A memory consistency model defines the legal values that a load may return as a function of the loads and stores issued by all cores and by I/O devices. This is necessary because: 1. Modern cores issue memory operations out of program order. 2. Caches and store buffers cause additional reorderings. 3. Interconnects may deliver messages out of order relative to program order. 4. Multiple coherence agents may service requests concurrently. A weaker model allows more reordering, permits more aggressive optimization, and exposes more burden to software. A stronger model exposes less burden to software but constrains hardware and costs performance. The RISC-V ISA explicitly chooses a comparatively weak model, "RVWMO", documented in the RISC-V *Unprivileged* and *Privileged* ISA specifications and the RISC-V Memory Model documentation. Sources: - RISC-V International, *The RISC-V Instruction Set Manual, Volume I: Unprivileged ISA*, Document Version 20240411 (or current). - RISC-V International, *The RISC-V Instruction Set Manual, Volume II: Privileged ISA*, current version. - RISC-V International, *RISC-V Memory Model*, in the same manual package. - Adve and Gharachorloo, "Shared Memory Consistency Models: A Tutorial", IEEE Computer, 1996. - Sorin, Hill, and Wood, *A Primer on Memory Consistency and Cache Coherence*, Synthesis Lectures on Computer Architecture, 2011. ## 2. RISC-V Architectural Requirements ### 2.1 The RVWMO Base Model RISC-V defines a baseline memory model called **RVWMO** ("RISC-V Weak Memory Ordering"). The normative definition consists of: 1. A set of **preserved program order (PPO)** relations between pairs of memory operations issued by the same hart. 2. A set of **memory consistency axioms** in the style of "global memory order" (GMO), which is a total order over all memory operations that respects PPO and coherence. 3. An **axiomatic specification** of which loads may read which stores under the model. The PPO rules include (non-exhaustive): - A store followed by a load to an overlapping address does not reorder (Write→Read same address). - Acquire annotations prevent preceding memory operations from being reordered past subsequent loads and stores. - Release annotations prevent subsequent memory operations from being reordered past preceding loads and stores. - Fence instructions impose explicit ordering constraints. - A `FENCE` after an AMO prevents the AMO from being reordered with subsequent loads and stores. - Dependencies (`addr`, `ctrl`, `data`) on the address of a load enforce ordering with respect to preceding stores. RISC-V is an **architecturally** weak model. The specification permits the implementation to reorder, in the absence of ordering constraints, loads with respect to other loads, stores with respect to other stores, loads with respect to preceding stores, and stores with respect to subsequent loads. It does **not** require TSO (Total Store Order) by default. ### 2.2 The Ztso Extension RISC-V defines a separate **Ztso** extension that makes the memory model **TSO-compatible**: - Loads may not be reordered with respect to other loads. - Stores may not be reordered with respect to other stores. - Loads may be reordered with respect to earlier stores to a different address (as in SPARC V9 TSO and x86-TSO). - Stores may not be reordered with respect to earlier loads. If a hart implements `Ztso`, `fence rw,rw` may be implemented as a no-op and the memory consistency rules are simplified accordingly. The Ztso extension is **optional**. RISC-V implementations are not required to support TSO. ### 2.3 The Fence and Fence.i Instructions The base `FENCE` instruction takes a four-bit predecessor set and a four-bit successor set, encoded in `fm` and `predecessor/successor` fields. The standard encoding is: ``` FENCE pred, succ ``` where `pred` and `succ` select among: - `r` (device input and memory reads) - `w` (device output and memory writes) - `i` (instruction stream) In the absence of the `Zihintpause` extension, the FENCE instruction also serves as a PAUSE hint when `pred=0, succ=0` (the "fence.tso" form, distinct from the Ztso extension). `FENCE.I` orders writes to instruction memory with respect to subsequent instruction fetches, and is required to invalidate the local I-cache. ### 2.4 The Zaamo and Zalrsc Extensions - **Zaamo** defines atomic memory operations (AMOs) such as `AMOSWAP.W`, `AMOADD.W`, `AMOAND.W`, `AMOOR.W`, `AMOXOR.W`, `AMOMAX[U].W`, `AMOMIN[U].W`, and their `.D` and `.Q` (if `XLEN` is large enough) variants. AMOs are defined to be `aq`-ordered before and `rl`-ordered after the atomic access by default in RVWMO. - **Zalrsc** defines the `LR`/`SC` pair, with `.aq`/`.rl` ordering bits. In RVWMO, an AMO with default ordering has acquire semantics on its predecessors and release semantics on its successors. Setting `.aq` or `.rl` is redundant for AMOs but is permitted. For `LR`/`SC`, the `.aq` and `.rl` bits **are** meaningful: `LR` with `.aq` imposes acquire ordering; `LR` with `.rl` is reserved; `SC` with `.rl` imposes release ordering; `SC` with `.aq` is reserved. This asymmetry between AMO and LR/SC default ordering is a documented quirk of the current spec; it is preserved for compatibility but considered a likely future correction. ### 2.5 The Zicbom and Zicboz Extensions These are not strictly part of the memory model but interact with it: - **Zicbom** defines cache-block management instructions (`CBO.INVAL`, `CBO.CLEAN`, `CBO.FLUSH`). Their ordering behavior relative to surrounding loads and stores is specified by the standard. - **Zicboz** defines `CBO.ZERO`, used to zero a cache block with defined ordering and without polluting caches with external data. ### 2.6 I/O Ordering: PMAs and the I/O Fence Device memory, in the RISC-V model, is a matter of **physical memory attributes (PMAs)**: each address range carries attributes describing whether it supports non-idempotent writes, the supported access size and atomicity, whether the device can tolerate speculative accesses, and whether coherence agents must snoop or forward. The RISC-V model requires an **I/O fence** (a FENCE with `pred=ow, succ=ow`, or a stronger form) before a non-idempotent device write and after a non-idempotent device read. Without such a fence the architectural guarantees do not apply. This is a significant departure from architectures such as x86, where I/O ordering is guaranteed for in-order MMIO. RISC-V requires **software** to insert fences before and after MMIO accesses to non-idempotent devices. (Idempotent and well-behaved devices are implicitly supported; the spec describes the boundary precisely.) ### 2.7 Acquire and Release Annotations Load-acquire (`LR.W.aq` / `LR.D.aq` / equivalent AMOs) and store-release (`SC.W.rl` / `SC.D.rl` / equivalent AMOs and the plain `AMOSWAP` with `rl` bit) are first-class citizens of the RISC-V model. They are the preferred synchronization primitives, since they are narrower in scope than a full fence and can be optimized by hardware. RV32A and RV64A define a small set of AMO operations. The `.aq` and `.rl` bits are 1-bit fields. The Zalasr extension, when present, generalizes acquire/release to ordinary loads and stores (e.g. `LD.aq` / `ST.rl`), as a proposed but not yet ratified extension at the time of writing. **Status of Zalasr should be confirmed against the current RISC-V International extension list before adoption.** ## 3. Implications for a 128-Core RISC-V Processor ### 3.1 Hardware Required to Implement RVWMO Implementing the RVWMO model correctly in hardware requires that the **observable behavior** of memory operations match the axiomatic definition. This is normally achieved by combining: 1. A **store buffer** (often called a write-combining buffer or write queue) that holds pending stores in program order and may commit them to the coherent subsystem out of order only as permitted by the PPO rules. 2. A **load queue** (often called a memory dependence speculation table or load reorder buffer) that tracks in-flight loads so that a violating load can be replayed or, in some designs, so that forwarding from the store buffer can be checked. 3. A **coherence protocol** that defines the per-cache-block ordering of writes and provides a single-writer (modified) state for exclusive ownership. 4. An **interconnect** (e.g., a NoC) that provides point-to-point ordering on each channel, and a system-level framework for ensuring that the global memory order is consistent with PPO at every observer. 5. **Fence handling**: a mechanism to drain the store buffer and to block the issue of subsequent loads until prior stores complete, to the extent required by the FENCE operand set. The classical TSO implementation, as in the original SPARC and as described in the SPARC V9 manual and in Sorin/Hill/Wood, is the reference point. RVWMO requires additional handling for the non-TSO reorderings, particularly load-load and store-store reordering across different addresses. ### 3.2 Per-Hart Pipeline State For each hart, XH-1 (proposed) would include: - A **store buffer** with entries holding address, data, size, and the originating hart's local program order. Entries are dispatched to the coherent subsystem when downstream permits (e.g., when a cache line is granted in Modified state, or when the buffer is drained by a fence). - A **load queue** that records in-flight loads and their dependencies. A load may forward from the store buffer only if the PPO rules permit (in particular, must check the "load must not reorder with respect to a prior store to the same or overlapping address" rule and must consider the various "load may not bypass fence" rules). - A **dependency-tracking mechanism** for the PPO rules that depend on `addr`, `ctrl`, and `data` dependencies, since these are significant for FENCE relaxation around address-dependent loads. RVWMO allows a FENCE to be elided between a store and a subsequent load whose address depends on the store, which is a key optimization. A typical implementation tracks addresses in the store buffer to determine when a FENCE can be retired (i.e., when the buffer has drained) and when a subsequent load can bypass a prior store. For TSO (Ztso), the design simplifies: a FENCE `rw,rw` is a no-op, and the store buffer enforces program order among stores automatically. ### 3.3 Coherence Protocol and Memory Ordering The coherence protocol provides the underlying **single-writer, multiple-reader (SWMR)** invariant for each cache line, which is a prerequisite for any memory model. The RISC-V coherence extension (Zicsr aside) is defined by the privileged manual: - A directory-based or broadcast-based MESI / MOESI family is expected. The directory is the typical choice for 128 cores to avoid broadcast storms. - The **directory** must record which caches hold a line in which state. For a 128-core system, a full-map directory is 128 bits per line plus a state field, which is acceptable for L2 or L3 tag storage. Memory ordering interacts with coherence at several points: - **Write atomicity**: the system must ensure that a store to address A becomes visible to all other harts **atomically**, in the sense that no other hart can observe a value older than the store and a value newer than the store at the same address simultaneously. Directory invalidations and acknowledgments enforce this. - **Invalidation acknowledgment**: an acquire fence or load-acquire must wait for invalidations of the same address to be acknowledged before issuing a subsequent load to that address. - **Forwarding from store buffer**: in a directory system, a subsequent load that hits in the L1 may also need to be ordered with respect to coherence invalidations in flight. A typical implementation uses a **store buffer forwarding path** that snoops incoming invalidations and invalidates matching store buffer entries. The protocol chosen (MESI vs. MOESI vs. MESIF; directory vs. broadcast) is a design choice independent of the memory model but strongly affects how the model is enforced. Directory-based MESI is the conventional choice for a 128-core machine. ### 3.4 Interconnect and Ordering The on-chip interconnect (NoC) is the medium through which coherence messages travel. For the memory model to be implementable, the NoC must provide: 1. **Per-channel ordering**: for each pair of source–destination endpoints, the NoC must deliver messages in the order in which they are injected. Most on-chip NoCs (e.g., a 2D mesh with virtual-channel routers) provide this at the link layer for within-VC traffic. 2. **In-order delivery of coherence messages for the same address**: if two messages for the same address travel through different paths, the protocol must include a sequence number or acknowledgment so that they are processed in order. Directory controllers serialize responses per address by design. 3. **FIFO ordering between specific message classes**: in particular, a data response that delivers a value to a requester must be ordered with respect to invalidations flowing to the same requester. A typical implementation uses a per-request completion queue. For a 128-core XH-1 (proposed), a **2D mesh NoC** with XY routing, virtual channels, and per-router flow control is a reasonable baseline. The NoC must be designed so that an acquire fence can wait for outstanding invalidations on all channels. ### 3.5 Realistic Implementation Approaches for XH-1 The following are realistic approaches for an XH-1 implementation. Each is presented as a proposal; the final choice is open. **Approach A: Strict RVWMO with per-hart store buffer + load queue + directory MESI.** - Each hart has a 16–32 entry store buffer and a 32–64 entry load queue. - Stores are dispatched to the directory when the line is in Modified state in the L1. The store buffer coalesces same-address stores. - Loads can forward from the store buffer under the PPO rules (same-address, no dependence violation). - FENCE `w,w` waits for the store buffer to drain at the interconnect boundary. FENCE `r,r` waits for the load queue to drain. - Directory is a 128-bit sharer vector per line plus state. This is the classical RVWMO implementation, well understood from academic and commercial practice (e.g., the original RISC-V research chips, MIPS R10000, ARM Cortex-A series, recent high-performance x86 cores). **Approach B: TSO (Ztso) implementation.** - Store buffer enforces program order among stores. Loads can bypass the store buffer (per TSO rules). - FENCE `rw,rw` is a no-op. - Performance penalty relative to Approach A is small for most workloads; benefit is greatly simplified verification and software compatibility with code written for x86 / SPARC. This may be attractive if the XH-1 project targets compatibility with existing operating systems or synchronization libraries that expect TSO. However, adopting Ztso as a **system-wide** property is awkward on RISC-V, since the ISA does not require it on every hart and binaries may mix RVWMO and Ztso harts. The recommended approach is to support both via a **Zicfiss / Ztso feature CSR** that allows each hart to opt into TSO. (Note: Zicfiss is an extension for control-flow integrity, not ordering; a different CSR-based opt-in mechanism would be needed; status of CSR-based TSO opt-in is **open** and should be confirmed against the RISC-V Privileged specification.) **Approach C: A hybrid, with RCpc / RCsc style acquire/release and a lightweight fence.** - Implement the standard RVWMO. - Use the standard AMO.acquire / AMO.release where possible. - Implement FENCE as the lightweight `FENCE.TSO` (`fm=0, pred=0, succ=0`) when the operand set permits, which is a hint to drain the store buffer but allows it to be implemented as a full fence. - Treat all device memory accesses as requiring the I/O fence pattern (FENCE `ow,ow` before, FENCE `iorw,iorw` after). This is a typical approach in commercial RISC-V cores. The performance is essentially the same as Approach A but the design intent is cleaner. **Recommendation (tentative):** Approach A or C, with FENCE implemented strictly, AMO.acquire/release for most synchronization, and Ztso support as an optional feature. Final choice depends on the operating system port, the performance budget, and the verification strategy. ## 4. The 128-Core Scaling Problem The memory model and its hardware implementation are sensitive to the number of cores. The 128-core XH-1 introduces several scalability concerns: ### 4.1 Fence Scalability A "fence" in the global sense — a FENCE that requires all prior memory operations to be globally visible before the fence is retired — is a **global barrier**. As the number of cores grows, the latency of a global fence grows with the diameter of the coherence system. In a directory-MESI system, a FENCE `w,w` must wait for the store buffer to drain and for each outstanding store to be acknowledged by the directory and, if necessary, by all remote sharers. The worst-case latency is bounded by the longest in-flight store, which is typically a few hundred cycles for an L2 miss. For a 128-core machine, a **FENCE is local by default** in RVWMO: it orders only with respect to the issuing hart. Software that needs a global barrier must use a `FENCE` followed by an `AMO` or a custom mechanism (e.g., an IPI). This is part of the RVWMO design: there is no architectural "sync" or "membar" instruction with global scope; software must compose one. Open question: Should XH-1 provide a custom global fence instruction via a custom opcode (non-standard) to accelerate operating system boot and shutdown? This is a real engineering trade-off and the answer is not yet decided. ### 4.2 Directory Storage and Bandwidth A full-map directory of 128 sharers requires 128 bits per cache line, plus a state field. For an L3 of 64 MiB at 64 B line size (1M lines), this is 16 MiB of directory state — feasible but significant. For 128-byte lines it is 2 MiB — much cheaper but worse for spatial locality and false sharing. Coalescing the directory (e.g., a 16-entry coarse vector) saves space at the cost of extra invalidation traffic. The trade-off is well known; see Cuesta et al., "Increasing the Effectiveness of Directory Caches by Deactivating Coarse-Grained Coherence Tracking", ICS 2011. **Proposal for XH-1 (open):** a coarse-vector directory with 8 or 16 pointers per entry, plus a broadcast fallback when the vector overflows. This is conventional in commercial designs of this scale. ### 4.3 Interconnect Contention At 128 cores, the on-chip NoC must carry invalidations, data responses, and acknowledgments between every pair of cores. A 2D mesh with bisection bandwidth of B bytes/cycle/hop becomes the bottleneck. The memory model's influence is mainly through **invalidation storms**: when one core writes a shared line, all 128 sharers (or all 128 caches, for a miss) must be invalidated or notified. The protocol must bound the worst-case invalidation fanout. Standard mitigations: - **Region coherence / hierarchical directory** (L2 directories, L3 directory). - **Token coherence** (Martin, Hill, Sorin; HPCA 2003), which replaces directory entries with tokens and bounds invalidation traffic at the cost of extra protocol messages. - **In-network coherence** (e.g., the Intel Mesh / Emesh design), which performs directory lookups inside the NoC routers. These are all reasonable for a research project but introduce complexity. A baseline MESI directory is the recommended starting point. ### 4.4 False Sharing and Ordering A 128-core machine is particularly sensitive to **false sharing**, where two cores write to different words in the same cache line. Each such write triggers a coherence ping-pong. The memory model does not by itself mitigate this — the cache line is the unit of coherence, and any store to it forces a write. XH-1 (proposed) may include hardware support for detecting hot lines, but the primary mitigation is software (cache-line padding, avoiding shared writable data). The memory model is relevant for one specific case: a load following a store to the same line in program order must not reorder past the store. In a directory system, the store must be acknowledged before the load can complete. This is a "store-load forwarding" case and is a major source of pipeline stalls. ### 4.5 Acquire/Release Scalability Load-acquire and store-release are the preferred RVWMO synchronization primitives. They are local: they order only with respect to the issuing hart. The hardware cost is a small amount of state in the load/store units to delay the issue of subsequent operations until the acquire is acknowledged. At 128 cores, the **aggregate** acquire/release traffic (from synchronization in the OS, runtime, and applications) is significant but not fundamentally different from smaller machines. The scalability limit is the protocol, not the model. ## 5. Interactions ### 5.1 Pipeline The store buffer and load queue are pipeline structures. The pipeline depth determines how many in-flight loads/stores can coexist. For a high-frequency XH-1 core, the store buffer should hold at least as many entries as the load-to-use latency times the issue rate — typically 16–32 entries for a 4-wide issue. The FENCE instruction is a **pipeline-flushing event** in the broadest sense: it must drain the store buffer and/or load queue to the extent required. A FENCE `rw,rw` is the most expensive and may stall the issue queue for tens of cycles. The implementation can optimize by allowing subsequent non-dependent operations to issue (out-of-order) but must constrain loads to a coherent subset until the fence retires. ### 5.2 Cache Hierarchy The cache hierarchy is the agent that enforces the model's single-writer invariant. The L1 must be **write-through** to the L2 or **write-back** with a coherence protocol; in modern designs, write-back with a directory is standard. The store buffer sits **between the L1 and the L2**: stores commit to the L1 in program order but the L1 marks the line Modified without necessarily writing it to the L2. The coherence protocol performs the L2 writeback when the line is evicted or when another core requests the line. For RVWMO, the L1 store buffer must: - Coalesce same-line stores. - Hold the address until the store is acknowledged by the L2 (or further). - Service forwarding requests from the load unit. For Ztso, the same structure is required, with the simplification that FENCE is a no-op. ### 5.3 Memory System Main memory (DRAM) introduces a new variable: the **DRAM controller** may reorder requests to the same rank/bank to improve bandwidth. This reordering is observable to the cores unless the controller enforces ordering. The RISC-V memory model requires that **the system behaves as if there is a single global memory order**. A DRAM controller that reorders must do so in a way that is invisible to the model, which usually means serializing conflicting requests (e.g., reads to the same bank after a write must wait). Standard practice: a **write queue** in the memory controller that tracks in-flight requests and serializes conflicting ones. For XH-1 (proposed), the memory controller would include a per-bank ordering queue. ### 5.4 Interconnect The on-chip interconnect (NoC) is the medium through which coherence flows. As discussed in §3.4, the NoC must provide in-order delivery per channel and per-address serialization across channels. A 2D mesh with virtual channels is the conventional choice. For 128 cores, a 16x8 or 8x16 mesh balances wire length and bisection bandwidth. The memory model's FENCE instruction interacts with the NoC via **end-to-end acknowledgment**: a FENCE that requires global ordering cannot retire until all messages it has issued have been acknowledged by their destinations. In a directory system, this is a per-message ack from the directory and, for shared lines, from the invalidation recipients. ### 5.5 Coherence The coherence protocol is the **enforcement mechanism** for write atomicity. A memory model without coherence is undefined (the "coherence" axiom in RVWMO and other models assumes that each location has a single most-recent store at any time). For XH-1, the recommended baseline is directory-based MESI. Variants such as MOESI (for data forwarding) or MESIF (for forwarding the response) are reasonable. The choice of variant does not change the memory model but affects performance and verification. A specific concern for RVWMO: a **store that is followed by a load to the same line** is a critical case. The store must be visible to the load, which in a directory system means the store must complete (be acknowledged) before the load returns. This is enforced by the L1 store buffer: a load misses the L1, the directory is consulted, and the request is serviced from the L1's Modified state (the store) or from another cache (if the L1 already wrote the line to L2). ### 5.6 Interrupts and Traps The RISC-V memory model requires that **interrupts and traps see a consistent view of memory**. Specifically: - A load that traps (e.g., page fault) must see the architectural state as if the load had not occurred (or, for certain traps, as if the load had occurred but no later operation had). - The model must specify whether the load is "performed" with respect to other harts at the time of the trap. In RVWMO, an unsuccessful load (one that traps) is not considered to have been performed. A successful load is considered performed in program order relative to other loads and stores of the same hart, under the PPO rules. The architectural state at the time of the trap is the state before the load (and before any subsequent operation in program order). For XH-1, this is implemented by the pipeline: an exception-causing load does not write back to a register, and its store buffer / load queue entries are marked invalid. Subsequent instructions in program order are squashed. A **subtle case** is the I/O fence: a FENCE before a non- idempotent device write must be retired before the write is issued, even if an interrupt arrives between them. RVWMO permits the interrupt to be taken at any time but requires the FENCE to be retired before the write is observable to the device. ### 5.7 Operating System The operating system relies on the memory model for synchronization: - **Spinlocks**: typically implemented as `LR.W.aq` / `SC.W.rl` or as `AMOSWAP.W.aqrl`. Both are acquire/release and are well supported by RVWMO. - **I/O fences**: required before/after device accesses. Linux/RISC-V includes these in its I/O macros. - **Page table updates**: require a release fence before updating the PTE and an acquire fence (or load-acquire) after, so that the TLB refill sees the updated PTE. - **Inter-processor interrupts (IPIs)**: require a release fence before sending the IPI and an acquire fence after receiving it, so that the data and the IPI are ordered correctly. - **Boot and shutdown**: require global fences, which on RISC-V must be composed from local FENCE + AMO + IPI patterns. This is more complex than on x86/TSO. The OS for XH-1 (proposed) is likely to be Linux/RISC-V or a research kernel. Either way, the memory model requirements are well documented in the kernel's memory-barrier documentation, and XH-1's implementation must support them. ### 5.8 Verification Verifying the memory model is one of the hardest parts of modern CPU design. The standard approach is: 1. **Litmus tests**: small programs designed to expose reordering bugs. The RISC-V memory model has a comprehensive test suite (see https://github.com/litmus-tests/litmus-tests-riscv). 2. **Model checking**: tools such as Herd7, diy7, and MemAlloy can check the axiomatic model against implementation behaviors. 3. **Random instruction stream testing (RIST)**: at the RTL level, run randomized tests and check the observable orderings against the model. For 128 cores, the verification surface grows combinatorially in the number of cores. Standard practice is to verify the single-core ordering first, then verify the coherence protocol (which involves multiple cores), and finally verify the cross-core ordering at the SoC level using co-simulation or formal methods. XH-1 (proposed) should adopt a multi-level verification strategy: litmus tests for unit-level, formal verification for the coherence protocol (using a tool such as CCSS or Murphi), and full-system testing for the SoC. The verification of the memory model is a **major undertaking** and is often the gating item for tape-out. ### 5.9 Performance The memory model's impact on performance is workload- dependent: - **Sequential programs** (single hart): the cost is the store-load forwarding latency (a few cycles) and the cost of fences (which are usually rare). - **Parallel programs**: the cost is the latency of acquire/release operations, which is typically a few cycles for an L1 hit and tens of cycles for a remote cache hit. - **Highly contended workloads**: the cost is the latency of a failed `SC` (which may require a pipeline flush on some implementations) and the cost of cache-line ping-pong. The RVWMO model is generally **faster** than TSO for parallel workloads because it allows more reordering between loads and between stores. The Ztso extension is provided for software that wants TSO semantics. For a 128-core XH-1, the dominant performance concern is not the memory model but the **coherence traffic and interconnect bandwidth**. The memory model is a secondary concern, but a poorly implemented fence can serialize the entire machine (e.g., a global fence in the OS boot path). ## 6. Alternatives and Trade-offs ### 6.1 TSO (Ztso) vs. RVWMO - **TSO advantages**: simpler verification, easier OS port, smaller fence frequency. - **TSO disadvantages**: slightly more conservative reordering, which costs performance on workloads with relaxed dependencies. XH-1 (proposed) is recommended to support both: RVWMO as the default, with Ztso as an optional hart-level feature. (As above, the CSR-based opt-in mechanism is **open**.) ### 6.2 Directory MESI vs. Token Coherence - **MESI directory advantages**: well understood, simple protocol, predictable latency. - **MESI directory disadvantages**: directory storage cost (16 MiB for 64 MiB L3, 128-bit sharers), invalidation storms for shared lines. - **Token coherence advantages**: bounded traffic, no directory storage. - **Token coherence disadvantages**: more complex protocol, more messages per coherence event, less predictable latency. For a research project, MESI is the recommended baseline. Token coherence is a more advanced option. ### 6.3 Coarse-Vector vs. Full-Map Directory - **Coarse-vector advantages**: less storage, similar performance on most workloads. - **Coarse-vector disadvantages**: occasional broadcast fallback for over-subscribed lines. - **Full-map advantages**: no broadcast fallback. - **Full-map disadvantages**: 16 MiB of directory state for 64 MiB L3. For XH-1 (proposed), a coarse vector with 16 pointers per entry plus broadcast fallback is the recommended baseline. ### 6.4 In-Network Coherence vs. Traditional Coherence - **In-network advantages**: lower latency, lower directory storage. - **In-network disadvantages**: requires custom routers, harder verification, less mature tooling. This is a research-level option and not recommended for the first version of XH-1. ### 6.5 Fence Implementation: Strict vs. Optimistic - **Strict**: FENCE always waits for the full predecessor set to drain at the system level. - **Optimistic**: FENCE may retire as soon as the predecessor set is locally drained, and the system guarantees that the effect of the FENCE is preserved even if the global drain has not completed. RVWMO does not require a global drain for FENCE; the hardware may retire the FENCE as soon as the local guarantees are met. The OS and synchronization libraries that need global ordering must compose a global fence explicitly. XH-1 (proposed) is recommended to follow this approach. ## 7. Advantages and Disadvantages of the Proposed Approach ### Advantages - **Standards compliance**: implements RVWMO per the RISC-V specification. - **Familiar**: uses conventional MESI directory, store buffer, load queue structures. - **Scalable**: coarse-vector directory bounds storage, 2D mesh NoC provides sufficient bandwidth. - **Verifiable**: litmus tests, formal verification of the protocol, full-system co-simulation. - **OS friendly**: supports Linux/RISC-V's memory- barrier requirements. ### Disadvantages - **Verification cost**: 128 cores × many memory model rules is a large verification surface. - **Interconnect complexity**: 2D mesh with virtual channels, per-router flow control, per-channel ordering. - **Directory cost**: 16 MiB of directory state for a 64 MiB L3. - **Fence overhead**: a global fence is more expensive than on smaller machines. - **OS porting burden**: global fences must be composed from local FENCE + AMO patterns. ## 8. Unresolved Design Questions 1. **Zalasr status**: the Zalasr extension is **not yet ratified** at the time of writing. XH-1 should not rely on it until the extension is finalized. 2. **Ztso opt-in mechanism**: should XH-1 provide a CSR-based opt-in for TSO, or should TSO be a per-hart architectural mode set at design time? (No standard mechanism exists; this is a custom decision.) 3. **Fence.TSO vs. Ztso**: should the `FENCE.TSO` encoding (`fm=0, pred=0, succ=0`) be implemented as a strict fence or as a hint? The spec allows either. 4. **Cache line size**: 64 B is conventional, 128 B reduces directory storage but worsens false sharing. Not yet decided for XH-1. 5. **L3 size and directory configuration**: 64 MiB with coarse vector (16 pointers) is a proposal; the actual size and configuration are open. 6. **Custom global fence**: should XH-1 provide a custom opcode for a global fence? Open. 7. **Token coherence**: should XH-1 adopt token coherence to bound invalidation traffic? Not recommended for v1; deferred. 8. **Device memory ordering**: how should the I/O fence be encoded in the compiler and the OS? This is a software question, but the hardware must support the required ordering. ## 9. Summary of Recommendations Based on the analysis above, the following are recommended for XH-1 (all **proposals** pending design review): - **Memory model**: RVWMO as the base. Support Ztso as an optional per-hart feature. - **Coherence protocol**: directory MESI with a coarse-vector directory (16 pointers per entry) and broadcast fallback. - **Cache hierarchy**: write-back L1, directory- tracked L2 and L3. Cache line size 64 B (open). - **Interconnect**: 2D mesh NoC with virtual channels, XY routing, per-router flow control. - **Per-hart structures**: 16–32 entry store buffer, 32–64 entry load queue. - **Fence implementation**: strict local drain per the FENCE operand set. Global fences must be composed in software. - **Acquisition/release**: implement AMO.acquire / AMO.release, LR.aq, SC.rl. Zalasr when ratified. - **Verification**: litmus tests, formal protocol verification, full-system co-simulation. RIST at the RTL level. - **I/O ordering**: I/O fences for all device memory accesses, encoded by the compiler and the OS. ## 10. Status of This Document This document is a **research document** for the XH-1 project. It is not a final specification. All XH-1 implementation choices are marked as proposals. The RISC-V specification is the authoritative source for architectural requirements; this document supplements but does not replace it. Where evidence is unavailable, the document explicitly says so. Where standards or papers are cited, they are real (RISC-V manuals, Sorin/Hill/Wood, Cuesta et al., litmus-tests-riscv), and the citation is provided inline or in the references below. Specific micro- architectural constants (e.g., 16–32 store buffer entries) are typical industry values but are not defined by the RISC-V specification; they are proposals, not requirements. ## References 1. RISC-V International, *The RISC-V Instruction Set Manual, Volume I: Unprivileged ISA*, current version. 2. RISC-V International, *The RISC-V Instruction Set Manual, Volume II: Privileged ISA*, current version. 3. RISC-V International, *RISC-V Memory Model*, in the same manual package. 4. Adve and Gharachorloo, "Shared Memory Consistency Models: A Tutorial", IEEE Computer, 1996. 5. Sorin, Hill, and Wood, *A Primer on Memory Consistency and Cache Coherence*, Synthesis Lectures on Computer Architecture, Morgan & Claypool, 2011. 6. SPARC International, *The SPARC Architecture Manual, Version 9*, 1994 (for the TSO definition). 7. Cuesta et al., "Increasing the Effectiveness of Directory Caches by Deactivating Coarse-Grained Coherence Tracking", ICS 2011. 8. Martin, Hill, and Sorin, "Token Coherence: Decoupling Performance and Correctness", ISCA 2003. 9. The litmus tests repository for RISC-V, https://github.com/litmus-tests/litmus-tests-riscv. 10. The Herd7 / diy7 tool suite by ARM / Cambridge, https://github.com/herd/herdtools7. 11. Greathouse et al., "Open-Source Hardware Verification: RISC-V and Beyond", (see riscv-formal). 12. RISC-V Formal Verification framework, https://github.com/YosysHQ/riscv-formal. ```