40 KiB
Memory Ordering
Status: Research proposal for the XH-1 128-core RISC-V processor. Nothing in this document is a finalized architectural commitment unless explicitly stated. Items marked [Proposed] are engineering recommendations under evaluation; items marked [Open] are unresolved design questions.
1. Scope and Definitions
Memory ordering defines the set of rules that determine when memory operations issued by one or more hardware threads become visible to other hardware threads, and in what order such visibility is required to occur.
In a single-core, single-threaded system, "ordering" is trivial: program order matches execution order matches visibility order. As soon as a processor exposes:
- out-of-order execution,
- multiple hardware threads (SMT),
- multiple cores with private caches,
- coherent or non-coherent interconnects, or
- devices that perform DMA,
…memory ordering becomes a first-class architectural and microarchitectural concern. For a 128-core design like XH-1, it is one of the dominant correctness and performance constraints.
The RISC-V ISA delegates most ordering policy to a Memory Consistency Model (MCM) specification rather than baking it into the base ISA. The current published MCM is documented in the RISC-V specification volume “RISC-V Instruction Set Manual, Volume I: Unprivileged ISA”, Appendix A — "RVWMO: The RISC-V Weak Memory Ordering Model". RVWMO is the canonical model; "Zam" (the "A" extension candidate) is a candidate formalization of the same weak-ordering contract with additional operational clarifications.
This document covers:
- The RVWMO contract XH-1 must honor.
- Microarchitectural mechanisms needed to enforce that contract.
- Scalability concerns specific to 128 cores.
- Interactions with coherence, the interconnect, the pipeline, the OS, and verification.
2. The RVWMO Contract XH-1 Must Honor
RVWMO is defined by:
- Preserved Program Order (PPO): a set of (microarchitectural)
orderings between pairs of memory operations from the same hart that
the implementation must never violate. PPO rules include the obvious
ones (a load following a store to the same address, register
dependencies, fences such as
FENCE,FENCE.I, etc.) and the subtler "axiomatic" PPO rules covering acquire/release semantics on ordinary AMOs. - Memory Model Axioms: a set of forbidden load–load, load–store, and store–store reorderings when no PPO rule requires otherwise (e.g., "no load may be reordered with a store to a different location" except under specific conditions).
- Syntactic Dependencies: register dependencies via the destination register of a load are sufficient to constrain certain reorderings, mirroring the original SPARC RMO/TSO literature.
The base RVWMO does not mandate TSO; an implementation may run hart-local memory accesses in any order consistent with the PPO rules, with a few mandated orderings (e.g., loads after stores to the same address are preserved).
Key architectural levers available to software:
| Construct | Effect (informal) |
|---|---|
FENCE (various pred/succ/fm) |
Orders memory operations as specified. |
FENCE.I |
Orders stores to instruction memory w.r.t. subsequent instruction fetches on the same hart. |
FENCE.VMA |
Orders SFENCE.VMA-style TLB management. |
SFENCE.VMA |
Orders stores w.r.t. subsequent TLB refills on the same hart. |
| Acquire/release on AMOs | Provides PPO ordering for paired RMW operations. |
LR/SC (LR/SC pair) |
Ordering anchored to the reservation. |
Zalrsc / Zaamo |
Standard atomic extensions. |
Zifencei |
Required for FENCE.I. |
Zihintpause |
Hint only; no ordering effect. |
RV32/64 fence on devices |
Device ordering via PTE bits (see §10.3). |
XH-1 requirement: the ISA-visible contract of every memory and
synchronization instruction listed in the Unprivileged ISA and any
extension XH-1 implements (A, Zifencei, Zihintpause, Zalrsc,
Zaamo, Ztso if adopted) must be honored by the combined
core/LSU/coherence/interconnect subsystem.
Open: Whether XH-1 advertises
Ztso(an optional RVWMO relaxation that grants TSO at every hart) or only RVWMO.Ztsosimplifies software reasoning but slightly constrains hardware ordering flexibility. Recommendation deferred until microarchitecture is fixed.
3. Required Fences and Their Microarchitectural Meaning
The FENCE instruction is parameterized by:
pred: which predecessor operations are ordered (R, W).succ: which successor operations are ordered (R, W).fmfield: indicates "fence-mode"; value0is the standard fence, value>=1is reserved for implementations and is treated as a no-op by standard tools. Custom XH-1 fence-mode encodings are [Open].
XH-1 must ensure that a FENCE with predecessor set P and successor
set S:
- All memory operations of type P issued by this hart prior to the
FENCEare ordered before all memory operations of type S issued by this hart after theFENCE, in the global memory order, to the extent that the coherence protocol is willing to serialize them.
Microarchitecturally, the FENCE typically drains in-flight operations
of the predecessor classes through the load/store unit (LSU), the
coherence point, and the write-completion path. On XH-1 with private L1
and a shared L2 slice, "the coherence point" is the directory controller
on the L2 slice; see §7.
4. Why 128 Cores Changes the Problem
A small-count many-core can often "get away" with conservative ordering: a single shared bus or ring, broadcast snoops, and a full-fence on every store buffer drain. At 128 cores these approaches do not scale:
-
Bandwidth: A serialized global fence requires all in-flight traffic from the issuing core to drain, and the directory to observe all acknowledgments before subsequent operations are injected. With 128 cores contending for the directory and the interconnect, the latency of a single
FENCEcan grow from ~10–30 cycles (small system) to several hundred cycles. This directly impacts lock-based critical sections, RCU grace periods, andatomicoperations in the OS. -
Queue occupancy: Every core holds per-message-class queues (read requests, write requests, invalidations, data responses, writeback data). A fence must account for occupancy of queues between the core and the L2 directory and between directories across the mesh. Without back-pressure, an XH-1 fence could complete locally but leave a "hole" in the global order.
-
Coherence protocol overhead: Directory-based coherence (assumed for XH-1; see the coherence research document) replaces broadcast snoop traffic with per-line directory lookups, but introduces ack/grant ordering dependencies that the fence must not skip past.
-
Scalable coherence introduces non-atomic transitions. Directory state moves from M→I→S→M through intermediate invalidation acknowledgment states. A store-conditional (
SC) on one core that has been racing with a remoteLRmay observe the line in intermediate states; the memory model must be specified for these. -
Contention of fence-induced drains: When 128 cores all execute
FENCEin a short window (e.g., a barrier in a parallel workload), the interconnect can become congested with drain acknowledgments. The microarchitecture must avoid feedback-loop deadlock between the queues and the fence-acknowledgment path.
Proposed XH-1 stance: Treat memory ordering as a system-level performance constraint, not just a per-core correctness one. Evaluate the cost of every
FENCEsite in the OS, runtime, and synchronization libraries at the system level, not at single-core simulation.
5. Microarchitectural Mechanisms
The classic toolkit, applied to XH-1:
5.1 Load-Store Unit (LSU) Pipeline
A typical in-order XH-1 core will have:
- A store buffer (STB) between the integer pipeline and the L1 data cache.
- A load queue (LDQ) that issues loads past stores (speculatively) and validates the result at commit, rolling back on a misspeculation.
- A memory dependence predictor (optional; see §9.1) to reduce the cost of Ld→St forwarding checks.
The LSU's job relative to ordering is twofold:
- Honor PPO locally — enforce that a load following a store to the same address returns the store's value (or a younger value), and that the AMO-acquire/release rules produce the required PPO edges.
- Provide a clean drain point for
FENCE— a place where the LSU can mark "all pred-class operations up to this point have completed globally."
5.2 Store Buffer Drain
A FENCE with pred including W must drain the store buffer. In a
directory-coherent system the relevant completion is not "store reached
L1" but "store has acquired the directory ownership token (e.g., M
response) and the directory has acknowledged writeback/Invalidate
acknowledgments for any invalidated peers."
For XH-1:
- L1 hit, no coherence traffic: store can complete in ~1–3 cycles after reaching L1.
- L1 miss, L2 hit, no sharers: invalidation may be skipped (silent upgrade to M); store completes when directory grants M.
- L1 miss, L2 hit, sharers exist: directory generates invalidation messages to all sharers; the store completes when all invalidation acknowledgments are received.
- L1 miss, L2 miss: directory forwards to home node; home node replies with data and grants M; the store completes when data and the M grant are both received.
Only the last three cases impose significant drain latency.
Proposed XH-1 stance: The store buffer must be sized to hold stores pending completion, but the fence-drain path must not depend on the store buffer draining in order — older stores in the buffer that have already received their M response should not block the fence on younger stores still in flight, unless the fence successor set requires it.
5.3 Load Issue Rules
Loads are the harder side. RVWMO allows loads to be reordered with prior stores to different addresses. XH-1's LSU will speculatively issue loads past older stores; the load is "validated" at commit by checking that the cache line has not been invalidated and that no older store wrote to the same address.
A load that needs to "see" the result of an older store on the same hart uses store-to-load forwarding through the store buffer; this must respect the size and sign-extension rules of the load.
A load that needs to "see" the result of an older store on a different
hart (which is what FENCE/FENCE.RW enforces) must wait for that
remote store to complete at the directory and update the local
view via the coherence protocol. The directory's response is the
"happens-before" edge.
5.4 Acquire/Release on AMOs
A standard AMO* operation in RVWMO has the ordering semantics of an
acquire (load part) and release (store part) when used in aq/rl
positions. In RV32A/RV64A, the instructions LR and SC carry
acquire/release bits; AMOs (e.g., AMOSWAP.W, AMOADD.D) also carry
these bits.
For XH-1, an aq AMO must order prior loads/stores from the issuing
hart before the AMO, and the AMO before subsequent loads/stores.
Microarchitecturally:
aq: drain the load and store queues (or at least establish PPO edges) before issuing the AMO to the directory.rl: drain the load and store queues after the AMO completes.aq+rl: drain on both sides; this is the strong fence.
Open: Whether XH-1 implements
aq/rlpurely in the LSU (with no global fence message) or with a lightweight fence message to the directory. Purely-local enforcement is faster but requires that the directory serialize the AMO at the line state machine.
5.5 LR/SC
LR loads with reservation; SC stores conditionally if the
reservation is still valid. XH-1's reservation representation is
proposed as a small set of reservation-set entries per core, indexed
by physical address, anchored in the L1. Coherence actions that
invalidate the reserved line (a snoop/Invalidate that the L1 sees)
must invalidate the reservation.
Proposed XH-1 reservation-granularity choice: XH-1 reserves at cache-line granularity (the natural L1 line size) and treats forward-progress as the OS's responsibility. This is the conservative default; some designs use sub-line reservation to improve SC success rates, but the gain on typical lock code is modest.
5.6 FENCE.I and Self-Modifying Code
FENCE.I ensures that stores from the issuing hart that preceded the
FENCE.I are visible to subsequent instruction fetches on the same
hart. It does not order stores from other harts.
For XH-1, FENCE.I typically:
- Drains the store buffer (in the W sense).
- Invalidates the L1 I-cache (or the local portion of a unified L1).
- Optionally issues a
FENCEinternally so that the store is visible at the directory before the I-fetch resumes.
A cross-hart self-modifying-code sequence (hart A writes code, hart B
executes it) requires an explicit FENCE on hart A and a FENCE.I on
hart B (per the RISC-V recommendation), plus coherence to propagate
the new data. This is awkward and XH-1 will document it in the
programmer's manual.
5.7 SFENCE.VMA
SFENCE.VMA orders prior stores to PTEs before subsequent
address-translation results on the same hart. XH-1's TLB refill
engine must observe the store-completion signal; a typical
implementation stalls the pipeline at the SFENCE.VMA until the
store buffer drains.
6. System-Level Ordering Devices
Beyond per-hart mechanisms, XH-1 needs system-level ordering for I/O and devices. The RISC-V privileged specification provides two mechanisms:
6.1 I/O Ordering Bits in PTE
A PTE can have the PBMT field (Page-Based Memory Types), which can be marked as:
- Non-cacheable, idempotent, weakly-ordered (NC, IO).
- Non-cacheable, idempotent, strongly-ordered (NC, I/O-strong).
XH-1 must honor these by:
- Routing accesses to the appropriate device port on the interconnect.
- For weakly-ordered I/O, treating the access as device but allowing reordering with respect to other weakly-ordered I/O of the same type.
- For strongly-ordered I/O, treating the access as a full-fence — the load or store completes in program order with respect to all other accesses from the same hart.
6.2 Memory-Mapped I/O Fences
XH-1 will provide a FENCE variant (the standard FENCE pred=rw, succ=rw, fm=0) that software uses to order I/O accesses. The
microarchitecture must ensure the drain endpoint includes the
device-port queue, not just the L1/L2.
6.3 Interrupt Acknowledgment and Ordering
RISC-V interrupt acknowledgment is itself a memory-mapped I/O
write. The microarchitecture must guarantee that the write to the
Platform-Level Interrupt Controller (PLIC) is ordered with respect
to the I/O operations it is acknowledging. The standard idiom is a
FENCE before the claim/complete operation. XH-1's PLIC integration
must be designed so that the FENCE does not require extra
microarchitectural support beyond the standard fence.
Open: Whether XH-1 implements a "fence fast path" for accesses to specific device memory regions (e.g., the PLIC) that bypasses the directory and uses a dedicated sideband signal. Such optimizations can reduce interrupt latency by tens of cycles but complicate coherence reasoning. Not recommended for the first silicon.
7. Coherence Interaction
XH-1's coherence protocol (described in the coherence research document) is directory-based, MSI or MESI, with a distributed directory across the L2 slices. Memory ordering depends on the coherence protocol's notion of "completion."
For a store to be considered "globally performed":
- The line is held in Modified state by the requesting core.
- All prior sharers have sent Invalidate Acknowledgments to the directory (or the directory has confirmed there were no sharers).
- The directory has logged the new owner.
For a load to be considered "globally performed":
- The line is held in Shared (or Modified if exclusive) by the requesting core.
- The data response from the directory (or forwarded from a peer) is the value the directory currently authorizes.
Important subtlety: A
FENCE W,Wbetween two stores on the same hart must ensure that store 1 has acquired its M grant and the directory has processed the invalidations before store 2 acquires its M grant. In directory terms, store 2 cannot be issued until store 1's M grant has been observed at the directory. XH-1 must implement this as an ordering dependency at the directory's L2 controller, not as a per-line serialization at the directory (which would be too coarse and too slow).
7.1 Coherence-Induced Ordering Pitfalls
Common pitfalls in directory-based designs:
- Directory ack starvation under fence pressure: when many cores fence simultaneously, the directory must serialize invalidation acknowledgments without deadlock.
- Coherence downgrade during fence: a core that downgrades M→S while a fence is in progress must not let younger loads proceed past the fence.
- Silent upgrade (no-invalidate transition) on store: a clean line upgraded to M without invalidation must still be ordered with respect to other cores' loads that observed the line in S. The directory must explicitly synchronize this transition.
7.2 Ordering at the Interconnect
XH-1's interconnect is a mesh (proposed; see interconnect research document). The mesh delivers messages in-order between a given pair of routers, but globally the order is not guaranteed. The directory controller is the serialization point.
For ordering, XH-1's mesh must:
- Not reorder different message classes on the same virtual channel.
- Support a "tail" or "credit" signal so that a fence-acknowledgment message can be generated only when all prior messages on the same path have been delivered.
8. Implementation Approaches Compared
This section lists the realistic implementation strategies for an in-order XH-1 core. The choice of in-order vs. out-of-order pipeline is discussed in §9.1; here we focus on ordering-specific mechanisms.
8.1 Approach A: Conservative Per-Core Fence (Baseline)
- The LSU drains the relevant queue classes fully on every
FENCE. - The store buffer's "drain complete" is defined as "every store has received its directory M grant."
- A
FENCEis implemented as a special instruction that stalls the pipeline until drain completes. - Acquire/release AMOs are implemented as full
FENCEs around the AMO at the LSU. SFENCE.VMAstalls the pipeline until the store buffer drains.
Pros:
- Simple to verify. The fence semantics maps directly to "queue empty."
- Predictable performance for small core counts.
- Easy to reason about coherence interactions.
Cons:
- Poor scalability. At 128 cores, the worst-case
FENCElatency grows because the drain point is global, not local. - Overly conservative for
aq/rlAMOs, which do not require a full global drain on either side — only ordering relative to surrounding ordinary accesses.
8.2 Approach B: Token-Based Completion
- The LSU attaches a monotonically increasing completion token to each outstanding operation.
- A
FENCEwith pred P, succ S becomes: "wait until all operations of type P issued before thisFENCEhave produced a token greater than or equal to the fence's token." - The token is generated when the directory acknowledges the operation; this is exactly the "global completion" notion.
- Acquire/Release is implemented as a fence of the appropriate type around the AMO at the LSU.
Pros:
- Fences do not require the entire store buffer to drain; an empty store buffer is sufficient, but so is a store buffer where all older stores have completed.
- Naturally parallel: many stores can be outstanding, and the fence only waits for the oldest of them.
- Slightly more complex but a well-understood technique (used in academic and commercial cores).
Cons:
- Per-message token bookkeeping is more expensive than a simple "queue empty" check.
- Verification requires modeling tokens at the testbench, which is a non-trivial simulation infrastructure investment.
8.3 Approach C: Directory-Assisted Ordering
- A special "fence-ack" message class is added to the coherence protocol.
- A
FENCEinstruction sends a fence-ack request to the directory; the directory replies only after all prior messages from the issuing core have been serialized. - The directory tracks per-core "last-seen" sequence numbers to arbitrate fence-ack requests against ongoing traffic.
Pros:
- The directory is the natural serialization point; leveraging it avoids per-core reinvention of global ordering.
- Scales well to large core counts because the directory already handles 128-core traffic.
- Allows the per-core LSU to continue issuing younger operations while the fence is in progress, as long as they are ordered correctly.
Cons:
- Adds protocol complexity (a new message class, directory state, ack/grant logic).
- Risk of fence-ack starvation under heavy traffic.
- Requires careful design to avoid deadlock with normal coherence traffic (classic message-class-dependency cycle).
8.4 Approach D: Speculative OoO + MOB
This is the out-of-order approach. The LSU becomes a Memory Ordering Buffer (MOB) that tracks per-operation completion and issues fences by tracking the "oldest in-flight" of each class. This is how high-performance OoO cores implement fences; the trade-off is pipeline depth and complexity.
XH-1 is proposed as in-order for power and predictability; see §9.1. OoO is listed for completeness.
Pros:
- Highest performance on workloads with deep memory-level parallelism.
- Fences can be "free" if no relevant older operation is in flight.
Cons:
- Significantly larger area and power.
- More verification effort, especially for the MOB's speculative reordering.
- Diminishing returns on memory-bound workloads at 128 cores (the bottleneck is the interconnect, not the per-core issue rate).
8.5 Summary of Approaches
| Approach | Perf (single) | Perf (128) | Area | Verif. effort | Risk |
|---|---|---|---|---|---|
| A: Cons. | Low | Lowest | Lowest | Lowest | Low |
| B: Token | Medium | Medium | Low | Medium | Low |
| C: Dir-Assisted | Medium-high | High | Medium | High | Medium |
| D: OoO+MOB | High | High | High | Highest | High |
Proposed XH-1 implementation: Approach C (Directory-Assisted Ordering) as the primary, with Approach B (Token-Based Completion) as the per-core implementation strategy. The directory handles global serialization, the LSU handles local token tracking. This combination is the most natural for a directory-coherent 128-core design.
9. Interactions with the Pipeline and Other Subsystems
9.1 Pipeline Choice (In-Order vs. Out-of-Order)
XH-1 is [Proposed] in-order for the first silicon. The reasons:
- Power envelope. 128 cores at high frequency is a thermal challenge; an in-order core can run cooler.
- Predictable timing. Real-time and safety-critical workloads
benefit from deterministic
FENCEandFENCE.Ilatencies. - Verification cost. An OoO core's MOB is a major verification burden; a 128-core OoO design is a multi-year project.
An in-order pipeline interacts with ordering as follows:
- The LSU is non-speculative on the store side (stores are committed when the integer pipeline commits them) and lightly speculative on the load side (loads may issue before older stores, with a structural dependency check at commit).
FENCEstalls the pipeline until drain; the cost is a few tens to a few hundred cycles depending on directory state.- Acquire/release AMOs generate local fence events in the LSU.
An OoO pipeline would:
- Issue stores from the STB to L1, with the MOB tracking completion.
- Issue loads speculatively past stores with prediction; the MOB validates at commit.
FENCEbecomes a "wait for oldest in-flight of pred class to complete" check; usually very fast.
Open: Whether a future XH-2 could be OoO. The research project is not yet at that question.
9.2 Cache Hierarchy
Memory ordering is shaped by the cache hierarchy:
- L1 D-cache: the load-side ordering reference. A load that hits in L1 sees a value that is at least as recent as the cache's coherence state indicates. A load that misses in L1 must go to L2, where the directory arbitrates the order.
- L2 slice (per tile): home of the directory entry; serializes ordering for the addresses it owns.
- L3 (shared, distributed): a victim cache and last-level coherence point. XH-1's directory is the L2 directory; L3 is treated as a backing store. Invalidation ordering for XH-1 therefore ends at L2.
- L1 I-cache: the
FENCE.Itarget. The I-cache must be flushed or invalidated to ensure stores are visible.
The cache hierarchy directly affects FENCE latency because the
"completion" point is the L2 directory. A FENCE W,W in a system
with L1-only ordering would be much cheaper (a few cycles) but
incorrect; the L2 directory must be in the loop.
9.3 Memory System
The memory system (DRAM controllers, on-chip memory, scratchpads) must respect the same ordering. For XH-1, accesses to on-chip scratchpad memory bypass the coherence protocol entirely; they are ordered by the scratchpad controller.
Open: Whether XH-1 has a per-tile scratchpad or a shared on-chip SRAM region. If shared, ordering between scratchpad and cache-coherent regions is a research question.
9.4 Interconnect
The mesh interconnect has ordering properties (per-link in-order, globally non-ordered) that determine what the directory and endpoints must do to maintain memory ordering. The "virtual channel" abstraction in the mesh is the unit of ordering; XH-1 [Proposed] uses:
- VC0: Coherence requests (in-order).
- VC1: Coherence responses (in-order).
- VC2: Data messages (in-order).
- VC3: Fence-ack and other ordering-sensitive messages (in-order).
This separation avoids deadlock and simplifies ordering reasoning.
9.5 Interrupts
RISC-V interrupts are delivered to a hart based on the global interrupt-enable bits and the PLIC. Ordering constraints:
- A hart that disables interrupts and then writes to a shared
variable does not need a
FENCEbefore the interrupt-disable write; the write tomstatusorsstatusis itself a CSR write and is ordered architecturally. - A hart that claims an interrupt at the PLIC and then reads a
device register must
FENCEbetween the claim and the read to ensure the claim is globally visible (the claim is a load from the PLIC's claim register; the read is a load to the device). - The interrupt-return path uses
mret/sret, which is architecturally ordered with respect to prior CSR writes. - Trap delivery and
xRETdo not include implicit memory fences; the OS must insert them.
XH-1's interrupt controller integration must ensure that the interrupt-acknowledge signal is recognized only after the relevant memory accesses are ordered.
9.6 Operating System
The OS uses FENCE heavily:
- Spinlocks and mutexes: the lock acquire is typically a
LR/SCsequence; the lock release is anSCor aFENCEfollowed by aSWto the lock variable. - Page-table manipulation:
SFENCE.VMAafter every PTE write. - I/O:
FENCEbefore/after device accesses (see §6). - Cross-core TLB shootdown: an IPI sender writes the
invalidation request to a shared memory location, then sends an
IPI, and the receiver must see the request after the IPI — this
requires a
FENCEon the sender side. - RCU grace periods: RCU relies on the memory model's
guarantees; a weak model like RVWMO requires explicit
FENCEat grace-period boundaries. - Boot and SMP startup: the boot hart must
FENCEbefore starting secondary harts to ensure the trampoline code and data are visible.
XH-1's OS port (Linux, or a research kernel) must insert FENCE
at all these sites. The cost of a FENCE in a 128-core system
makes the boot-time and idle-time costs small in absolute terms
but potentially large in relative terms; an in-order core that
spends 200 cycles per FENCE will spend a noticeable fraction of
its time in fences if it idles on a spinlock.
Open: Whether XH-1 ports a vanilla Linux or a research kernel. The memory-ordering analysis in this document assumes vanilla Linux semantics but applies in both cases.
9.7 Verification
Memory ordering is famously difficult to verify because:
- The bug is typically a very rare, non-deterministic interleaving.
- The legal reorderings are many; enumerating them is hard.
- The coherence protocol and the memory model interact in subtle ways.
XH-1's verification strategy [Proposed]:
- Litmus tests for the memory model. Run a corpus of litmus tests (based on the RISC-V memory-model litmus test suite, plus XH-1-specific tests) on RTL simulation. The tests must include the published RVWMO tests and tests specific to XH-1's coherence protocol.
- Coherence protocol model checking. Use a formal model (e.g., Murphi, TLA+) of the directory protocol; prove invariant properties.
- Pipeline formal verification. Verify the LSU's local reordering rules with bounded model checking.
- System-level simulation. Boot Linux on the RTL or a
high-level model; run stress workloads (e.g.,
perf bench futex, RCU torture test). - Stress with random delay injection. Vary latency on each interconnect path randomly to expose ordering bugs.
The published herd7 tool from Cambridge and the diy7 litmus-test framework are the standard tools for RVWMO verification. Operational models like Zam are more amenable to formal reasoning than axiomatic models; XH-1 should [Proposed] support both for verification.
Open: Whether XH-1's verification flow includes formal proofs of the memory ordering contract end-to-end, or whether simulation-based testing is the primary line of defense. A research project may choose the latter; a commercial project should aim for the former.
9.8 Performance
Memory ordering has two performance axes:
- Steady-state throughput: how fast can the system process ordinary memory traffic. This is dominated by coherence and cache behavior, not by ordering per se.
- Fence/atomic latency and throughput: how fast is a synchronization primitive. This is where ordering directly impacts performance.
For a 128-core system, the second axis is the more important
contributor to scalability. A workload with frequent
synchronization (e.g., a parallel database, a many-core
machine-learning inference engine) is dominated by FENCE and
AMO latency.
Quantitative note (rough estimate): Assuming a 1 GHz in-order core, a directory-coherent L2 slice with a hop latency of ~10 ns per mesh traversal, and an average of 3–5 hops per coherence transaction, a single uncontended
FENCEon a hot line might take ~30–50 ns (30–50 cycles). A contendedFENCE(multiple sharers) requiring invalidation acks could take 100–300 ns (100–300 cycles). On a 128-core workload that fences once per critical section, this is the dominant cost.
XH-1's performance evaluation [Proposed]:
- Measure
FENCElatency distribution under varying contention and across the mesh (corner vs. center tile). - Measure AMO latency under contention.
- Measure the effect of
aq/rlvs. explicitFENCEin the OS. - Compare against an ideal "no ordering cost" baseline.
10. Specific RISC-V Considerations
10.1 The Ztso Extension (Optional)
Ztso provides TSO (Total Store Order) at every hart. With
Ztso, the implementation guarantees that stores are not
reordered with respect to subsequent loads, simplifying the
software model.
Proposed XH-1 stance: Do not advertise
Ztsofor the first silicon. The hardware cost of supporting TSO at every hart is a tighter store-load ordering in the LSU, and the benefit is only felt by software that is willing to be incompatible with weak-ordering implementations. The standard RISC-V Linux kernel is RVWMO-compatible and does not requireZtso. Re-evaluate if a major software target requires it.
10.2 The Zam Extension (If Adopted)
Zam is a proposed extension that provides an operational
memory model in addition to the axiomatic RVWMO. It does not
change the contract; it provides operational clarity. XH-1
[Proposed] will document its memory model in terms of Zam
where possible because operational models are easier to
implement and verify.
10.3 The PTE-Based Memory Type (PBMT)
XH-1's MMU must support the PBMT field and route accesses to the appropriate ordering class. Specifically:
- PMA = NC (non-cacheable): accesses go to the device port; load/store ordering is enforced by the device-port controller.
- PMA = IO (weakly ordered I/O): same as NC, but with explicit reordering rules.
- PMA = WEAK (cacheable but weakly ordered): accesses are cached but ordered weakly; useful for write-combining buffers.
- PMA = STRONG (cacheable, strongly ordered): same as cacheable but with fence semantics on every access; rarely used.
Open: Whether XH-1 supports all PBMT values or a subset. A research project may support all four; a commercial project may support only NC and cacheable.
10.4 The fence.i and Instruction-Fetch Coherence
FENCE.I orders stores w.r.t. instruction fetches on the same
hart. Cross-hart, the software idiom is:
// hart A:
write_code(addr, code);
__sync_synchronize(); // FENCE
send_ipi(hart_B);
// hart B:
handle_ipi();
__sync_synchronize(); // FENCE
__fence_i(); // FENCE.I
icache_flush();
jump_to(addr);
XH-1's I-cache is a VIPT cache; on an aliasing configuration (small pages, large cache), the OS must use the recommended sequence (or the larger page size). The architecture team will document this.
10.5 TSO and RVWMO Differences for Software
Although the focus here is hardware, it is worth noting for completeness:
- Under TSO, ordinary loads and stores have program order semantics. Under RVWMO, they do not.
- A program that is correct under TSO may not be correct under
RVWMO without additional fences. The Linux kernel's
memory-model code (
tools/memory-model/) contains the required fences for RVWMO.
XH-1's port will use the RVWMO-correct fences; no special
support is required in hardware beyond honoring the standard
FENCE.
11. Unresolved Design Questions
The following questions remain open for the XH-1 architecture:
- In-order vs. out-of-order pipeline. In-order is proposed,
but the decision is not final. An OoO pipeline changes the
MOB design and may improve
FENCElatency. Ztsoadoption. Whether to support TSO at every hart.Zamadoption. Whether to require the operational memory model in the specification.- Coherence protocol. MSI vs. MESI vs. MOESI. The choice affects the number of transitions and therefore the number of coherence-induced ordering edges. (See the coherence research document.)
- Interconnect topology. Mesh vs. ring vs. hierarchical
bus+mesh. The choice affects
FENCElatency and the directory's role. - PBMT support scope. Which PTE memory types XH-1 implements.
- Per-line vs. per-word reservation. Affects
LR/SCsuccess rates. - Fence-ack message class. Whether to add a dedicated fence-ack message (Approach C) or rely entirely on per-core token tracking (Approach B).
- Verification strategy. Simulation-based vs. formal; end-to-end memory model proofs.
- OS port. Linux vs. a research kernel; this affects
the set of
FENCEsites the system must support.
12. Recommendations
Where evidence supports a recommendation, the following are XH-1's proposed positions:
- Adopt RVWMO as the memory model. This is the standard RISC-V model and is the most widely supported by software.
- Implement an in-order pipeline with Approach C (Directory-Assisted Ordering) and Approach B (Per-Core Token Completion). This combination scales to 128 cores while keeping the per-core pipeline simple.
- Use a directory-based MESI coherence protocol at the L2 slice. The directory is the natural serialization point for memory ordering. (See the coherence research document for the full argument.)
- Add a dedicated fence-ack message class to the coherence protocol. This avoids overloading existing messages and reduces verification ambiguity.
- Honour the standard
FENCE,FENCE.I,SFENCE.VMA,LR/SC, and AMO semantics. No architectural deviations. - Document the cross-hart self-modifying-code sequence in the programmer's manual. Use the standard fence pair.
- Implement the PTE-based memory types (PBMT) for at least NC and cacheable. Other types are deferred.
- Provide a memory-model litmus-test suite as part of the verification deliverables. This is essential for catching ordering bugs.
- Measure
FENCEand AMO latency across the mesh in performance evaluation. Corner tiles have different latency from center tiles; the asymmetry is a real performance concern.
13. References
The following sources informed this document. URLs and publication details are given where known; the reader should verify currency, as the RISC-V ecosystem is rapidly evolving.
- RISC-V International, RISC-V Instruction Set Manual, Volume I: Unprivileged ISA. Appendix A, "RVWMO: The RISC-V Weak Memory Ordering Model."
- RISC-V International, RISC-V Instruction Set Manual, Volume II: Privileged ISA. Sections on fence and memory types (PBMT).
- Pulte, C., et al. "Zam: A Complete and Verifiable Operational Memory Model for RISC-V." (Reference to the operational memory model paper.)
- Owens, S., Sarkar, S., Sewell, P. "A Better x86 Memory
Model: x86-TSO." (TSO reference; the TSO model is the
inspiration for
Ztso.) - Lamport, L. "How to Make a Multiprocessor Computer That Correctly Executes Multiprocess Programs." IEEE TC, 1979. (The foundational sequential consistency paper.)
- Adve, S. V., Gharachorloo, K. "Shared Memory Consistency Models: A Tutorial." IEEE Computer, 1996. (Tutorial on memory models.)
- Martin, M. M. K., et al. "Token Coherence." (Token-based ordering reference, applicable to Approach B.)
- ACM SIGOPS Operating Systems Review. Linux Kernel Memory
Model Documentation. (
tools/memory-model/in the Linux kernel source tree.) - AMD. AMD64 Architecture Programmer's Manual, Volume 2. (TSO reference for comparison.)
- Intel. Intel 64 and IA-32 Architectures Software Developer's Manual, Volume 3. (TSO reference for comparison.)
- ARM. ARM Architecture Reference Manual. (Comparison with ARM's weak memory model.)
Specific tool references:
- herd7: litmus-test framework for weak memory models, by Cambridge. https://github.com/herd/herdtools7
- diy7: companion tool for herd7.
- Murphi, TLA+: protocol model checkers used for coherence protocol verification.
The RISC-V Memory Model working group's published documentation, including the operational model papers and the litmus-test corpus, is the primary source for the RVWMO contract. Specifics of XH-1's coherence protocol are out of scope for this document and are addressed in the coherence research document.
14. Document Status
This document is a research proposal. It does not represent a finalized XH-1 architectural commitment. Items marked [Proposed] are engineering recommendations under evaluation; items marked [Open] are unresolved design questions.
The next revisions of this document should:
- Reflect decisions on the open questions in §11.
- Add quantitative
FENCElatency measurements once RTL or high-level models are available. - Incorporate results from the verification litmus-test suite.
- Align with the final coherence protocol choice.
- Document the OS-level
FENCEinsertion policy.
End of document.