Files

997 lines
40 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Memory Ordering
> **Status:** Research proposal for the XH-1 128-core RISC-V processor.
> Nothing in this document is a finalized architectural commitment unless
> explicitly stated. Items marked **[Proposed]** are engineering
> recommendations under evaluation; items marked **[Open]** are unresolved
> design questions.
---
## 1. Scope and Definitions
**Memory ordering** defines the set of rules that determine when memory
operations issued by one or more hardware threads become visible to other
hardware threads, and in what order such visibility is required to occur.
In a single-core, single-threaded system, "ordering" is trivial: program
order matches execution order matches visibility order. As soon as a
processor exposes:
- out-of-order execution,
- multiple hardware threads (SMT),
- multiple cores with private caches,
- coherent or non-coherent interconnects, or
- devices that perform DMA,
…memory ordering becomes a first-class architectural and microarchitectural
concern. For a 128-core design like XH-1, it is one of the dominant
correctness and performance constraints.
The RISC-V ISA delegates most ordering policy to a **Memory Consistency
Model (MCM)** specification rather than baking it into the base ISA. The
current published MCM is documented in the RISC-V specification volume
*“RISC-V Instruction Set Manual, Volume I: Unprivileged ISA”*, Appendix A —
*"RVWMO: The RISC-V Weak Memory Ordering Model"*. RVWMO is the canonical
model; "Zam" (the "A" extension candidate) is a candidate formalization
of the same weak-ordering contract with additional operational
clarifications.
This document covers:
1. The RVWMO contract XH-1 must honor.
2. Microarchitectural mechanisms needed to enforce that contract.
3. Scalability concerns specific to 128 cores.
4. Interactions with coherence, the interconnect, the pipeline, the
OS, and verification.
---
## 2. The RVWMO Contract XH-1 Must Honor
RVWMO is defined by:
- **Preserved Program Order (PPO):** a set of (microarchitectural)
orderings between pairs of memory operations from the same hart that
the implementation must never violate. PPO rules include the obvious
ones (a load following a store to the same address, register
dependencies, fences such as `FENCE`, `FENCE.I`, etc.) and the
subtler "axiomatic" PPO rules covering acquire/release semantics on
ordinary AMOs.
- **Memory Model Axioms:** a set of forbidden loadload, loadstore,
and storestore reorderings when no PPO rule requires otherwise
(e.g., "no load may be reordered with a store to a different
location" except under specific conditions).
- **Syntactic Dependencies:** register dependencies via the
destination register of a load are sufficient to constrain certain
reorderings, mirroring the original SPARC RMO/TSO literature.
The base RVWMO does not mandate TSO; an implementation may run hart-local
memory accesses in any order consistent with the PPO rules, with a few
mandated orderings (e.g., loads after stores to the *same* address are
preserved).
Key architectural levers available to software:
| Construct | Effect (informal) |
| -------------------------- | ------------------------------------------------------- |
| `FENCE` (various `pred`/`succ`/fm) | Orders memory operations as specified. |
| `FENCE.I` | Orders stores to instruction memory w.r.t. subsequent instruction fetches on the same hart. |
| `FENCE.VMA` | Orders SFENCE.VMA-style TLB management. |
| `SFENCE.VMA` | Orders stores w.r.t. subsequent TLB refills on the same hart. |
| Acquire/release on AMOs | Provides PPO ordering for paired RMW operations. |
| `LR`/`SC` (LR/SC pair) | Ordering anchored to the reservation. |
| `Zalrsc` / `Zaamo` | Standard atomic extensions. |
| `Zifencei` | Required for `FENCE.I`. |
| `Zihintpause` | Hint only; no ordering effect. |
| RV32/64 `fence` on devices | Device ordering via PTE bits (see §10.3). |
**XH-1 requirement:** the ISA-visible contract of every memory and
synchronization instruction listed in the *Unprivileged ISA* and any
extension XH-1 implements (`A`, `Zifencei`, `Zihintpause`, `Zalrsc`,
`Zaamo`, `Ztso` if adopted) must be honored by the combined
core/LSU/coherence/interconnect subsystem.
> **Open:** Whether XH-1 advertises `Ztso` (an optional RVWMO relaxation
> that grants TSO at every hart) or only RVWMO. `Ztso` simplifies
> software reasoning but slightly constrains hardware ordering
> flexibility. Recommendation deferred until microarchitecture is
> fixed.
---
## 3. Required Fences and Their Microarchitectural Meaning
The `FENCE` instruction is parameterized by:
- `pred`: which predecessor operations are ordered (R, W).
- `succ`: which successor operations are ordered (R, W).
- `fm` field: indicates "fence-mode"; value `0` is the standard fence,
value `>=1` is reserved for implementations and is treated as a no-op
by standard tools. Custom XH-1 fence-mode encodings are **[Open]**.
XH-1 must ensure that a `FENCE` with predecessor set *P* and successor
set *S*:
- All memory operations of type *P* issued by this hart prior to the
`FENCE` are ordered before all memory operations of type *S* issued by
this hart after the `FENCE`, in the global memory order, to the extent
that the coherence protocol is willing to serialize them.
Microarchitecturally, the `FENCE` typically drains in-flight operations
of the predecessor classes through the load/store unit (LSU), the
coherence point, and the write-completion path. On XH-1 with private L1
and a shared L2 slice, "the coherence point" is the directory controller
on the L2 slice; see §7.
---
## 4. Why 128 Cores Changes the Problem
A small-count many-core can often "get away" with conservative
ordering: a single shared bus or ring, broadcast snoops, and a
full-fence on every store buffer drain. At 128 cores these approaches
do not scale:
1. **Bandwidth:** A serialized global fence requires all in-flight
traffic from the issuing core to drain, *and* the directory to
observe all acknowledgments before subsequent operations are
injected. With 128 cores contending for the directory and the
interconnect, the latency of a single `FENCE` can grow from
~1030 cycles (small system) to several hundred cycles. This
directly impacts lock-based critical sections, RCU grace periods,
and `atomic` operations in the OS.
2. **Queue occupancy:** Every core holds per-message-class queues
(read requests, write requests, invalidations, data responses,
writeback data). A fence must account for occupancy of queues
between the core and the L2 directory *and* between directories
across the mesh. Without back-pressure, an XH-1 fence could
complete locally but leave a "hole" in the global order.
3. **Coherence protocol overhead:** Directory-based coherence
(assumed for XH-1; see the coherence research document) replaces
broadcast snoop traffic with per-line directory lookups, but
introduces ack/grant ordering dependencies that the fence must
not skip past.
4. **Scalable coherence introduces non-atomic transitions.**
Directory state moves from M→I→S→M through intermediate invalidation
acknowledgment states. A store-conditional (`SC`) on one core that
has been racing with a remote `LR` may observe the line in
intermediate states; the memory model must be specified for these.
5. **Contention of fence-induced drains:** When 128 cores all execute
`FENCE` in a short window (e.g., a barrier in a parallel workload),
the interconnect can become congested with drain acknowledgments.
The microarchitecture must avoid feedback-loop deadlock between
the queues and the fence-acknowledgment path.
> **Proposed XH-1 stance:** Treat memory ordering as a system-level
> performance constraint, not just a per-core correctness one. Evaluate
> the cost of every `FENCE` site in the OS, runtime, and synchronization
> libraries at the system level, not at single-core simulation.
---
## 5. Microarchitectural Mechanisms
The classic toolkit, applied to XH-1:
### 5.1 Load-Store Unit (LSU) Pipeline
A typical in-order XH-1 core will have:
- A **store buffer (STB)** between the integer pipeline and the L1
data cache.
- A **load queue (LDQ)** that issues loads past stores (speculatively)
and validates the result at commit, rolling back on a misspeculation.
- A **memory dependence predictor** (optional; see §9.1) to reduce the
cost of Ld→St forwarding checks.
The LSU's job relative to ordering is twofold:
1. **Honor PPO locally** — enforce that a load following a store to the
same address returns the store's value (or a younger value), and
that the AMO-acquire/release rules produce the required PPO edges.
2. **Provide a clean drain point** for `FENCE` — a place where the LSU
can mark "all pred-class operations up to this point have completed
globally."
### 5.2 Store Buffer Drain
A `FENCE` with `pred` including W must drain the store buffer. In a
directory-coherent system the relevant completion is not "store reached
L1" but "store has acquired the directory ownership token (e.g., `M`
response) and the directory has acknowledged writeback/Invalidate
acknowledgments for any invalidated peers."
For XH-1:
- **L1 hit, no coherence traffic:** store can complete in ~13 cycles
after reaching L1.
- **L1 miss, L2 hit, no sharers:** invalidation may be skipped (silent
upgrade to M); store completes when directory grants M.
- **L1 miss, L2 hit, sharers exist:** directory generates invalidation
messages to all sharers; the store completes when all invalidation
acknowledgments are received.
- **L1 miss, L2 miss:** directory forwards to home node; home node
replies with data and grants M; the store completes when data and
the M grant are both received.
Only the last three cases impose significant drain latency.
> **Proposed XH-1 stance:** The store buffer must be sized to hold
> stores pending completion, but the fence-drain path must not depend
> on the store buffer draining in order — older stores in the buffer
> that have already received their M response should not block the
> fence on younger stores still in flight, *unless* the fence
> successor set requires it.
### 5.3 Load Issue Rules
Loads are the harder side. RVWMO allows loads to be reordered with
prior stores to different addresses. XH-1's LSU will speculatively
issue loads past older stores; the load is "validated" at commit by
checking that the cache line has not been invalidated and that no
older store wrote to the same address.
A load that needs to "see" the result of an older store on the same
hart uses **store-to-load forwarding** through the store buffer; this
must respect the size and sign-extension rules of the load.
A load that needs to "see" the result of an older store on a *different*
hart (which is what `FENCE`/`FENCE.RW` enforces) must wait for that
remote store to complete at the directory and update the local
view via the coherence protocol. The directory's response is the
"happens-before" edge.
### 5.4 Acquire/Release on AMOs
A standard `AMO*` operation in RVWMO has the ordering semantics of an
acquire (load part) and release (store part) when used in `aq`/`rl`
positions. In RV32A/RV64A, the instructions `LR` and `SC` carry
acquire/release bits; AMOs (e.g., `AMOSWAP.W`, `AMOADD.D`) also carry
these bits.
For XH-1, an `aq` AMO must order prior loads/stores from the issuing
hart before the AMO, and the AMO before subsequent loads/stores.
Microarchitecturally:
- `aq`: drain the load and store queues (or at least establish PPO
edges) before issuing the AMO to the directory.
- `rl`: drain the load and store queues *after* the AMO completes.
- `aq+rl`: drain on both sides; this is the strong fence.
> **Open:** Whether XH-1 implements `aq`/`rl` purely in the LSU
> (with no global fence message) or with a lightweight fence message
> to the directory. Purely-local enforcement is faster but requires
> that the directory serialize the AMO at the line state machine.
### 5.5 `LR`/`SC`
`LR` loads with reservation; `SC` stores conditionally if the
reservation is still valid. XH-1's reservation representation is
proposed as a small set of reservation-set entries per core, indexed
by physical address, anchored in the L1. Coherence actions that
invalidate the reserved line (a snoop/Invalidate that the L1 sees)
must invalidate the reservation.
> **Proposed XH-1 reservation-granularity choice:** XH-1 reserves at
> cache-line granularity (the natural L1 line size) and treats
> forward-progress as the OS's responsibility. This is the
> conservative default; some designs use sub-line reservation to
> improve SC success rates, but the gain on typical lock code is
> modest.
### 5.6 `FENCE.I` and Self-Modifying Code
`FENCE.I` ensures that stores from the issuing hart that preceded the
`FENCE.I` are visible to subsequent instruction fetches on the same
hart. It does **not** order stores from other harts.
For XH-1, `FENCE.I` typically:
1. Drains the store buffer (in the W sense).
2. Invalidates the L1 I-cache (or the local portion of a unified L1).
3. Optionally issues a `FENCE` internally so that the store is
visible at the directory before the I-fetch resumes.
A cross-hart self-modifying-code sequence (hart A writes code, hart B
executes it) requires an explicit `FENCE` on hart A and a `FENCE.I` on
hart B (per the RISC-V recommendation), plus coherence to propagate
the new data. This is awkward and XH-1 will document it in the
programmer's manual.
### 5.7 `SFENCE.VMA`
`SFENCE.VMA` orders prior stores to PTEs before subsequent
address-translation results on the same hart. XH-1's TLB refill
engine must observe the store-completion signal; a typical
implementation stalls the pipeline at the `SFENCE.VMA` until the
store buffer drains.
---
## 6. System-Level Ordering Devices
Beyond per-hart mechanisms, XH-1 needs system-level ordering for
I/O and devices. The RISC-V privileged specification provides two
mechanisms:
### 6.1 I/O Ordering Bits in PTE
A PTE can have the **PBMT** field (Page-Based Memory Types), which can
be marked as:
- **Non-cacheable, idempotent, weakly-ordered (NC, IO).**
- **Non-cacheable, idempotent, strongly-ordered (NC, I/O-strong).**
XH-1 must honor these by:
- Routing accesses to the appropriate device port on the interconnect.
- For weakly-ordered I/O, treating the access as device but allowing
reordering with respect to other weakly-ordered I/O of the same type.
- For strongly-ordered I/O, treating the access as a full-fence —
the load or store completes in program order with respect to all
other accesses from the same hart.
### 6.2 Memory-Mapped I/O Fences
XH-1 will provide a `FENCE` variant (the standard `FENCE pred=rw,
succ=rw, fm=0`) that software uses to order I/O accesses. The
microarchitecture must ensure the drain endpoint includes the
device-port queue, not just the L1/L2.
### 6.3 Interrupt Acknowledgment and Ordering
RISC-V interrupt acknowledgment is itself a memory-mapped I/O
write. The microarchitecture must guarantee that the write to the
*Platform-Level Interrupt Controller* (PLIC) is ordered with respect
to the I/O operations it is acknowledging. The standard idiom is a
`FENCE` before the claim/complete operation. XH-1's PLIC integration
must be designed so that the `FENCE` does not require extra
microarchitectural support beyond the standard fence.
> **Open:** Whether XH-1 implements a **"fence fast path"** for
> accesses to specific device memory regions (e.g., the PLIC) that
> bypasses the directory and uses a dedicated sideband signal.
> Such optimizations can reduce interrupt latency by tens of cycles
> but complicate coherence reasoning. Not recommended for the first
> silicon.
---
## 7. Coherence Interaction
XH-1's coherence protocol (described in the coherence research
document) is directory-based, MSI or MESI, with a distributed
directory across the L2 slices. Memory ordering depends on the
coherence protocol's notion of "completion."
For a store to be considered "globally performed":
- The line is held in **Modified** state by the requesting core.
- All prior sharers have sent Invalidate Acknowledgments to the
directory (or the directory has confirmed there were no sharers).
- The directory has logged the new owner.
For a load to be considered "globally performed":
- The line is held in **Shared** (or **Modified** if exclusive) by
the requesting core.
- The data response from the directory (or forwarded from a peer) is
the value the directory currently authorizes.
> **Important subtlety:** A `FENCE W,W` between two stores on the
> same hart must ensure that store 1 has acquired its M grant *and*
> the directory has processed the invalidations *before* store 2
> acquires its M grant. In directory terms, store 2 cannot be issued
> until store 1's M grant has been observed at the directory. XH-1
> must implement this as an ordering dependency at the directory's
> L2 controller, not as a per-line serialization at the directory
> (which would be too coarse and too slow).
### 7.1 Coherence-Induced Ordering Pitfalls
Common pitfalls in directory-based designs:
- **Directory ack starvation under fence pressure:** when many cores
fence simultaneously, the directory must serialize invalidation
acknowledgments without deadlock.
- **Coherence downgrade during fence:** a core that downgrades M→S
while a fence is in progress must not let younger loads proceed
past the fence.
- **Silent upgrade (no-invalidate transition) on store:** a clean
line upgraded to M without invalidation must still be ordered
with respect to other cores' loads that observed the line in S.
The directory must explicitly synchronize this transition.
### 7.2 Ordering at the Interconnect
XH-1's interconnect is a mesh (proposed; see interconnect research
document). The mesh delivers messages in-order *between a given pair
of routers*, but globally the order is not guaranteed. The directory
controller is the serialization point.
For ordering, XH-1's mesh must:
- Not reorder different message classes on the same virtual channel.
- Support a "tail" or "credit" signal so that a fence-acknowledgment
message can be generated only when all prior messages on the same
path have been delivered.
---
## 8. Implementation Approaches Compared
This section lists the realistic implementation strategies for an
in-order XH-1 core. The choice of in-order vs. out-of-order pipeline
is discussed in §9.1; here we focus on ordering-specific mechanisms.
### 8.1 Approach A: Conservative Per-Core Fence (Baseline)
- The LSU drains the relevant queue classes fully on every `FENCE`.
- The store buffer's "drain complete" is defined as "every store has
received its directory M grant."
- A `FENCE` is implemented as a special instruction that stalls the
pipeline until drain completes.
- Acquire/release AMOs are implemented as full `FENCE`s around the
AMO at the LSU.
- `SFENCE.VMA` stalls the pipeline until the store buffer drains.
**Pros:**
- Simple to verify. The fence semantics maps directly to "queue
empty."
- Predictable performance for small core counts.
- Easy to reason about coherence interactions.
**Cons:**
- Poor scalability. At 128 cores, the worst-case `FENCE` latency
grows because the drain point is global, not local.
- Overly conservative for `aq`/`rl` AMOs, which do not require a
full global drain on either side — only ordering relative to
surrounding ordinary accesses.
### 8.2 Approach B: Token-Based Completion
- The LSU attaches a monotonically increasing **completion token**
to each outstanding operation.
- A `FENCE` with pred P, succ S becomes: "wait until all operations
of type P issued before this `FENCE` have produced a token greater
than or equal to the fence's token."
- The token is generated when the directory acknowledges the
operation; this is exactly the "global completion" notion.
- Acquire/Release is implemented as a fence of the appropriate
type around the AMO at the LSU.
**Pros:**
- Fences do not require the entire store buffer to drain; an empty
store buffer is sufficient, but so is a store buffer where all
older stores have completed.
- Naturally parallel: many stores can be outstanding, and the
fence only waits for the *oldest* of them.
- Slightly more complex but a well-understood technique (used in
academic and commercial cores).
**Cons:**
- Per-message token bookkeeping is more expensive than a simple
"queue empty" check.
- Verification requires modeling tokens at the testbench, which is
a non-trivial simulation infrastructure investment.
### 8.3 Approach C: Directory-Assisted Ordering
- A special "fence-ack" message class is added to the coherence
protocol.
- A `FENCE` instruction sends a fence-ack request to the directory;
the directory replies only after all prior messages from the
issuing core have been serialized.
- The directory tracks per-core "last-seen" sequence numbers to
arbitrate fence-ack requests against ongoing traffic.
**Pros:**
- The directory is the natural serialization point; leveraging it
avoids per-core reinvention of global ordering.
- Scales well to large core counts because the directory already
handles 128-core traffic.
- Allows the per-core LSU to continue issuing younger operations
while the fence is in progress, as long as they are ordered
correctly.
**Cons:**
- Adds protocol complexity (a new message class, directory state,
ack/grant logic).
- Risk of fence-ack starvation under heavy traffic.
- Requires careful design to avoid deadlock with normal coherence
traffic (classic message-class-dependency cycle).
### 8.4 Approach D: Speculative OoO + MOB
This is the out-of-order approach. The LSU becomes a Memory
Ordering Buffer (MOB) that tracks per-operation completion and
issues fences by tracking the "oldest in-flight" of each class.
This is how high-performance OoO cores implement fences; the
trade-off is pipeline depth and complexity.
XH-1 is proposed as **in-order** for power and predictability; see
§9.1. OoO is listed for completeness.
**Pros:**
- Highest performance on workloads with deep memory-level
parallelism.
- Fences can be "free" if no relevant older operation is in flight.
**Cons:**
- Significantly larger area and power.
- More verification effort, especially for the MOB's speculative
reordering.
- Diminishing returns on memory-bound workloads at 128 cores (the
bottleneck is the interconnect, not the per-core issue rate).
### 8.5 Summary of Approaches
| Approach | Perf (single) | Perf (128) | Area | Verif. effort | Risk |
| -------- | ------------- | ---------- | ---- | ------------- | ---- |
| A: Cons. | Low | Lowest | Lowest | Lowest | Low |
| B: Token | Medium | Medium | Low | Medium | Low |
| C: Dir-Assisted | Medium-high | High | Medium | High | Medium |
| D: OoO+MOB | High | High | High | Highest | High |
> **Proposed XH-1 implementation:** **Approach C (Directory-Assisted
> Ordering)** as the primary, with **Approach B (Token-Based
> Completion)** as the per-core implementation strategy. The
> directory handles global serialization, the LSU handles local
> token tracking. This combination is the most natural for a
> directory-coherent 128-core design.
---
## 9. Interactions with the Pipeline and Other Subsystems
### 9.1 Pipeline Choice (In-Order vs. Out-of-Order)
XH-1 is **[Proposed] in-order** for the first silicon. The reasons:
- Power envelope. 128 cores at high frequency is a thermal
challenge; an in-order core can run cooler.
- Predictable timing. Real-time and safety-critical workloads
benefit from deterministic `FENCE` and `FENCE.I` latencies.
- Verification cost. An OoO core's MOB is a major verification
burden; a 128-core OoO design is a multi-year project.
An in-order pipeline interacts with ordering as follows:
- The LSU is non-speculative on the store side (stores are
committed when the integer pipeline commits them) and lightly
speculative on the load side (loads may issue before older
stores, with a structural dependency check at commit).
- `FENCE` stalls the pipeline until drain; the cost is a few tens
to a few hundred cycles depending on directory state.
- Acquire/release AMOs generate local fence events in the LSU.
An OoO pipeline would:
- Issue stores from the STB to L1, with the MOB tracking
completion.
- Issue loads speculatively past stores with prediction; the MOB
validates at commit.
- `FENCE` becomes a "wait for oldest in-flight of pred class to
complete" check; usually very fast.
> **Open:** Whether a future XH-2 could be OoO. The research
> project is not yet at that question.
### 9.2 Cache Hierarchy
Memory ordering is shaped by the cache hierarchy:
- **L1 D-cache:** the load-side ordering reference. A load that hits
in L1 sees a value that is at least as recent as the cache's
coherence state indicates. A load that misses in L1 must go to
L2, where the directory arbitrates the order.
- **L2 slice (per tile):** home of the directory entry; serializes
ordering for the addresses it owns.
- **L3 (shared, distributed):** a victim cache and last-level
coherence point. XH-1's directory is the L2 directory; L3 is
treated as a backing store. Invalidation ordering for XH-1
therefore ends at L2.
- **L1 I-cache:** the `FENCE.I` target. The I-cache must be flushed
or invalidated to ensure stores are visible.
The cache hierarchy directly affects `FENCE` latency because the
"completion" point is the L2 directory. A `FENCE W,W` in a system
with L1-only ordering would be much cheaper (a few cycles) but
incorrect; the L2 directory must be in the loop.
### 9.3 Memory System
The memory system (DRAM controllers, on-chip memory, scratchpads)
must respect the same ordering. For XH-1, accesses to on-chip
scratchpad memory bypass the coherence protocol entirely; they
are ordered by the scratchpad controller.
> **Open:** Whether XH-1 has a per-tile scratchpad or a shared
> on-chip SRAM region. If shared, ordering between scratchpad and
cache-coherent regions is a research question.
### 9.4 Interconnect
The mesh interconnect has ordering properties (per-link in-order,
globally non-ordered) that determine what the directory and
endpoints must do to maintain memory ordering. The "virtual
channel" abstraction in the mesh is the unit of ordering; XH-1
**[Proposed]** uses:
- **VC0:** Coherence requests (in-order).
- **VC1:** Coherence responses (in-order).
- **VC2:** Data messages (in-order).
- **VC3:** Fence-ack and other ordering-sensitive messages
(in-order).
This separation avoids deadlock and simplifies ordering reasoning.
### 9.5 Interrupts
RISC-V interrupts are delivered to a hart based on the *global
interrupt-enable* bits and the PLIC. Ordering constraints:
- A hart that disables interrupts and then writes to a shared
variable does not need a `FENCE` before the interrupt-disable
write; the write to `mstatus` or `sstatus` is itself a CSR write
and is ordered architecturally.
- A hart that claims an interrupt at the PLIC and then reads a
device register must `FENCE` between the claim and the read to
ensure the claim is globally visible (the claim is a load from
the PLIC's claim register; the read is a load to the device).
- The interrupt-return path uses `mret`/`sret`, which is
architecturally ordered with respect to prior CSR writes.
- Trap delivery and `xRET` do not include implicit memory fences;
the OS must insert them.
XH-1's interrupt controller integration must ensure that the
interrupt-acknowledge signal is recognized only after the
relevant memory accesses are ordered.
### 9.6 Operating System
The OS uses `FENCE` heavily:
- **Spinlocks and mutexes:** the lock acquire is typically a
`LR`/`SC` sequence; the lock release is an `SC` or a `FENCE`
followed by a `SW` to the lock variable.
- **Page-table manipulation:** `SFENCE.VMA` after every PTE write.
- **I/O:** `FENCE` before/after device accesses (see §6).
- **Cross-core TLB shootdown:** an IPI sender writes the
invalidation request to a shared memory location, then sends an
IPI, and the receiver must see the request after the IPI — this
requires a `FENCE` on the sender side.
- **RCU grace periods:** RCU relies on the memory model's
guarantees; a weak model like RVWMO requires explicit
`FENCE` at grace-period boundaries.
- **Boot and SMP startup:** the boot hart must `FENCE` before
starting secondary harts to ensure the trampoline code and data
are visible.
XH-1's OS port (Linux, or a research kernel) must insert `FENCE`
at all these sites. The cost of a `FENCE` in a 128-core system
makes the boot-time and idle-time costs small in absolute terms
but potentially large in *relative* terms; an in-order core that
spends 200 cycles per `FENCE` will spend a noticeable fraction of
its time in fences if it idles on a spinlock.
> **Open:** Whether XH-1 ports a vanilla Linux or a research
> kernel. The memory-ordering analysis in this document assumes
> vanilla Linux semantics but applies in both cases.
### 9.7 Verification
Memory ordering is famously difficult to verify because:
- The bug is typically a very rare, non-deterministic interleaving.
- The legal reorderings are many; enumerating them is hard.
- The coherence protocol and the memory model interact in subtle
ways.
XH-1's verification strategy **[Proposed]**:
1. **Litmus tests** for the memory model. Run a corpus of litmus
tests (based on the RISC-V memory-model litmus test suite, plus
XH-1-specific tests) on RTL simulation. The tests must include
the published RVWMO tests and tests specific to XH-1's
coherence protocol.
2. **Coherence protocol model checking.** Use a formal model
(e.g., Murphi, TLA+) of the directory protocol; prove
invariant properties.
3. **Pipeline formal verification.** Verify the LSU's local
reordering rules with bounded model checking.
4. **System-level simulation.** Boot Linux on the RTL or a
high-level model; run stress workloads (e.g., `perf bench
futex`, RCU torture test).
5. **Stress with random delay injection.** Vary latency on each
interconnect path randomly to expose ordering bugs.
The published **herd7** tool from Cambridge and the **diy7**
litmus-test framework are the standard tools for RVWMO
verification. **Operational models** like Zam are more amenable
to formal reasoning than axiomatic models; XH-1 should
**[Proposed]** support both for verification.
> **Open:** Whether XH-1's verification flow includes formal
> proofs of the memory ordering contract end-to-end, or whether
> simulation-based testing is the primary line of defense. A
> research project may choose the latter; a commercial project
> should aim for the former.
### 9.8 Performance
Memory ordering has two performance axes:
1. **Steady-state throughput:** how fast can the system process
ordinary memory traffic. This is dominated by coherence and
cache behavior, not by ordering per se.
2. **Fence/atomic latency and throughput:** how fast is a
synchronization primitive. This is where ordering directly
impacts performance.
For a 128-core system, the second axis is the more important
contributor to scalability. A workload with frequent
synchronization (e.g., a parallel database, a many-core
machine-learning inference engine) is dominated by `FENCE` and
AMO latency.
> **Quantitative note (rough estimate):** Assuming a 1 GHz in-order
> core, a directory-coherent L2 slice with a hop latency of ~10 ns
> per mesh traversal, and an average of 35 hops per coherence
> transaction, a single uncontended `FENCE` on a hot line might
> take ~3050 ns (3050 cycles). A contended `FENCE` (multiple
> sharers) requiring invalidation acks could take 100300 ns
> (100300 cycles). On a 128-core workload that fences once per
> critical section, this is the dominant cost.
XH-1's performance evaluation **[Proposed]**:
- Measure `FENCE` latency distribution under varying contention
and across the mesh (corner vs. center tile).
- Measure AMO latency under contention.
- Measure the effect of `aq`/`rl` vs. explicit `FENCE` in the
OS.
- Compare against an ideal "no ordering cost" baseline.
---
## 10. Specific RISC-V Considerations
### 10.1 The `Ztso` Extension (Optional)
`Ztso` provides TSO (Total Store Order) at every hart. With
`Ztso`, the implementation guarantees that stores are not
reordered with respect to subsequent loads, simplifying the
software model.
> **Proposed XH-1 stance:** Do not advertise `Ztso` for the
> first silicon. The hardware cost of supporting TSO at every
> hart is a tighter store-load ordering in the LSU, and the
> benefit is only felt by software that is willing to be
> incompatible with weak-ordering implementations. The standard
> RISC-V Linux kernel is RVWMO-compatible and does not require
> `Ztso`. Re-evaluate if a major software target requires it.
### 10.2 The `Zam` Extension (If Adopted)
`Zam` is a proposed extension that provides an operational
memory model in addition to the axiomatic RVWMO. It does not
change the contract; it provides operational clarity. XH-1
**[Proposed]** will document its memory model in terms of `Zam`
where possible because operational models are easier to
implement and verify.
### 10.3 The PTE-Based Memory Type (PBMT)
XH-1's MMU must support the PBMT field and route accesses to
the appropriate ordering class. Specifically:
- **PMA = NC (non-cacheable):** accesses go to the device port;
load/store ordering is enforced by the device-port controller.
- **PMA = IO (weakly ordered I/O):** same as NC, but with
explicit reordering rules.
- **PMA = WEAK (cacheable but weakly ordered):** accesses are
cached but ordered weakly; useful for write-combining buffers.
- **PMA = STRONG (cacheable, strongly ordered):** same as
cacheable but with fence semantics on every access; rarely
used.
> **Open:** Whether XH-1 supports all PBMT values or a subset.
> A research project may support all four; a commercial project
> may support only NC and cacheable.
### 10.4 The `fence.i` and Instruction-Fetch Coherence
`FENCE.I` orders stores w.r.t. instruction fetches on the same
hart. Cross-hart, the software idiom is:
```c
// hart A:
write_code(addr, code);
__sync_synchronize(); // FENCE
send_ipi(hart_B);
// hart B:
handle_ipi();
__sync_synchronize(); // FENCE
__fence_i(); // FENCE.I
icache_flush();
jump_to(addr);
```
XH-1's I-cache is a VIPT cache; on an aliasing configuration
(small pages, large cache), the OS must use the recommended
sequence (or the larger page size). The architecture team will
document this.
### 10.5 TSO and RVWMO Differences for Software
Although the focus here is hardware, it is worth noting for
completeness:
- Under TSO, ordinary loads and stores have program order
semantics. Under RVWMO, they do not.
- A program that is correct under TSO may not be correct under
RVWMO without additional fences. The Linux kernel's
memory-model code (`tools/memory-model/`) contains the
required fences for RVWMO.
XH-1's port will use the RVWMO-correct fences; no special
support is required in hardware beyond honoring the standard
`FENCE`.
---
## 11. Unresolved Design Questions
The following questions remain open for the XH-1 architecture:
1. **In-order vs. out-of-order pipeline.** In-order is proposed,
but the decision is not final. An OoO pipeline changes the
MOB design and may improve `FENCE` latency.
2. **`Ztso` adoption.** Whether to support TSO at every hart.
3. **`Zam` adoption.** Whether to require the operational memory
model in the specification.
4. **Coherence protocol.** MSI vs. MESI vs. MOESI. The choice
affects the number of transitions and therefore the number
of coherence-induced ordering edges. (See the coherence
research document.)
5. **Interconnect topology.** Mesh vs. ring vs. hierarchical
bus+mesh. The choice affects `FENCE` latency and the
directory's role.
6. **PBMT support scope.** Which PTE memory types XH-1
implements.
7. **Per-line vs. per-word reservation.** Affects `LR`/`SC`
success rates.
8. **Fence-ack message class.** Whether to add a dedicated
fence-ack message (Approach C) or rely entirely on
per-core token tracking (Approach B).
9. **Verification strategy.** Simulation-based vs. formal;
end-to-end memory model proofs.
10. **OS port.** Linux vs. a research kernel; this affects
the set of `FENCE` sites the system must support.
---
## 12. Recommendations
Where evidence supports a recommendation, the following are
XH-1's proposed positions:
1. **Adopt RVWMO as the memory model.** This is the standard
RISC-V model and is the most widely supported by software.
2. **Implement an in-order pipeline with Approach C
(Directory-Assisted Ordering) and Approach B (Per-Core
Token Completion).** This combination scales to 128 cores
while keeping the per-core pipeline simple.
3. **Use a directory-based MESI coherence protocol at the
L2 slice.** The directory is the natural serialization
point for memory ordering. (See the coherence research
document for the full argument.)
4. **Add a dedicated fence-ack message class to the coherence
protocol.** This avoids overloading existing messages and
reduces verification ambiguity.
5. **Honour the standard `FENCE`, `FENCE.I`, `SFENCE.VMA`,
`LR`/`SC`, and AMO semantics.** No architectural deviations.
6. **Document the cross-hart self-modifying-code sequence in
the programmer's manual.** Use the standard fence pair.
7. **Implement the PTE-based memory types (PBMT) for at
least NC and cacheable.** Other types are deferred.
8. **Provide a memory-model litmus-test suite as part of
the verification deliverables.** This is essential for
catching ordering bugs.
9. **Measure `FENCE` and AMO latency across the mesh in
performance evaluation.** Corner tiles have different
latency from center tiles; the asymmetry is a real
performance concern.
---
## 13. References
The following sources informed this document. URLs and
publication details are given where known; the reader should
verify currency, as the RISC-V ecosystem is rapidly evolving.
- RISC-V International, *RISC-V Instruction Set Manual,
Volume I: Unprivileged ISA*. Appendix A, "RVWMO: The RISC-V
Weak Memory Ordering Model."
- RISC-V International, *RISC-V Instruction Set Manual,
Volume II: Privileged ISA*. Sections on fence and memory
types (PBMT).
- Pulte, C., et al. *"Zam: A Complete and Verifiable
Operational Memory Model for RISC-V."* (Reference to the
operational memory model paper.)
- Owens, S., Sarkar, S., Sewell, P. *"A Better x86 Memory
Model: x86-TSO."* (TSO reference; the TSO model is the
inspiration for `Ztso`.)
- Lamport, L. *"How to Make a Multiprocessor Computer That
Correctly Executes Multiprocess Programs."* IEEE TC, 1979.
(The foundational sequential consistency paper.)
- Adve, S. V., Gharachorloo, K. *"Shared Memory Consistency
Models: A Tutorial."* IEEE Computer, 1996. (Tutorial on
memory models.)
- Martin, M. M. K., et al. *"Token Coherence."* (Token-based
ordering reference, applicable to Approach B.)
- ACM SIGOPS Operating Systems Review. *Linux Kernel Memory
Model Documentation.* (`tools/memory-model/` in the Linux
kernel source tree.)
- AMD. *AMD64 Architecture Programmer's Manual*, Volume 2.
(TSO reference for comparison.)
- Intel. *Intel 64 and IA-32 Architectures Software
Developer's Manual*, Volume 3. (TSO reference for
comparison.)
- ARM. *ARM Architecture Reference Manual*. (Comparison
with ARM's weak memory model.)
Specific tool references:
- **herd7**: litmus-test framework for weak memory models,
by Cambridge. https://github.com/herd/herdtools7
- **diy7**: companion tool for herd7.
- **Murphi**, **TLA+**: protocol model checkers used for
coherence protocol verification.
The RISC-V Memory Model working group's published
documentation, including the operational model papers and
the litmus-test corpus, is the primary source for the
RVWMO contract. Specifics of XH-1's coherence protocol are
out of scope for this document and are addressed in the
coherence research document.
---
## 14. Document Status
This document is a **research proposal**. It does not represent
a finalized XH-1 architectural commitment. Items marked
**[Proposed]** are engineering recommendations under evaluation;
items marked **[Open]** are unresolved design questions.
The next revisions of this document should:
- Reflect decisions on the open questions in §11.
- Add quantitative `FENCE` latency measurements once RTL or
high-level models are available.
- Incorporate results from the verification litmus-test
suite.
- Align with the final coherence protocol choice.
- Document the OS-level `FENCE` insertion policy.
---
*End of document.*