mirror of
https://github.com/allexanderbergmns/xh1-research.git
synced 2026-08-26 23:47:01 +00:00
997 lines
40 KiB
Markdown
997 lines
40 KiB
Markdown
# Memory Ordering
|
||
|
||
> **Status:** Research proposal for the XH-1 128-core RISC-V processor.
|
||
> Nothing in this document is a finalized architectural commitment unless
|
||
> explicitly stated. Items marked **[Proposed]** are engineering
|
||
> recommendations under evaluation; items marked **[Open]** are unresolved
|
||
> design questions.
|
||
|
||
---
|
||
|
||
## 1. Scope and Definitions
|
||
|
||
**Memory ordering** defines the set of rules that determine when memory
|
||
operations issued by one or more hardware threads become visible to other
|
||
hardware threads, and in what order such visibility is required to occur.
|
||
|
||
In a single-core, single-threaded system, "ordering" is trivial: program
|
||
order matches execution order matches visibility order. As soon as a
|
||
processor exposes:
|
||
|
||
- out-of-order execution,
|
||
- multiple hardware threads (SMT),
|
||
- multiple cores with private caches,
|
||
- coherent or non-coherent interconnects, or
|
||
- devices that perform DMA,
|
||
|
||
…memory ordering becomes a first-class architectural and microarchitectural
|
||
concern. For a 128-core design like XH-1, it is one of the dominant
|
||
correctness and performance constraints.
|
||
|
||
The RISC-V ISA delegates most ordering policy to a **Memory Consistency
|
||
Model (MCM)** specification rather than baking it into the base ISA. The
|
||
current published MCM is documented in the RISC-V specification volume
|
||
*“RISC-V Instruction Set Manual, Volume I: Unprivileged ISA”*, Appendix A —
|
||
*"RVWMO: The RISC-V Weak Memory Ordering Model"*. RVWMO is the canonical
|
||
model; "Zam" (the "A" extension candidate) is a candidate formalization
|
||
of the same weak-ordering contract with additional operational
|
||
clarifications.
|
||
|
||
This document covers:
|
||
|
||
1. The RVWMO contract XH-1 must honor.
|
||
2. Microarchitectural mechanisms needed to enforce that contract.
|
||
3. Scalability concerns specific to 128 cores.
|
||
4. Interactions with coherence, the interconnect, the pipeline, the
|
||
OS, and verification.
|
||
|
||
---
|
||
|
||
## 2. The RVWMO Contract XH-1 Must Honor
|
||
|
||
RVWMO is defined by:
|
||
|
||
- **Preserved Program Order (PPO):** a set of (microarchitectural)
|
||
orderings between pairs of memory operations from the same hart that
|
||
the implementation must never violate. PPO rules include the obvious
|
||
ones (a load following a store to the same address, register
|
||
dependencies, fences such as `FENCE`, `FENCE.I`, etc.) and the
|
||
subtler "axiomatic" PPO rules covering acquire/release semantics on
|
||
ordinary AMOs.
|
||
- **Memory Model Axioms:** a set of forbidden load–load, load–store,
|
||
and store–store reorderings when no PPO rule requires otherwise
|
||
(e.g., "no load may be reordered with a store to a different
|
||
location" except under specific conditions).
|
||
- **Syntactic Dependencies:** register dependencies via the
|
||
destination register of a load are sufficient to constrain certain
|
||
reorderings, mirroring the original SPARC RMO/TSO literature.
|
||
|
||
The base RVWMO does not mandate TSO; an implementation may run hart-local
|
||
memory accesses in any order consistent with the PPO rules, with a few
|
||
mandated orderings (e.g., loads after stores to the *same* address are
|
||
preserved).
|
||
|
||
Key architectural levers available to software:
|
||
|
||
| Construct | Effect (informal) |
|
||
| -------------------------- | ------------------------------------------------------- |
|
||
| `FENCE` (various `pred`/`succ`/fm) | Orders memory operations as specified. |
|
||
| `FENCE.I` | Orders stores to instruction memory w.r.t. subsequent instruction fetches on the same hart. |
|
||
| `FENCE.VMA` | Orders SFENCE.VMA-style TLB management. |
|
||
| `SFENCE.VMA` | Orders stores w.r.t. subsequent TLB refills on the same hart. |
|
||
| Acquire/release on AMOs | Provides PPO ordering for paired RMW operations. |
|
||
| `LR`/`SC` (LR/SC pair) | Ordering anchored to the reservation. |
|
||
| `Zalrsc` / `Zaamo` | Standard atomic extensions. |
|
||
| `Zifencei` | Required for `FENCE.I`. |
|
||
| `Zihintpause` | Hint only; no ordering effect. |
|
||
| RV32/64 `fence` on devices | Device ordering via PTE bits (see §10.3). |
|
||
|
||
**XH-1 requirement:** the ISA-visible contract of every memory and
|
||
synchronization instruction listed in the *Unprivileged ISA* and any
|
||
extension XH-1 implements (`A`, `Zifencei`, `Zihintpause`, `Zalrsc`,
|
||
`Zaamo`, `Ztso` if adopted) must be honored by the combined
|
||
core/LSU/coherence/interconnect subsystem.
|
||
|
||
> **Open:** Whether XH-1 advertises `Ztso` (an optional RVWMO relaxation
|
||
> that grants TSO at every hart) or only RVWMO. `Ztso` simplifies
|
||
> software reasoning but slightly constrains hardware ordering
|
||
> flexibility. Recommendation deferred until microarchitecture is
|
||
> fixed.
|
||
|
||
---
|
||
|
||
## 3. Required Fences and Their Microarchitectural Meaning
|
||
|
||
The `FENCE` instruction is parameterized by:
|
||
|
||
- `pred`: which predecessor operations are ordered (R, W).
|
||
- `succ`: which successor operations are ordered (R, W).
|
||
- `fm` field: indicates "fence-mode"; value `0` is the standard fence,
|
||
value `>=1` is reserved for implementations and is treated as a no-op
|
||
by standard tools. Custom XH-1 fence-mode encodings are **[Open]**.
|
||
|
||
XH-1 must ensure that a `FENCE` with predecessor set *P* and successor
|
||
set *S*:
|
||
|
||
- All memory operations of type *P* issued by this hart prior to the
|
||
`FENCE` are ordered before all memory operations of type *S* issued by
|
||
this hart after the `FENCE`, in the global memory order, to the extent
|
||
that the coherence protocol is willing to serialize them.
|
||
|
||
Microarchitecturally, the `FENCE` typically drains in-flight operations
|
||
of the predecessor classes through the load/store unit (LSU), the
|
||
coherence point, and the write-completion path. On XH-1 with private L1
|
||
and a shared L2 slice, "the coherence point" is the directory controller
|
||
on the L2 slice; see §7.
|
||
|
||
---
|
||
|
||
## 4. Why 128 Cores Changes the Problem
|
||
|
||
A small-count many-core can often "get away" with conservative
|
||
ordering: a single shared bus or ring, broadcast snoops, and a
|
||
full-fence on every store buffer drain. At 128 cores these approaches
|
||
do not scale:
|
||
|
||
1. **Bandwidth:** A serialized global fence requires all in-flight
|
||
traffic from the issuing core to drain, *and* the directory to
|
||
observe all acknowledgments before subsequent operations are
|
||
injected. With 128 cores contending for the directory and the
|
||
interconnect, the latency of a single `FENCE` can grow from
|
||
~10–30 cycles (small system) to several hundred cycles. This
|
||
directly impacts lock-based critical sections, RCU grace periods,
|
||
and `atomic` operations in the OS.
|
||
|
||
2. **Queue occupancy:** Every core holds per-message-class queues
|
||
(read requests, write requests, invalidations, data responses,
|
||
writeback data). A fence must account for occupancy of queues
|
||
between the core and the L2 directory *and* between directories
|
||
across the mesh. Without back-pressure, an XH-1 fence could
|
||
complete locally but leave a "hole" in the global order.
|
||
|
||
3. **Coherence protocol overhead:** Directory-based coherence
|
||
(assumed for XH-1; see the coherence research document) replaces
|
||
broadcast snoop traffic with per-line directory lookups, but
|
||
introduces ack/grant ordering dependencies that the fence must
|
||
not skip past.
|
||
|
||
4. **Scalable coherence introduces non-atomic transitions.**
|
||
Directory state moves from M→I→S→M through intermediate invalidation
|
||
acknowledgment states. A store-conditional (`SC`) on one core that
|
||
has been racing with a remote `LR` may observe the line in
|
||
intermediate states; the memory model must be specified for these.
|
||
|
||
5. **Contention of fence-induced drains:** When 128 cores all execute
|
||
`FENCE` in a short window (e.g., a barrier in a parallel workload),
|
||
the interconnect can become congested with drain acknowledgments.
|
||
The microarchitecture must avoid feedback-loop deadlock between
|
||
the queues and the fence-acknowledgment path.
|
||
|
||
> **Proposed XH-1 stance:** Treat memory ordering as a system-level
|
||
> performance constraint, not just a per-core correctness one. Evaluate
|
||
> the cost of every `FENCE` site in the OS, runtime, and synchronization
|
||
> libraries at the system level, not at single-core simulation.
|
||
|
||
---
|
||
|
||
## 5. Microarchitectural Mechanisms
|
||
|
||
The classic toolkit, applied to XH-1:
|
||
|
||
### 5.1 Load-Store Unit (LSU) Pipeline
|
||
|
||
A typical in-order XH-1 core will have:
|
||
|
||
- A **store buffer (STB)** between the integer pipeline and the L1
|
||
data cache.
|
||
- A **load queue (LDQ)** that issues loads past stores (speculatively)
|
||
and validates the result at commit, rolling back on a misspeculation.
|
||
- A **memory dependence predictor** (optional; see §9.1) to reduce the
|
||
cost of Ld→St forwarding checks.
|
||
|
||
The LSU's job relative to ordering is twofold:
|
||
|
||
1. **Honor PPO locally** — enforce that a load following a store to the
|
||
same address returns the store's value (or a younger value), and
|
||
that the AMO-acquire/release rules produce the required PPO edges.
|
||
2. **Provide a clean drain point** for `FENCE` — a place where the LSU
|
||
can mark "all pred-class operations up to this point have completed
|
||
globally."
|
||
|
||
### 5.2 Store Buffer Drain
|
||
|
||
A `FENCE` with `pred` including W must drain the store buffer. In a
|
||
directory-coherent system the relevant completion is not "store reached
|
||
L1" but "store has acquired the directory ownership token (e.g., `M`
|
||
response) and the directory has acknowledged writeback/Invalidate
|
||
acknowledgments for any invalidated peers."
|
||
|
||
For XH-1:
|
||
|
||
- **L1 hit, no coherence traffic:** store can complete in ~1–3 cycles
|
||
after reaching L1.
|
||
- **L1 miss, L2 hit, no sharers:** invalidation may be skipped (silent
|
||
upgrade to M); store completes when directory grants M.
|
||
- **L1 miss, L2 hit, sharers exist:** directory generates invalidation
|
||
messages to all sharers; the store completes when all invalidation
|
||
acknowledgments are received.
|
||
- **L1 miss, L2 miss:** directory forwards to home node; home node
|
||
replies with data and grants M; the store completes when data and
|
||
the M grant are both received.
|
||
|
||
Only the last three cases impose significant drain latency.
|
||
|
||
> **Proposed XH-1 stance:** The store buffer must be sized to hold
|
||
> stores pending completion, but the fence-drain path must not depend
|
||
> on the store buffer draining in order — older stores in the buffer
|
||
> that have already received their M response should not block the
|
||
> fence on younger stores still in flight, *unless* the fence
|
||
> successor set requires it.
|
||
|
||
### 5.3 Load Issue Rules
|
||
|
||
Loads are the harder side. RVWMO allows loads to be reordered with
|
||
prior stores to different addresses. XH-1's LSU will speculatively
|
||
issue loads past older stores; the load is "validated" at commit by
|
||
checking that the cache line has not been invalidated and that no
|
||
older store wrote to the same address.
|
||
|
||
A load that needs to "see" the result of an older store on the same
|
||
hart uses **store-to-load forwarding** through the store buffer; this
|
||
must respect the size and sign-extension rules of the load.
|
||
|
||
A load that needs to "see" the result of an older store on a *different*
|
||
hart (which is what `FENCE`/`FENCE.RW` enforces) must wait for that
|
||
remote store to complete at the directory and update the local
|
||
view via the coherence protocol. The directory's response is the
|
||
"happens-before" edge.
|
||
|
||
### 5.4 Acquire/Release on AMOs
|
||
|
||
A standard `AMO*` operation in RVWMO has the ordering semantics of an
|
||
acquire (load part) and release (store part) when used in `aq`/`rl`
|
||
positions. In RV32A/RV64A, the instructions `LR` and `SC` carry
|
||
acquire/release bits; AMOs (e.g., `AMOSWAP.W`, `AMOADD.D`) also carry
|
||
these bits.
|
||
|
||
For XH-1, an `aq` AMO must order prior loads/stores from the issuing
|
||
hart before the AMO, and the AMO before subsequent loads/stores.
|
||
Microarchitecturally:
|
||
|
||
- `aq`: drain the load and store queues (or at least establish PPO
|
||
edges) before issuing the AMO to the directory.
|
||
- `rl`: drain the load and store queues *after* the AMO completes.
|
||
- `aq+rl`: drain on both sides; this is the strong fence.
|
||
|
||
> **Open:** Whether XH-1 implements `aq`/`rl` purely in the LSU
|
||
> (with no global fence message) or with a lightweight fence message
|
||
> to the directory. Purely-local enforcement is faster but requires
|
||
> that the directory serialize the AMO at the line state machine.
|
||
|
||
### 5.5 `LR`/`SC`
|
||
|
||
`LR` loads with reservation; `SC` stores conditionally if the
|
||
reservation is still valid. XH-1's reservation representation is
|
||
proposed as a small set of reservation-set entries per core, indexed
|
||
by physical address, anchored in the L1. Coherence actions that
|
||
invalidate the reserved line (a snoop/Invalidate that the L1 sees)
|
||
must invalidate the reservation.
|
||
|
||
> **Proposed XH-1 reservation-granularity choice:** XH-1 reserves at
|
||
> cache-line granularity (the natural L1 line size) and treats
|
||
> forward-progress as the OS's responsibility. This is the
|
||
> conservative default; some designs use sub-line reservation to
|
||
> improve SC success rates, but the gain on typical lock code is
|
||
> modest.
|
||
|
||
### 5.6 `FENCE.I` and Self-Modifying Code
|
||
|
||
`FENCE.I` ensures that stores from the issuing hart that preceded the
|
||
`FENCE.I` are visible to subsequent instruction fetches on the same
|
||
hart. It does **not** order stores from other harts.
|
||
|
||
For XH-1, `FENCE.I` typically:
|
||
|
||
1. Drains the store buffer (in the W sense).
|
||
2. Invalidates the L1 I-cache (or the local portion of a unified L1).
|
||
3. Optionally issues a `FENCE` internally so that the store is
|
||
visible at the directory before the I-fetch resumes.
|
||
|
||
A cross-hart self-modifying-code sequence (hart A writes code, hart B
|
||
executes it) requires an explicit `FENCE` on hart A and a `FENCE.I` on
|
||
hart B (per the RISC-V recommendation), plus coherence to propagate
|
||
the new data. This is awkward and XH-1 will document it in the
|
||
programmer's manual.
|
||
|
||
### 5.7 `SFENCE.VMA`
|
||
|
||
`SFENCE.VMA` orders prior stores to PTEs before subsequent
|
||
address-translation results on the same hart. XH-1's TLB refill
|
||
engine must observe the store-completion signal; a typical
|
||
implementation stalls the pipeline at the `SFENCE.VMA` until the
|
||
store buffer drains.
|
||
|
||
---
|
||
|
||
## 6. System-Level Ordering Devices
|
||
|
||
Beyond per-hart mechanisms, XH-1 needs system-level ordering for
|
||
I/O and devices. The RISC-V privileged specification provides two
|
||
mechanisms:
|
||
|
||
### 6.1 I/O Ordering Bits in PTE
|
||
|
||
A PTE can have the **PBMT** field (Page-Based Memory Types), which can
|
||
be marked as:
|
||
|
||
- **Non-cacheable, idempotent, weakly-ordered (NC, IO).**
|
||
- **Non-cacheable, idempotent, strongly-ordered (NC, I/O-strong).**
|
||
|
||
XH-1 must honor these by:
|
||
|
||
- Routing accesses to the appropriate device port on the interconnect.
|
||
- For weakly-ordered I/O, treating the access as device but allowing
|
||
reordering with respect to other weakly-ordered I/O of the same type.
|
||
- For strongly-ordered I/O, treating the access as a full-fence —
|
||
the load or store completes in program order with respect to all
|
||
other accesses from the same hart.
|
||
|
||
### 6.2 Memory-Mapped I/O Fences
|
||
|
||
XH-1 will provide a `FENCE` variant (the standard `FENCE pred=rw,
|
||
succ=rw, fm=0`) that software uses to order I/O accesses. The
|
||
microarchitecture must ensure the drain endpoint includes the
|
||
device-port queue, not just the L1/L2.
|
||
|
||
### 6.3 Interrupt Acknowledgment and Ordering
|
||
|
||
RISC-V interrupt acknowledgment is itself a memory-mapped I/O
|
||
write. The microarchitecture must guarantee that the write to the
|
||
*Platform-Level Interrupt Controller* (PLIC) is ordered with respect
|
||
to the I/O operations it is acknowledging. The standard idiom is a
|
||
`FENCE` before the claim/complete operation. XH-1's PLIC integration
|
||
must be designed so that the `FENCE` does not require extra
|
||
microarchitectural support beyond the standard fence.
|
||
|
||
> **Open:** Whether XH-1 implements a **"fence fast path"** for
|
||
> accesses to specific device memory regions (e.g., the PLIC) that
|
||
> bypasses the directory and uses a dedicated sideband signal.
|
||
> Such optimizations can reduce interrupt latency by tens of cycles
|
||
> but complicate coherence reasoning. Not recommended for the first
|
||
> silicon.
|
||
|
||
---
|
||
|
||
## 7. Coherence Interaction
|
||
|
||
XH-1's coherence protocol (described in the coherence research
|
||
document) is directory-based, MSI or MESI, with a distributed
|
||
directory across the L2 slices. Memory ordering depends on the
|
||
coherence protocol's notion of "completion."
|
||
|
||
For a store to be considered "globally performed":
|
||
|
||
- The line is held in **Modified** state by the requesting core.
|
||
- All prior sharers have sent Invalidate Acknowledgments to the
|
||
directory (or the directory has confirmed there were no sharers).
|
||
- The directory has logged the new owner.
|
||
|
||
For a load to be considered "globally performed":
|
||
|
||
- The line is held in **Shared** (or **Modified** if exclusive) by
|
||
the requesting core.
|
||
- The data response from the directory (or forwarded from a peer) is
|
||
the value the directory currently authorizes.
|
||
|
||
> **Important subtlety:** A `FENCE W,W` between two stores on the
|
||
> same hart must ensure that store 1 has acquired its M grant *and*
|
||
> the directory has processed the invalidations *before* store 2
|
||
> acquires its M grant. In directory terms, store 2 cannot be issued
|
||
> until store 1's M grant has been observed at the directory. XH-1
|
||
> must implement this as an ordering dependency at the directory's
|
||
> L2 controller, not as a per-line serialization at the directory
|
||
> (which would be too coarse and too slow).
|
||
|
||
### 7.1 Coherence-Induced Ordering Pitfalls
|
||
|
||
Common pitfalls in directory-based designs:
|
||
|
||
- **Directory ack starvation under fence pressure:** when many cores
|
||
fence simultaneously, the directory must serialize invalidation
|
||
acknowledgments without deadlock.
|
||
- **Coherence downgrade during fence:** a core that downgrades M→S
|
||
while a fence is in progress must not let younger loads proceed
|
||
past the fence.
|
||
- **Silent upgrade (no-invalidate transition) on store:** a clean
|
||
line upgraded to M without invalidation must still be ordered
|
||
with respect to other cores' loads that observed the line in S.
|
||
The directory must explicitly synchronize this transition.
|
||
|
||
### 7.2 Ordering at the Interconnect
|
||
|
||
XH-1's interconnect is a mesh (proposed; see interconnect research
|
||
document). The mesh delivers messages in-order *between a given pair
|
||
of routers*, but globally the order is not guaranteed. The directory
|
||
controller is the serialization point.
|
||
|
||
For ordering, XH-1's mesh must:
|
||
|
||
- Not reorder different message classes on the same virtual channel.
|
||
- Support a "tail" or "credit" signal so that a fence-acknowledgment
|
||
message can be generated only when all prior messages on the same
|
||
path have been delivered.
|
||
|
||
---
|
||
|
||
## 8. Implementation Approaches Compared
|
||
|
||
This section lists the realistic implementation strategies for an
|
||
in-order XH-1 core. The choice of in-order vs. out-of-order pipeline
|
||
is discussed in §9.1; here we focus on ordering-specific mechanisms.
|
||
|
||
### 8.1 Approach A: Conservative Per-Core Fence (Baseline)
|
||
|
||
- The LSU drains the relevant queue classes fully on every `FENCE`.
|
||
- The store buffer's "drain complete" is defined as "every store has
|
||
received its directory M grant."
|
||
- A `FENCE` is implemented as a special instruction that stalls the
|
||
pipeline until drain completes.
|
||
- Acquire/release AMOs are implemented as full `FENCE`s around the
|
||
AMO at the LSU.
|
||
- `SFENCE.VMA` stalls the pipeline until the store buffer drains.
|
||
|
||
**Pros:**
|
||
|
||
- Simple to verify. The fence semantics maps directly to "queue
|
||
empty."
|
||
- Predictable performance for small core counts.
|
||
- Easy to reason about coherence interactions.
|
||
|
||
**Cons:**
|
||
|
||
- Poor scalability. At 128 cores, the worst-case `FENCE` latency
|
||
grows because the drain point is global, not local.
|
||
- Overly conservative for `aq`/`rl` AMOs, which do not require a
|
||
full global drain on either side — only ordering relative to
|
||
surrounding ordinary accesses.
|
||
|
||
### 8.2 Approach B: Token-Based Completion
|
||
|
||
- The LSU attaches a monotonically increasing **completion token**
|
||
to each outstanding operation.
|
||
- A `FENCE` with pred P, succ S becomes: "wait until all operations
|
||
of type P issued before this `FENCE` have produced a token greater
|
||
than or equal to the fence's token."
|
||
- The token is generated when the directory acknowledges the
|
||
operation; this is exactly the "global completion" notion.
|
||
- Acquire/Release is implemented as a fence of the appropriate
|
||
type around the AMO at the LSU.
|
||
|
||
**Pros:**
|
||
|
||
- Fences do not require the entire store buffer to drain; an empty
|
||
store buffer is sufficient, but so is a store buffer where all
|
||
older stores have completed.
|
||
- Naturally parallel: many stores can be outstanding, and the
|
||
fence only waits for the *oldest* of them.
|
||
- Slightly more complex but a well-understood technique (used in
|
||
academic and commercial cores).
|
||
|
||
**Cons:**
|
||
|
||
- Per-message token bookkeeping is more expensive than a simple
|
||
"queue empty" check.
|
||
- Verification requires modeling tokens at the testbench, which is
|
||
a non-trivial simulation infrastructure investment.
|
||
|
||
### 8.3 Approach C: Directory-Assisted Ordering
|
||
|
||
- A special "fence-ack" message class is added to the coherence
|
||
protocol.
|
||
- A `FENCE` instruction sends a fence-ack request to the directory;
|
||
the directory replies only after all prior messages from the
|
||
issuing core have been serialized.
|
||
- The directory tracks per-core "last-seen" sequence numbers to
|
||
arbitrate fence-ack requests against ongoing traffic.
|
||
|
||
**Pros:**
|
||
|
||
- The directory is the natural serialization point; leveraging it
|
||
avoids per-core reinvention of global ordering.
|
||
- Scales well to large core counts because the directory already
|
||
handles 128-core traffic.
|
||
- Allows the per-core LSU to continue issuing younger operations
|
||
while the fence is in progress, as long as they are ordered
|
||
correctly.
|
||
|
||
**Cons:**
|
||
|
||
- Adds protocol complexity (a new message class, directory state,
|
||
ack/grant logic).
|
||
- Risk of fence-ack starvation under heavy traffic.
|
||
- Requires careful design to avoid deadlock with normal coherence
|
||
traffic (classic message-class-dependency cycle).
|
||
|
||
### 8.4 Approach D: Speculative OoO + MOB
|
||
|
||
This is the out-of-order approach. The LSU becomes a Memory
|
||
Ordering Buffer (MOB) that tracks per-operation completion and
|
||
issues fences by tracking the "oldest in-flight" of each class.
|
||
This is how high-performance OoO cores implement fences; the
|
||
trade-off is pipeline depth and complexity.
|
||
|
||
XH-1 is proposed as **in-order** for power and predictability; see
|
||
§9.1. OoO is listed for completeness.
|
||
|
||
**Pros:**
|
||
|
||
- Highest performance on workloads with deep memory-level
|
||
parallelism.
|
||
- Fences can be "free" if no relevant older operation is in flight.
|
||
|
||
**Cons:**
|
||
|
||
- Significantly larger area and power.
|
||
- More verification effort, especially for the MOB's speculative
|
||
reordering.
|
||
- Diminishing returns on memory-bound workloads at 128 cores (the
|
||
bottleneck is the interconnect, not the per-core issue rate).
|
||
|
||
### 8.5 Summary of Approaches
|
||
|
||
| Approach | Perf (single) | Perf (128) | Area | Verif. effort | Risk |
|
||
| -------- | ------------- | ---------- | ---- | ------------- | ---- |
|
||
| A: Cons. | Low | Lowest | Lowest | Lowest | Low |
|
||
| B: Token | Medium | Medium | Low | Medium | Low |
|
||
| C: Dir-Assisted | Medium-high | High | Medium | High | Medium |
|
||
| D: OoO+MOB | High | High | High | Highest | High |
|
||
|
||
> **Proposed XH-1 implementation:** **Approach C (Directory-Assisted
|
||
> Ordering)** as the primary, with **Approach B (Token-Based
|
||
> Completion)** as the per-core implementation strategy. The
|
||
> directory handles global serialization, the LSU handles local
|
||
> token tracking. This combination is the most natural for a
|
||
> directory-coherent 128-core design.
|
||
|
||
---
|
||
|
||
## 9. Interactions with the Pipeline and Other Subsystems
|
||
|
||
### 9.1 Pipeline Choice (In-Order vs. Out-of-Order)
|
||
|
||
XH-1 is **[Proposed] in-order** for the first silicon. The reasons:
|
||
|
||
- Power envelope. 128 cores at high frequency is a thermal
|
||
challenge; an in-order core can run cooler.
|
||
- Predictable timing. Real-time and safety-critical workloads
|
||
benefit from deterministic `FENCE` and `FENCE.I` latencies.
|
||
- Verification cost. An OoO core's MOB is a major verification
|
||
burden; a 128-core OoO design is a multi-year project.
|
||
|
||
An in-order pipeline interacts with ordering as follows:
|
||
|
||
- The LSU is non-speculative on the store side (stores are
|
||
committed when the integer pipeline commits them) and lightly
|
||
speculative on the load side (loads may issue before older
|
||
stores, with a structural dependency check at commit).
|
||
- `FENCE` stalls the pipeline until drain; the cost is a few tens
|
||
to a few hundred cycles depending on directory state.
|
||
- Acquire/release AMOs generate local fence events in the LSU.
|
||
|
||
An OoO pipeline would:
|
||
|
||
- Issue stores from the STB to L1, with the MOB tracking
|
||
completion.
|
||
- Issue loads speculatively past stores with prediction; the MOB
|
||
validates at commit.
|
||
- `FENCE` becomes a "wait for oldest in-flight of pred class to
|
||
complete" check; usually very fast.
|
||
|
||
> **Open:** Whether a future XH-2 could be OoO. The research
|
||
> project is not yet at that question.
|
||
|
||
### 9.2 Cache Hierarchy
|
||
|
||
Memory ordering is shaped by the cache hierarchy:
|
||
|
||
- **L1 D-cache:** the load-side ordering reference. A load that hits
|
||
in L1 sees a value that is at least as recent as the cache's
|
||
coherence state indicates. A load that misses in L1 must go to
|
||
L2, where the directory arbitrates the order.
|
||
- **L2 slice (per tile):** home of the directory entry; serializes
|
||
ordering for the addresses it owns.
|
||
- **L3 (shared, distributed):** a victim cache and last-level
|
||
coherence point. XH-1's directory is the L2 directory; L3 is
|
||
treated as a backing store. Invalidation ordering for XH-1
|
||
therefore ends at L2.
|
||
- **L1 I-cache:** the `FENCE.I` target. The I-cache must be flushed
|
||
or invalidated to ensure stores are visible.
|
||
|
||
The cache hierarchy directly affects `FENCE` latency because the
|
||
"completion" point is the L2 directory. A `FENCE W,W` in a system
|
||
with L1-only ordering would be much cheaper (a few cycles) but
|
||
incorrect; the L2 directory must be in the loop.
|
||
|
||
### 9.3 Memory System
|
||
|
||
The memory system (DRAM controllers, on-chip memory, scratchpads)
|
||
must respect the same ordering. For XH-1, accesses to on-chip
|
||
scratchpad memory bypass the coherence protocol entirely; they
|
||
are ordered by the scratchpad controller.
|
||
|
||
> **Open:** Whether XH-1 has a per-tile scratchpad or a shared
|
||
> on-chip SRAM region. If shared, ordering between scratchpad and
|
||
cache-coherent regions is a research question.
|
||
|
||
### 9.4 Interconnect
|
||
|
||
The mesh interconnect has ordering properties (per-link in-order,
|
||
globally non-ordered) that determine what the directory and
|
||
endpoints must do to maintain memory ordering. The "virtual
|
||
channel" abstraction in the mesh is the unit of ordering; XH-1
|
||
**[Proposed]** uses:
|
||
|
||
- **VC0:** Coherence requests (in-order).
|
||
- **VC1:** Coherence responses (in-order).
|
||
- **VC2:** Data messages (in-order).
|
||
- **VC3:** Fence-ack and other ordering-sensitive messages
|
||
(in-order).
|
||
|
||
This separation avoids deadlock and simplifies ordering reasoning.
|
||
|
||
### 9.5 Interrupts
|
||
|
||
RISC-V interrupts are delivered to a hart based on the *global
|
||
interrupt-enable* bits and the PLIC. Ordering constraints:
|
||
|
||
- A hart that disables interrupts and then writes to a shared
|
||
variable does not need a `FENCE` before the interrupt-disable
|
||
write; the write to `mstatus` or `sstatus` is itself a CSR write
|
||
and is ordered architecturally.
|
||
- A hart that claims an interrupt at the PLIC and then reads a
|
||
device register must `FENCE` between the claim and the read to
|
||
ensure the claim is globally visible (the claim is a load from
|
||
the PLIC's claim register; the read is a load to the device).
|
||
- The interrupt-return path uses `mret`/`sret`, which is
|
||
architecturally ordered with respect to prior CSR writes.
|
||
- Trap delivery and `xRET` do not include implicit memory fences;
|
||
the OS must insert them.
|
||
|
||
XH-1's interrupt controller integration must ensure that the
|
||
interrupt-acknowledge signal is recognized only after the
|
||
relevant memory accesses are ordered.
|
||
|
||
### 9.6 Operating System
|
||
|
||
The OS uses `FENCE` heavily:
|
||
|
||
- **Spinlocks and mutexes:** the lock acquire is typically a
|
||
`LR`/`SC` sequence; the lock release is an `SC` or a `FENCE`
|
||
followed by a `SW` to the lock variable.
|
||
- **Page-table manipulation:** `SFENCE.VMA` after every PTE write.
|
||
- **I/O:** `FENCE` before/after device accesses (see §6).
|
||
- **Cross-core TLB shootdown:** an IPI sender writes the
|
||
invalidation request to a shared memory location, then sends an
|
||
IPI, and the receiver must see the request after the IPI — this
|
||
requires a `FENCE` on the sender side.
|
||
- **RCU grace periods:** RCU relies on the memory model's
|
||
guarantees; a weak model like RVWMO requires explicit
|
||
`FENCE` at grace-period boundaries.
|
||
- **Boot and SMP startup:** the boot hart must `FENCE` before
|
||
starting secondary harts to ensure the trampoline code and data
|
||
are visible.
|
||
|
||
XH-1's OS port (Linux, or a research kernel) must insert `FENCE`
|
||
at all these sites. The cost of a `FENCE` in a 128-core system
|
||
makes the boot-time and idle-time costs small in absolute terms
|
||
but potentially large in *relative* terms; an in-order core that
|
||
spends 200 cycles per `FENCE` will spend a noticeable fraction of
|
||
its time in fences if it idles on a spinlock.
|
||
|
||
> **Open:** Whether XH-1 ports a vanilla Linux or a research
|
||
> kernel. The memory-ordering analysis in this document assumes
|
||
> vanilla Linux semantics but applies in both cases.
|
||
|
||
### 9.7 Verification
|
||
|
||
Memory ordering is famously difficult to verify because:
|
||
|
||
- The bug is typically a very rare, non-deterministic interleaving.
|
||
- The legal reorderings are many; enumerating them is hard.
|
||
- The coherence protocol and the memory model interact in subtle
|
||
ways.
|
||
|
||
XH-1's verification strategy **[Proposed]**:
|
||
|
||
1. **Litmus tests** for the memory model. Run a corpus of litmus
|
||
tests (based on the RISC-V memory-model litmus test suite, plus
|
||
XH-1-specific tests) on RTL simulation. The tests must include
|
||
the published RVWMO tests and tests specific to XH-1's
|
||
coherence protocol.
|
||
2. **Coherence protocol model checking.** Use a formal model
|
||
(e.g., Murphi, TLA+) of the directory protocol; prove
|
||
invariant properties.
|
||
3. **Pipeline formal verification.** Verify the LSU's local
|
||
reordering rules with bounded model checking.
|
||
4. **System-level simulation.** Boot Linux on the RTL or a
|
||
high-level model; run stress workloads (e.g., `perf bench
|
||
futex`, RCU torture test).
|
||
5. **Stress with random delay injection.** Vary latency on each
|
||
interconnect path randomly to expose ordering bugs.
|
||
|
||
The published **herd7** tool from Cambridge and the **diy7**
|
||
litmus-test framework are the standard tools for RVWMO
|
||
verification. **Operational models** like Zam are more amenable
|
||
to formal reasoning than axiomatic models; XH-1 should
|
||
**[Proposed]** support both for verification.
|
||
|
||
> **Open:** Whether XH-1's verification flow includes formal
|
||
> proofs of the memory ordering contract end-to-end, or whether
|
||
> simulation-based testing is the primary line of defense. A
|
||
> research project may choose the latter; a commercial project
|
||
> should aim for the former.
|
||
|
||
### 9.8 Performance
|
||
|
||
Memory ordering has two performance axes:
|
||
|
||
1. **Steady-state throughput:** how fast can the system process
|
||
ordinary memory traffic. This is dominated by coherence and
|
||
cache behavior, not by ordering per se.
|
||
2. **Fence/atomic latency and throughput:** how fast is a
|
||
synchronization primitive. This is where ordering directly
|
||
impacts performance.
|
||
|
||
For a 128-core system, the second axis is the more important
|
||
contributor to scalability. A workload with frequent
|
||
synchronization (e.g., a parallel database, a many-core
|
||
machine-learning inference engine) is dominated by `FENCE` and
|
||
AMO latency.
|
||
|
||
> **Quantitative note (rough estimate):** Assuming a 1 GHz in-order
|
||
> core, a directory-coherent L2 slice with a hop latency of ~10 ns
|
||
> per mesh traversal, and an average of 3–5 hops per coherence
|
||
> transaction, a single uncontended `FENCE` on a hot line might
|
||
> take ~30–50 ns (30–50 cycles). A contended `FENCE` (multiple
|
||
> sharers) requiring invalidation acks could take 100–300 ns
|
||
> (100–300 cycles). On a 128-core workload that fences once per
|
||
> critical section, this is the dominant cost.
|
||
|
||
XH-1's performance evaluation **[Proposed]**:
|
||
|
||
- Measure `FENCE` latency distribution under varying contention
|
||
and across the mesh (corner vs. center tile).
|
||
- Measure AMO latency under contention.
|
||
- Measure the effect of `aq`/`rl` vs. explicit `FENCE` in the
|
||
OS.
|
||
- Compare against an ideal "no ordering cost" baseline.
|
||
|
||
---
|
||
|
||
## 10. Specific RISC-V Considerations
|
||
|
||
### 10.1 The `Ztso` Extension (Optional)
|
||
|
||
`Ztso` provides TSO (Total Store Order) at every hart. With
|
||
`Ztso`, the implementation guarantees that stores are not
|
||
reordered with respect to subsequent loads, simplifying the
|
||
software model.
|
||
|
||
> **Proposed XH-1 stance:** Do not advertise `Ztso` for the
|
||
> first silicon. The hardware cost of supporting TSO at every
|
||
> hart is a tighter store-load ordering in the LSU, and the
|
||
> benefit is only felt by software that is willing to be
|
||
> incompatible with weak-ordering implementations. The standard
|
||
> RISC-V Linux kernel is RVWMO-compatible and does not require
|
||
> `Ztso`. Re-evaluate if a major software target requires it.
|
||
|
||
### 10.2 The `Zam` Extension (If Adopted)
|
||
|
||
`Zam` is a proposed extension that provides an operational
|
||
memory model in addition to the axiomatic RVWMO. It does not
|
||
change the contract; it provides operational clarity. XH-1
|
||
**[Proposed]** will document its memory model in terms of `Zam`
|
||
where possible because operational models are easier to
|
||
implement and verify.
|
||
|
||
### 10.3 The PTE-Based Memory Type (PBMT)
|
||
|
||
XH-1's MMU must support the PBMT field and route accesses to
|
||
the appropriate ordering class. Specifically:
|
||
|
||
- **PMA = NC (non-cacheable):** accesses go to the device port;
|
||
load/store ordering is enforced by the device-port controller.
|
||
- **PMA = IO (weakly ordered I/O):** same as NC, but with
|
||
explicit reordering rules.
|
||
- **PMA = WEAK (cacheable but weakly ordered):** accesses are
|
||
cached but ordered weakly; useful for write-combining buffers.
|
||
- **PMA = STRONG (cacheable, strongly ordered):** same as
|
||
cacheable but with fence semantics on every access; rarely
|
||
used.
|
||
|
||
> **Open:** Whether XH-1 supports all PBMT values or a subset.
|
||
> A research project may support all four; a commercial project
|
||
> may support only NC and cacheable.
|
||
|
||
### 10.4 The `fence.i` and Instruction-Fetch Coherence
|
||
|
||
`FENCE.I` orders stores w.r.t. instruction fetches on the same
|
||
hart. Cross-hart, the software idiom is:
|
||
|
||
```c
|
||
// hart A:
|
||
write_code(addr, code);
|
||
__sync_synchronize(); // FENCE
|
||
send_ipi(hart_B);
|
||
|
||
// hart B:
|
||
handle_ipi();
|
||
__sync_synchronize(); // FENCE
|
||
__fence_i(); // FENCE.I
|
||
icache_flush();
|
||
jump_to(addr);
|
||
```
|
||
|
||
XH-1's I-cache is a VIPT cache; on an aliasing configuration
|
||
(small pages, large cache), the OS must use the recommended
|
||
sequence (or the larger page size). The architecture team will
|
||
document this.
|
||
|
||
### 10.5 TSO and RVWMO Differences for Software
|
||
|
||
Although the focus here is hardware, it is worth noting for
|
||
completeness:
|
||
|
||
- Under TSO, ordinary loads and stores have program order
|
||
semantics. Under RVWMO, they do not.
|
||
- A program that is correct under TSO may not be correct under
|
||
RVWMO without additional fences. The Linux kernel's
|
||
memory-model code (`tools/memory-model/`) contains the
|
||
required fences for RVWMO.
|
||
|
||
XH-1's port will use the RVWMO-correct fences; no special
|
||
support is required in hardware beyond honoring the standard
|
||
`FENCE`.
|
||
|
||
---
|
||
|
||
## 11. Unresolved Design Questions
|
||
|
||
The following questions remain open for the XH-1 architecture:
|
||
|
||
1. **In-order vs. out-of-order pipeline.** In-order is proposed,
|
||
but the decision is not final. An OoO pipeline changes the
|
||
MOB design and may improve `FENCE` latency.
|
||
2. **`Ztso` adoption.** Whether to support TSO at every hart.
|
||
3. **`Zam` adoption.** Whether to require the operational memory
|
||
model in the specification.
|
||
4. **Coherence protocol.** MSI vs. MESI vs. MOESI. The choice
|
||
affects the number of transitions and therefore the number
|
||
of coherence-induced ordering edges. (See the coherence
|
||
research document.)
|
||
5. **Interconnect topology.** Mesh vs. ring vs. hierarchical
|
||
bus+mesh. The choice affects `FENCE` latency and the
|
||
directory's role.
|
||
6. **PBMT support scope.** Which PTE memory types XH-1
|
||
implements.
|
||
7. **Per-line vs. per-word reservation.** Affects `LR`/`SC`
|
||
success rates.
|
||
8. **Fence-ack message class.** Whether to add a dedicated
|
||
fence-ack message (Approach C) or rely entirely on
|
||
per-core token tracking (Approach B).
|
||
9. **Verification strategy.** Simulation-based vs. formal;
|
||
end-to-end memory model proofs.
|
||
10. **OS port.** Linux vs. a research kernel; this affects
|
||
the set of `FENCE` sites the system must support.
|
||
|
||
---
|
||
|
||
## 12. Recommendations
|
||
|
||
Where evidence supports a recommendation, the following are
|
||
XH-1's proposed positions:
|
||
|
||
1. **Adopt RVWMO as the memory model.** This is the standard
|
||
RISC-V model and is the most widely supported by software.
|
||
2. **Implement an in-order pipeline with Approach C
|
||
(Directory-Assisted Ordering) and Approach B (Per-Core
|
||
Token Completion).** This combination scales to 128 cores
|
||
while keeping the per-core pipeline simple.
|
||
3. **Use a directory-based MESI coherence protocol at the
|
||
L2 slice.** The directory is the natural serialization
|
||
point for memory ordering. (See the coherence research
|
||
document for the full argument.)
|
||
4. **Add a dedicated fence-ack message class to the coherence
|
||
protocol.** This avoids overloading existing messages and
|
||
reduces verification ambiguity.
|
||
5. **Honour the standard `FENCE`, `FENCE.I`, `SFENCE.VMA`,
|
||
`LR`/`SC`, and AMO semantics.** No architectural deviations.
|
||
6. **Document the cross-hart self-modifying-code sequence in
|
||
the programmer's manual.** Use the standard fence pair.
|
||
7. **Implement the PTE-based memory types (PBMT) for at
|
||
least NC and cacheable.** Other types are deferred.
|
||
8. **Provide a memory-model litmus-test suite as part of
|
||
the verification deliverables.** This is essential for
|
||
catching ordering bugs.
|
||
9. **Measure `FENCE` and AMO latency across the mesh in
|
||
performance evaluation.** Corner tiles have different
|
||
latency from center tiles; the asymmetry is a real
|
||
performance concern.
|
||
|
||
---
|
||
|
||
## 13. References
|
||
|
||
The following sources informed this document. URLs and
|
||
publication details are given where known; the reader should
|
||
verify currency, as the RISC-V ecosystem is rapidly evolving.
|
||
|
||
- RISC-V International, *RISC-V Instruction Set Manual,
|
||
Volume I: Unprivileged ISA*. Appendix A, "RVWMO: The RISC-V
|
||
Weak Memory Ordering Model."
|
||
- RISC-V International, *RISC-V Instruction Set Manual,
|
||
Volume II: Privileged ISA*. Sections on fence and memory
|
||
types (PBMT).
|
||
- Pulte, C., et al. *"Zam: A Complete and Verifiable
|
||
Operational Memory Model for RISC-V."* (Reference to the
|
||
operational memory model paper.)
|
||
- Owens, S., Sarkar, S., Sewell, P. *"A Better x86 Memory
|
||
Model: x86-TSO."* (TSO reference; the TSO model is the
|
||
inspiration for `Ztso`.)
|
||
- Lamport, L. *"How to Make a Multiprocessor Computer That
|
||
Correctly Executes Multiprocess Programs."* IEEE TC, 1979.
|
||
(The foundational sequential consistency paper.)
|
||
- Adve, S. V., Gharachorloo, K. *"Shared Memory Consistency
|
||
Models: A Tutorial."* IEEE Computer, 1996. (Tutorial on
|
||
memory models.)
|
||
- Martin, M. M. K., et al. *"Token Coherence."* (Token-based
|
||
ordering reference, applicable to Approach B.)
|
||
- ACM SIGOPS Operating Systems Review. *Linux Kernel Memory
|
||
Model Documentation.* (`tools/memory-model/` in the Linux
|
||
kernel source tree.)
|
||
- AMD. *AMD64 Architecture Programmer's Manual*, Volume 2.
|
||
(TSO reference for comparison.)
|
||
- Intel. *Intel 64 and IA-32 Architectures Software
|
||
Developer's Manual*, Volume 3. (TSO reference for
|
||
comparison.)
|
||
- ARM. *ARM Architecture Reference Manual*. (Comparison
|
||
with ARM's weak memory model.)
|
||
|
||
Specific tool references:
|
||
|
||
- **herd7**: litmus-test framework for weak memory models,
|
||
by Cambridge. https://github.com/herd/herdtools7
|
||
- **diy7**: companion tool for herd7.
|
||
- **Murphi**, **TLA+**: protocol model checkers used for
|
||
coherence protocol verification.
|
||
|
||
The RISC-V Memory Model working group's published
|
||
documentation, including the operational model papers and
|
||
the litmus-test corpus, is the primary source for the
|
||
RVWMO contract. Specifics of XH-1's coherence protocol are
|
||
out of scope for this document and are addressed in the
|
||
coherence research document.
|
||
|
||
---
|
||
|
||
## 14. Document Status
|
||
|
||
This document is a **research proposal**. It does not represent
|
||
a finalized XH-1 architectural commitment. Items marked
|
||
**[Proposed]** are engineering recommendations under evaluation;
|
||
items marked **[Open]** are unresolved design questions.
|
||
|
||
The next revisions of this document should:
|
||
|
||
- Reflect decisions on the open questions in §11.
|
||
- Add quantitative `FENCE` latency measurements once RTL or
|
||
high-level models are available.
|
||
- Incorporate results from the verification litmus-test
|
||
suite.
|
||
- Align with the final coherence protocol choice.
|
||
- Document the OS-level `FENCE` insertion policy.
|
||
|
||
---
|
||
|
||
*End of document.*
|