mirror of
https://github.com/allexanderbergmns/xh1-research.git
synced 2026-08-26 16:47:01 +00:00
TEST: Completed Review #2 | research/05-memory/memory-ordering.md
This commit is contained in:
@@ -320,3 +320,112 @@
|
||||
[2026-08-26T10:34:46Z] Maximum research rounds reached.
|
||||
[2026-08-26T10:34:46Z] Leaving original document unchanged.
|
||||
[2026-08-26T10:34:46Z] Marking topic failed for this run.
|
||||
[2026-08-26T12:11:25Z] Started run: 20260826T121125Z
|
||||
[2026-08-26T12:11:25Z] ==================================================
|
||||
[2026-08-26T12:11:25Z] Researching: research/05-memory/memory-ordering.md
|
||||
[2026-08-26T12:11:25Z] ==================================================
|
||||
[2026-08-26T12:11:25Z] Research round 1/3
|
||||
[2026-08-26T12:11:25Z] Running researcher.
|
||||
[2026-08-26T12:11:25Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=1
|
||||
[2026-08-26T12:11:25Z] API request failed.
|
||||
[2026-08-26T12:11:26Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=2
|
||||
[2026-08-26T12:11:26Z] API request failed.
|
||||
[2026-08-26T12:11:58Z] Started run: 20260826T121158Z
|
||||
[2026-08-26T12:11:58Z] ==================================================
|
||||
[2026-08-26T12:11:58Z] Researching: research/05-memory/memory-ordering.md
|
||||
[2026-08-26T12:11:58Z] ==================================================
|
||||
[2026-08-26T12:11:58Z] Research round 1/3
|
||||
[2026-08-26T12:11:58Z] Running researcher.
|
||||
[2026-08-26T12:11:58Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=1
|
||||
[2026-08-26T12:12:47Z] Response received: 19485 bytes
|
||||
[2026-08-26T12:12:47Z] Running reviewer.
|
||||
[2026-08-26T12:12:47Z] API: provider=ollama model=gemma4:12b attempt=1
|
||||
[2026-08-26T12:15:45Z] Response received: 1622 bytes
|
||||
[2026-08-26T12:15:45Z] Research rejected by reviewer.
|
||||
[2026-08-26T12:15:45Z] Preparing revision round 2.
|
||||
[2026-08-26T12:15:50Z] Research round 2/3
|
||||
[2026-08-26T12:15:50Z] Running revision agent.
|
||||
[2026-08-26T12:15:50Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=1
|
||||
[2026-08-26T12:15:51Z] API request failed.
|
||||
[2026-08-26T12:15:52Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=2
|
||||
[2026-08-26T12:16:25Z] Response received: 23244 bytes
|
||||
[2026-08-26T12:16:25Z] Running reviewer.
|
||||
[2026-08-26T12:16:25Z] API: provider=ollama model=gemma4:12b attempt=1
|
||||
[2026-08-26T12:19:32Z] Response received: 1 bytes
|
||||
[2026-08-26T12:19:32Z] Research rejected by reviewer.
|
||||
[2026-08-26T12:19:32Z] Preparing revision round 3.
|
||||
[2026-08-26T12:19:37Z] Research round 3/3
|
||||
[2026-08-26T12:19:37Z] Running revision agent.
|
||||
[2026-08-26T12:19:37Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=1
|
||||
[2026-08-26T12:19:38Z] API request failed.
|
||||
[2026-08-26T12:19:39Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=2
|
||||
[2026-08-26T12:19:44Z] API request failed.
|
||||
[2026-08-26T12:19:48Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=3
|
||||
[2026-08-26T12:20:25Z] Response received: 25475 bytes
|
||||
[2026-08-26T12:20:25Z] Running reviewer.
|
||||
[2026-08-26T12:20:25Z] API: provider=ollama model=gemma4:12b attempt=1
|
||||
[2026-08-26T12:23:32Z] Response received: 1 bytes
|
||||
[2026-08-26T12:23:32Z] Research rejected by reviewer.
|
||||
[2026-08-26T12:23:32Z] Maximum research rounds reached.
|
||||
[2026-08-26T12:23:32Z] Leaving original document unchanged.
|
||||
[2026-08-26T12:23:32Z] Marking topic failed for this run.
|
||||
[2026-08-26T12:25:59Z] Started run: 20260826T122559Z
|
||||
[2026-08-26T12:25:59Z] ==================================================
|
||||
[2026-08-26T12:25:59Z] Researching: research/05-memory/memory-ordering.md
|
||||
[2026-08-26T12:25:59Z] ==================================================
|
||||
[2026-08-26T12:25:59Z] Research round 1/3
|
||||
[2026-08-26T12:25:59Z] Running researcher.
|
||||
[2026-08-26T12:25:59Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=1
|
||||
[2026-08-26T12:26:00Z] API HTTP status: 429
|
||||
[2026-08-26T12:26:00Z] WARNING: Rate limited by openrouter.
|
||||
[2026-08-26T12:26:00Z] Waiting 5s before retry.
|
||||
[2026-08-26T12:26:05Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=2
|
||||
[2026-08-26T12:26:06Z] API HTTP status: 429
|
||||
[2026-08-26T12:26:06Z] WARNING: Rate limited by openrouter.
|
||||
[2026-08-26T12:26:06Z] Waiting 20s before retry.
|
||||
[2026-08-26T12:26:26Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=3
|
||||
[2026-08-26T12:27:36Z] Response received: 23159 bytes
|
||||
[2026-08-26T12:27:36Z] Running reviewer.
|
||||
[2026-08-26T12:27:37Z] API: provider=ollama model=gemma4:12b attempt=1
|
||||
[2026-08-26T12:30:29Z] Response received: 4699 bytes
|
||||
[2026-08-26T12:30:29Z] WARNING: Reviewer returned no valid VERDICT.
|
||||
[2026-08-26T12:30:29Z] WARNING: This is a reviewer failure, not a research failure.
|
||||
[2026-08-26T12:33:03Z] Started run: 20260826T123303Z
|
||||
[2026-08-26T12:33:03Z] ==================================================
|
||||
[2026-08-26T12:33:03Z] Researching: research/05-memory/memory-ordering.md
|
||||
[2026-08-26T12:33:03Z] ==================================================
|
||||
[2026-08-26T12:33:03Z] Research round 1/3
|
||||
[2026-08-26T12:33:03Z] Running researcher.
|
||||
[2026-08-26T12:33:03Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=1
|
||||
[2026-08-26T12:33:04Z] API HTTP status: 429
|
||||
[2026-08-26T12:33:04Z] WARNING: Rate limited by openrouter.
|
||||
[2026-08-26T12:33:04Z] Waiting 5s before retry.
|
||||
[2026-08-26T12:33:09Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=2
|
||||
[2026-08-26T12:35:05Z] Response received: 41239 bytes
|
||||
[2026-08-26T12:35:05Z] Running reviewer.
|
||||
[2026-08-26T12:35:05Z] API: provider=ollama model=gemma4:12b attempt=1
|
||||
[2026-08-26T12:38:10Z] Response received: 2000 bytes
|
||||
[2026-08-26T12:38:10Z] Reviewer did not provide a machine-readable verdict.
|
||||
[2026-08-26T12:38:10Z] Requesting verdict-format repair.
|
||||
[2026-08-26T12:38:10Z] API: provider=ollama model=gemma4:12b attempt=1
|
||||
[2026-08-26T12:38:52Z] Response received: 87 bytes
|
||||
[2026-08-26T12:38:52Z] WARNING: Reviewer returned no valid VERDICT.
|
||||
[2026-08-26T12:38:52Z] WARNING: This is a reviewer failure, not a research failure.
|
||||
[2026-08-26T12:39:54Z] Started run: 20260826T123954Z
|
||||
[2026-08-26T12:39:54Z] ==================================================
|
||||
[2026-08-26T12:39:54Z] Researching: research/05-memory/memory-ordering.md
|
||||
[2026-08-26T12:39:54Z] ==================================================
|
||||
[2026-08-26T12:39:54Z] Research round 1/3
|
||||
[2026-08-26T12:39:54Z] Running researcher.
|
||||
[2026-08-26T12:39:54Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=1
|
||||
[2026-08-26T12:41:33Z] Response received: 38665 bytes
|
||||
[2026-08-26T12:41:33Z] Running reviewer.
|
||||
[2026-08-26T12:41:33Z] API: provider=ollama model=gemma4:12b attempt=1
|
||||
[2026-08-26T12:44:38Z] Response received: 2266 bytes
|
||||
[2026-08-26T12:44:38Z] Reviewer did not provide a machine-readable verdict.
|
||||
[2026-08-26T12:44:38Z] Requesting verdict-format repair.
|
||||
[2026-08-26T12:44:38Z] API: provider=ollama model=gemma4:12b attempt=1
|
||||
[2026-08-26T12:45:19Z] Response received: 87 bytes
|
||||
[2026-08-26T12:45:19Z] Recovered reviewer verdict: PASS
|
||||
[2026-08-26T12:45:19Z] Research passed review.
|
||||
[2026-08-26T12:45:19Z] Accepted research document: research/05-memory/memory-ordering.md
|
||||
|
||||
@@ -0,0 +1 @@
|
||||
{"error":{"message":"Provider returned error","code":429,"metadata":{"raw":"minimax/minimax-m3:free is temporarily rate-limited upstream. Please retry shortly, or add your own key to accumulate your rate limits: https://openrouter.ai/settings/integrations","provider_name":"GMICloud","is_byok":false,"provider_error_code":"rate_limit_exceeded","limit_source":"upstream_provider_shared_pool","remedy_hint":"Retry shortly, add your own provider key (https://openrouter.ai/settings/integrations), or route to another provider with provider routing: https://openrouter.ai/docs/features/provider-routing","retry_after_seconds":60,"retry_after_seconds_raw":60,"headers":{"Retry-After":"60"}}},"user_id":"user_2tRWhLFlPACV1q6VcCZVeaq8af4"}
|
||||
@@ -0,0 +1 @@
|
||||
{"error":{"message":"Provider returned error","code":429,"metadata":{"raw":"minimax/minimax-m3:free is temporarily rate-limited upstream. Please retry shortly, or add your own key to accumulate your rate limits: https://openrouter.ai/settings/integrations","provider_name":"GMICloud","is_byok":false,"provider_error_code":"rate_limit_exceeded","limit_source":"upstream_provider_shared_pool","remedy_hint":"Retry shortly, add your own provider key (https://openrouter.ai/settings/integrations), or route to another provider with provider routing: https://openrouter.ai/docs/features/provider-routing","retry_after_seconds":60,"retry_after_seconds_raw":60,"headers":{"Retry-After":"60"}}},"user_id":"user_2tRWhLFlPACV1q6VcCZVeaq8af4"}
|
||||
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
@@ -0,0 +1 @@
|
||||
{"error":{"message":"Provider returned error","code":429,"metadata":{"raw":"minimax/minimax-m3:free is temporarily rate-limited upstream. Please retry shortly, or add your own key to accumulate your rate limits: https://openrouter.ai/settings/integrations","provider_name":"GMICloud","is_byok":false,"provider_error_code":"rate_limit_exceeded","limit_source":"upstream_provider_shared_pool","remedy_hint":"Retry shortly, add your own provider key (https://openrouter.ai/settings/integrations), or route to another provider with provider routing: https://openrouter.ai/docs/features/provider-routing","retry_after_seconds":60,"retry_after_seconds_raw":60,"headers":{"Retry-After":"60"}}},"user_id":"user_2tRWhLFlPACV1q6VcCZVeaq8af4"}
|
||||
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
@@ -0,0 +1 @@
|
||||
{"id":"chatcmpl-948","object":"chat.completion","created":1787747932,"model":"gemma4:12b","system_fingerprint":"fp_ollama","choices":[{"index":0,"message":{"role":"assistant","content":"{\n \"verdict\": \"PASS\",\n \"confidence\": \"HIGH\",\n \"issues\": [],\n \"required_fixes\": []\n}","reasoning":"* Input: A technical review of an XH-1 RISC-V CPU research document.\n * Task: Determine if the reviewer intended \"PASS\" or \"FAIL\".\n * Output Format: Strict JSON (no markdown, no explanation, no code fences).\n * JSON Structure:\n * PASS: `{\"verdict\": \"PASS\", \"confidence\": \"HIGH\", \"issues\": [], \"required_fixes\": []}`\n * FAIL: `{\"verdict\": \"FAIL\", \"confidence\": \"HIGH\", \"issues\": [\"...\"], \"required_fixes\": [\"...\"]}`\n\n * \"The document is a well-structured and technically sound research proposal...\" (Positive)\n * \"It correctly identifies the primary challenges...\" (Positive)\n * \"The proposal aligns with the RISC-V ecosystem...\" (Positive)\n * \"demonstrates a sophisticated understanding...\" (Positive)\n * \"The choice of a directory-based MESI protocol... is the industry standard...\" (Positive)\n * \"The recommendation to use a mesh topology is also appropriate...\" (Positive)\n * \"The inclusion of herd7, diy7, and formal tools... indicates a rigorous approach...\" (Positive)\n * \"The document correctly references the relevant parts of the RISC-V ISA manuals...\" (Positive)\n\n * The reviewer uses exclusively positive language.\n * There are no criticisms, \"but\" clauses, or \"needs improvement\" sections.\n * The verdict is clearly a \"PASS\".\n\n * Verdict: PASS\n * Confidence: HIGH\n * Issues: []\n * Required Fixes: []\n\n * No Markdown? Yes.\n * No Explanation? Yes.\n * No Code Fences? Yes.\n * Valid JSON? Yes."},"finish_reason":"stop"}],"usage":{"prompt_tokens":635,"completion_tokens":445,"total_tokens":1080}}
|
||||
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
@@ -0,0 +1 @@
|
||||
{"id":"chatcmpl-997","object":"chat.completion","created":1787748319,"model":"gemma4:12b","system_fingerprint":"fp_ollama","choices":[{"index":0,"message":{"role":"assistant","content":"{\n \"verdict\": \"PASS\",\n \"confidence\": \"HIGH\",\n \"issues\": [],\n \"required_fixes\": []\n}","reasoning":"* Input: A technical review of an XH-1 RISC-V CPU research document.\n * Task: Determine if the reviewer intended \"PASS\" or \"FAIL\".\n * Output Format: Strict JSON (no markdown, no explanation, no code fences).\n * JSON Structure:\n * PASS: `{\"verdict\": \"PASS\", \"confidence\": \"HIGH\", \"issues\": [], \"required_fixes\": []}`\n * FAIL: `{\"verdict\": \"FAIL\", \"confidence\": \"HIGH\", \"issues\": [\"...\"], \"required_fixes\": [\"...\"]}`\n\n * \"The document is technically sound and well-structured...\" (Positive)\n * \"The document correctly identifies...\" (Positive)\n * \"This demonstrates a high level of adherence...\" (Positive)\n * \"The choice of... is the industry standard.\" (Positive)\n * \"The inclusion of... is essential...\" (Positive)\n * \"The 'Unresolved' section shows a mature understanding...\" (Positive)\n * \"This is a valid concern...\" (Positive)\n * There are no \"Critical\" or \"Major\" issues mentioned. The \"Unresolved\" section is praised for showing maturity, not as a failure.\n\n * The reviewer is very positive.\n * Verdict: PASS.\n * Confidence: HIGH.\n * Issues: [].\n * Required Fixes: [].\n\n * No markdown.\n * No code fences.\n * No explanation.\n\n ```json\n {\n \"verdict\": \"PASS\",\n \"confidence\": \"HIGH\",\n \"issues\": [],\n \"required_fixes\": []\n }\n ```"},"finish_reason":"stop"}],"usage":{"prompt_tokens":730,"completion_tokens":427,"total_tokens":1157}}
|
||||
@@ -0,0 +1,6 @@
|
||||
2026-08-26T12:12:47Z research/05-memory/memory-ordering.md 1 research completed
|
||||
2026-08-26T12:15:45Z research/05-memory/memory-ordering.md 1 review
|
||||
2026-08-26T12:16:25Z research/05-memory/memory-ordering.md 2 revision completed
|
||||
2026-08-26T12:19:32Z research/05-memory/memory-ordering.md 2 review
|
||||
2026-08-26T12:20:25Z research/05-memory/memory-ordering.md 3 revision completed
|
||||
2026-08-26T12:23:32Z research/05-memory/memory-ordering.md 3 review
|
||||
@@ -0,0 +1 @@
|
||||
research/05-memory/memory-ordering.md
|
||||
@@ -0,0 +1,333 @@
|
||||
# Memory Ordering
|
||||
|
||||
## Status
|
||||
|
||||
Stub document. XH-1 memory ordering model has not yet been defined. The repository contains no source artifacts, RTL, ISA extension proposals, or reference implementations that establish a memory consistency model for XH-1. All quantitative or XH-1-specific claims in this document are flagged INSUFFICIENT EVIDENCE where applicable.
|
||||
|
||||
## Abstract
|
||||
|
||||
Memory ordering defines the rules by which loads and stores issued by one or more cores become visible to other cores in the system. For a 128-core RISC-V processor, the choice of memory model — and the hardware mechanisms that enforce it — is one of the most consequential architectural decisions, affecting correctness, performance, area, power, scalability, and software portability. This document surveys the design space, identifies the candidate models available within the RISC-V ISA framework, examines alternative hardware enforcement strategies, and frames the open questions for XH-1. No final recommendation is made because the repository does not establish pipeline style, coherency protocol, or interconnect topology, all of which constrain the memory ordering decision. A conditional tentative default is stated, contingent on the assumed constraints being correct.
|
||||
|
||||
## Research Question
|
||||
|
||||
What memory consistency model and hardware ordering mechanism should XH-1 adopt, given:
|
||||
|
||||
1. A 128-core replication target
|
||||
2. The RISC-V ISA baseline
|
||||
3. The need to balance software portability against hardware implementation cost
|
||||
4. Unknown microarchitecture style (in-order vs. out-of-order is not established in the repository)
|
||||
5. The relationship to the cache coherency protocol (defined in `cache-coherency.md`, currently SOON)
|
||||
|
||||
## Background
|
||||
|
||||
### Memory Consistency Models — Terminology
|
||||
|
||||
PROPOSAL — Use the following working taxonomy in this document:
|
||||
|
||||
- **Sequential Consistency (SC)**: All cores observe a single total order of memory operations that respects each core's program order.
|
||||
- **Total Store Order (TSO)**: Loads may bypass stores from the same core; stores from one core are seen in order by all cores. Used by x86.
|
||||
- **Relaxed / Weak Memory Models**: Many orderings are allowed; fences (barriers) are required to enforce constraints. Used by ARM, RISC-V, and PowerPC.
|
||||
- **Release Consistency (RC) / RCsc / RCpc / RCO**: A family of models with different fence semantics. Used by academic designs and the basis for much of the C11/C++ memory model formalization.
|
||||
|
||||
ASSUMPTION — This document uses "RC variants" as an umbrella term for RCsc, RCpc, and RCO. The specific semantics of each variant are not redefined here; readers are referred to the literature.
|
||||
|
||||
### The RISC-V "RVWMO" Model
|
||||
|
||||
FACT — The RISC-V ISA specification defines the **RISC-V Weak Memory Ordering (RVWMO)** model as the baseline. Under RVWMO:
|
||||
|
||||
- Loads and stores to the same address from the same hart are ordered.
|
||||
- A store followed by a load to a *different* address may be reordered.
|
||||
- The `FENCE` instruction is used to enforce orderings. The base FENCE provides `pred` (predecessor) and `succ` (successor) bits covering device I/O (I), memory reads (R), and memory writes (W). Device ordering follows a stronger "Preserved over Higher Privilege and PMA" model (PostModel).
|
||||
- The optional `Ztso` extension provides **RVTSO**, which gives x86-like TSO semantics with a relaxed form of the FENCE.
|
||||
|
||||
FACT — RVWMO and the `Ztso` extension are documented in the ratified RISC-V Unprivileged ISA specification. The repository does not state which revision of the specification is the reference for XH-1; this is an open question.
|
||||
|
||||
ASSUMPTION — RVWMO + the Ztso optional extension is the available ISA-level model space that XH-1 must select from. (The repository does not document any custom ISA extensions, so XH-1 is assumed to stay within ratified RISC-V.)
|
||||
|
||||
### Why Memory Ordering is Hard at 128 Cores
|
||||
|
||||
OPEN QUESTION — The repository does not establish:
|
||||
- Whether XH-1 is a single die, multi-die package, or chiplet-based system
|
||||
- The interconnect topology (mesh, ring, crossbar, NoC, etc.)
|
||||
- Whether coherence is directory-based, broadcast-based, or hybrid
|
||||
- The private-vs-shared L2/L3 split
|
||||
- The memory-side directory or snoop filter organization
|
||||
|
||||
These decisions have direct bearing on which ordering mechanisms are feasible. Any hardware ordering scheme (invalidation queues, store buffers, fence acknowledgments, coherence-based ordering) must be analyzed in the context of these unknowns.
|
||||
|
||||
## Existing Approaches
|
||||
|
||||
### Approach A — RISC-V RVWMO (Baseline)
|
||||
|
||||
The standard weak memory model provided by the RISC-V ISA. Requires implementation of `FENCE`, `FENCE.I`, and the I/O ordering semantics.
|
||||
|
||||
**Mechanism**: Architecture permits aggressive reordering. Hardware implementations commonly use:
|
||||
- Store buffers with forwarding
|
||||
- Load-load reordering queues
|
||||
- Coherence-based ordering: stores become visible at coherence invalidation/upgrade points
|
||||
- Fence implementation via global acknowledgment counters or by draining reorder queues
|
||||
|
||||
INSUFFICIENT EVIDENCE — Which of these mechanisms XH-1 will use is not established by the repository.
|
||||
|
||||
### Approach B — RVWMO + Ztso Extension
|
||||
|
||||
Adds RVTSO. Software that requires x86-like ordering can avoid fences in many cases. Still RISC-V compliant.
|
||||
|
||||
**Mechanism**: Stores from one hart are serialized in the interconnect; loads may still pass earlier stores to different addresses.
|
||||
|
||||
### Approach C — Sequential Consistency (SC)
|
||||
|
||||
All operations observed in a single global order respecting program order. Simplest software model; most expensive hardware in general.
|
||||
|
||||
**Mechanism**: Commonly implemented via store-buffer draining on loads, broadcast-based coherence, and conservative replay mechanisms. The qualitative characterization of "very costly" is a general literature claim, not an XH-1 measurement; whether SC is specifically costly at 128 cores is workload- and topology-dependent and is treated here as a hypothesis, not a known result.
|
||||
|
||||
### Approach D — Release Consistency variants (RCsc, RCpc, RCO)
|
||||
|
||||
Academic models where release/acquire ordering is cheap but stronger fences are required for full memory ordering. Basis of C11/C++ model.
|
||||
|
||||
**Mechanism**: Only release/acquire fences require acknowledgment; acquire-only loads are cheap. Alignment with C11 is a strong software benefit. These variants are not part of the ratified RISC-V ISA, so adoption would require a custom ISA extension.
|
||||
|
||||
## Alternative Designs
|
||||
|
||||
### Hardware Ordering Mechanisms
|
||||
|
||||
PROPOSAL — The following mechanisms should be considered for fence/ordering enforcement once coherency and interconnect are known:
|
||||
|
||||
1. **Invalidation-acknowledgment fences** — A fence blocks until all outstanding coherence requests have been acknowledged. Simple and common; scaling impact depends on directory hop count, aggregate invalidation traffic, and invalidation filter structure.
|
||||
|
||||
2. **Per-core store-buffer drain fences** — A fence stalls the issuing core until its store buffer is empty and all stores have been globally observed. Predictable latency; serializes through the coherence fabric.
|
||||
|
||||
3. **Tournament / Token-based fences** — Fences acquire a global token or wait for a "fence epoch" counter. Avoids worst-case drain, but adds global state.
|
||||
|
||||
4. **FIFO coherence ordering** — Leverages the in-order completion of coherence transactions to provide ordering without explicit fences. Requires in-order interconnect, which conflicts with typical latency-optimized NoC designs.
|
||||
|
||||
5. **Time-to-live / epoch schemes** — Each coherence request carries an epoch tag; ordering is enforced at the L2/L3 directory by queueing requests by epoch.
|
||||
|
||||
6. **Speculative load reordering with rollback** — Loads can execute speculatively past stores; on conflict, the load is replayed. Common in high-performance out-of-order cores (e.g., reported in the IBM POWER literature).
|
||||
|
||||
ASSUMPTION — XH-1's microarchitecture style is not yet established. If out-of-order, the "speculative load reordering with rollback" category becomes a candidate. If in-order, mechanism choices narrow considerably.
|
||||
|
||||
## Comparison
|
||||
|
||||
| Property | SC | RVWMO | RVWMO + Ztso | RC variants |
|
||||
|----------|----|----|----|----|
|
||||
| Software ease | Highest | Medium | Medium-High | Low-Medium |
|
||||
| Hardware cost (general) | Highest | Low | Low-Medium | Low |
|
||||
| Performance headroom (general) | Lowest | Highest | High | High |
|
||||
| Fence latency (general) | N/A | Variable | Variable | Variable |
|
||||
| Verification cost (general) | Lowest in principle | High | High | High |
|
||||
| C11/C++ alignment | Approximate | Approximate | Weak | Direct |
|
||||
| 128-core scalability | Poor (hypothesis) | Unmeasured | Unmeasured | Unmeasured |
|
||||
|
||||
INSUFFICIENT EVIDENCE — Quantitative fence-latency, coherence-traffic, and scalability numbers for XH-1 are unavailable. Several rows above contain qualitative judgments rather than measured results, and the "128-core scalability" and "verification cost" rows reflect general literature consensus, not XH-1 measurements.
|
||||
|
||||
## Advantages
|
||||
|
||||
### RVWMO (Baseline)
|
||||
- **Lowest hardware cost** of all candidates for the base case, as a general property of the model — most orderings are simply not enforced.
|
||||
- **Maximum performance** potential, as a general property — reordering is unconstrained.
|
||||
- **Familiar** to the RISC-V software ecosystem (Linux, RISC-V GCC, LLVM), subject to the open question of whether Linux is the target OS.
|
||||
- **Standardized** — concrete compliance test suites exist externally.
|
||||
|
||||
### RVWMO + Ztso
|
||||
- **x86-like TSO** for software that benefits from stronger ordering.
|
||||
- **Backwards compatible** — software not using Ztso still runs correctly under RVWMO.
|
||||
- **Modest hardware cost** — primarily requires store-store ordering at the coherence layer (qualitative claim, not measured for XH-1).
|
||||
|
||||
### Release Consistency Variants
|
||||
- **Excellent fit for C11/C++** atomics — acquire/release are the most-used fence types in real code, per the C/C++ standards.
|
||||
- **Cheapest common case** — acquire and release can be implemented with lightweight mechanisms, as a general property.
|
||||
- **Hardware cost is concentrated on the rare case** of full seq_cst fences, as a general property.
|
||||
|
||||
## Disadvantages
|
||||
|
||||
### RVWMO (Baseline)
|
||||
- **Fence latency depends on coherence round-trip** — at 128 cores, fence latency is expected to be high because global acknowledgment must traverse the coherence fabric. Specific cycle counts: INSUFFICIENT EVIDENCE.
|
||||
- **Verification complexity is high** — many legal reorderings; the litmus-test failure surface is large. This is a general property of weak models, not an XH-1 measurement.
|
||||
- **Software burden** — kernel and runtime code must insert fences correctly.
|
||||
|
||||
### RVWMO + Ztso
|
||||
- **Two consistency models in one chip** — adds documentation, validation, and software education cost.
|
||||
- **In-flight mixing** of ordering regimes in the same software is hard to reason about.
|
||||
|
||||
### Release Consistency Variants
|
||||
- **Non-standard for RISC-V** — would require a custom ISA extension. Cannot be recommended without a strong software-side driver.
|
||||
- **Verification cost is high** in general because of subtle fence semantics; whether it is the "highest" among candidates is a claim about the literature, not a measured XH-1 result.
|
||||
|
||||
### Sequential Consistency
|
||||
- **Performance cost is severe in the general case** — every load may stall on store-buffer drain; whether this is "severe at 128 cores" specifically is workload- and topology-dependent and is not established by the repository.
|
||||
- **No selective escape valve** — software cannot opt out of the strongest model.
|
||||
- **Does not match RISC-V ecosystem expectations** — surprising to RISC-V software developers, as a general observation.
|
||||
|
||||
## XH-1 Considerations
|
||||
|
||||
OPEN QUESTION — The following XH-1-internal questions are unresolved by the repository:
|
||||
|
||||
1. Is XH-1's pipeline in-order or out-of-order?
|
||||
2. What is the coherence protocol (MOESI, MESI, directory-based, broadcast)?
|
||||
3. Is the interconnect a NoC, ring, or crossbar? Does it preserve in-order delivery of coherence responses?
|
||||
4. What is the expected working set of a typical hart — does streaming-store optimization matter?
|
||||
5. Does XH-1 target HPC, server, embedded, or mixed workloads? (Workload affects fence frequency.)
|
||||
6. Is XH-1 a research vehicle (where SC is acceptable for simplicity) or a product (where RVWMO compliance is mandatory)?
|
||||
7. Does the memory map use non-coherent regions (e.g., DMA, I/O) that interact with the ordering model?
|
||||
|
||||
ASSUMPTION — If XH-1 targets RISC-V ecosystem compatibility (Linux, RISC-V GCC, RISC-V LLVM), then RVWMO is mandatory. Ztso may optionally be added for x86-software porting convenience.
|
||||
|
||||
## 128-Core Scaling Considerations
|
||||
|
||||
INSUFFICIENT EVIDENCE — Quantitative figures for XH-1 at 128 cores are unavailable. The qualitative observations below are hypotheses grounded in general literature, not XH-1 measurements.
|
||||
|
||||
### Fence Latency
|
||||
- **Fence global acknowledgment** latency is bounded below by the worst-case coherence round-trip time. In a 128-core system, this can plausibly reach many tens of cycles in a NoC and higher in a ring or multi-hop topology; specific XH-1 figures: INSUFFICIENT EVIDENCE.
|
||||
- **Store-buffer drain** latency is approximately proportional to the store buffer depth and the time to invalidate all sharers. Specific XH-1 figures: INSUFFICIENT EVIDENCE.
|
||||
- **Token / epoch fences** may scale better because they avoid draining the entire coherence fabric — they synchronize at a logical epoch boundary. This is a general property, not an XH-1 measurement.
|
||||
|
||||
### Coherence Traffic
|
||||
- **RVWMO fences** generate invalidation or acknowledgment traffic that competes with normal coherence traffic. The magnitude depends on directory organization, snoop-filter effectiveness, and share-set size.
|
||||
- **SC at 128 cores** can plausibly saturate a directory-based interconnect because every load's coherence transaction may stall pending store ordering; this is a hypothesis, not a measured XH-1 result.
|
||||
- **RC variants** concentrate traffic on the fence path only, as a general property.
|
||||
|
||||
### Verification Scalability
|
||||
- **State-space explosion**: at 128 cores, the number of concurrently observable memory operations grows combinatorially in the unmitigated model. Memory-ordering verification is a known scaling bottleneck in the literature, and the repository does not establish what reduction or abstraction techniques XH-1 will employ.
|
||||
- **RC variants** are commonly cited as among the hardest to verify because of subtle fence interactions; whether they are hardest in the absolute sense is a literature claim, not an XH-1 measurement.
|
||||
- **SC is easiest to verify in principle** because it admits fewer legal behaviors, but worst in performance.
|
||||
|
||||
INSUFFICIENT EVIDENCE — No quantitative fence-latency, coherence-traffic, or verification-scaling data is available for XH-1. Estimates cannot be made without knowing the interconnect, coherence protocol, and verification methodology.
|
||||
|
||||
## Performance Considerations
|
||||
|
||||
PROPOSAL — Performance-impact dimensions to evaluate once workload and microarchitecture are defined:
|
||||
|
||||
1. **Steady-state IPC** for memory-bound workloads — qualitatively, RVWMO permits the most reordering and SC the least; XH-1-specific IPC: INSUFFICIENT EVIDENCE.
|
||||
2. **Fence latency** — measured as cycles from fence issue to fence completion. Dominant for synchronization-heavy code. XH-1-specific: INSUFFICIENT EVIDENCE.
|
||||
3. **Scalable synchronization** — how the model interacts with lock implementations, RCU, and atomic primitives.
|
||||
4. **Load-use latency** — whether loads can bypass earlier stores (allowed in RVWMO and TSO, not in SC).
|
||||
5. **Coherence miss rate** interaction — does the ordering model inflate or deflate cache-line ping-pong? Model-dependent; XH-1-specific: INSUFFICIENT EVIDENCE.
|
||||
|
||||
ASSUMPTION — For HPC workloads with high-fence frequency (e.g., MPI, OpenMP, lock-heavy code), even small differences in fence latency will dominate. For server workloads with mixed locking, the cost of fences is amortized over longer critical sections. No XH-1 workload profile is documented in the repository.
|
||||
|
||||
## Area Considerations
|
||||
|
||||
INSUFFICIENT EVIDENCE — All area descriptions in this section are qualitative. Precise XH-1 area figures (in gates, µm², or mm²) are unavailable; no synthesis, layout, or RTL data exists in the repository.
|
||||
|
||||
PROPOSAL — Qualitative area cost of ordering mechanisms, as a function of microarchitectural choices:
|
||||
|
||||
| Mechanism | Approximate area cost per core (qualitative) |
|
||||
|-----------|----------------------------------------------|
|
||||
| Store buffer with forwarding | Small (few entries × cache-line width) |
|
||||
| Load-load reordering queue | Small to medium (grows with IPC and reorder depth) |
|
||||
| Fence-acknowledgment counter | Negligible (per-core counter) |
|
||||
| Global epoch / token state | Medium (one per chip, not per core) |
|
||||
| Speculative load reorder + rollback | Large (load-load queue, replay logic) |
|
||||
|
||||
PROPOSAL — At 128 cores, per-core area costs are replicated 128 times. Global state (epoch counters, fence acknowledgers) is O(1) per chip and therefore amortizes favorably on a per-core basis, but the critical-path and clock-distribution impact of global state is not established by the repository.
|
||||
|
||||
## Power and Energy Considerations
|
||||
|
||||
INSUFFICIENT EVIDENCE — No per-fence, per-core, or aggregate energy figures for XH-1 are available. The earlier version of this document included a "100 pJ-per-core fence" figure that was not substantiated by any repository artifact or external citation; that figure has been removed.
|
||||
|
||||
PROPOSAL — Qualitative energy implications, as general properties of the model classes:
|
||||
|
||||
- **Fences** generate global coherence traffic — energy cost scales with the number of cores that must acknowledge. The aggregate chip-level energy per fence is therefore expected to grow with core count. Quantitative XH-1 values: INSUFFICIENT EVIDENCE.
|
||||
- **Speculative load reordering** wastes energy on rollback when a violation is detected, as a general property.
|
||||
- **Strong ordering models** keep more coherence state in flight (more invalidations, more retries), as a general property.
|
||||
- **RVWMO** with lightweight fences (when software permits) is the lowest-energy regime in the general case.
|
||||
|
||||
## Implementation Considerations
|
||||
|
||||
PROPOSAL — Implementation order of dependencies (the repository does not establish these, but the research area suggests the following):
|
||||
|
||||
1. Define coherency protocol → `cache-coherency.md`
|
||||
2. Define interconnect topology → not yet in repository
|
||||
3. Define memory hierarchy → `memory-hierarchy.md`
|
||||
4. Define memory ordering model → this document
|
||||
5. Define atomic primitives → `atomics.md`
|
||||
|
||||
ASSUMPTION — The Ztso extension is a small RTL delta over RVWMO. The dominant implementation cost is verification, not area. No RTL exists in the repository to substantiate even a relative area claim, so this is treated as a plausible qualitative statement rather than a measured result.
|
||||
|
||||
OPEN QUESTION — Does XH-1 have a coherent accelerator fabric or non-coherent IO (e.g., CXL, DMA engines)? If so, the memory ordering model must define how these agents interact, and the repository does not currently do so.
|
||||
|
||||
## Verification Considerations
|
||||
|
||||
PROPOSAL — Verification challenges:
|
||||
|
||||
- **Litmus tests**: A standard RISC-V litmus test suite is available externally; whether XH-1 will adopt it as-is or define a custom suite is OPEN.
|
||||
- **State-space explosion**: 128 cores × N outstanding operations produces a combinatorial unmitigated state space. The repository does not establish whether XH-1 will use state-space reduction, abstraction, bisimulation, or formal methods. Any specific XH-1 claim: INSUFFICIENT EVIDENCE.
|
||||
- **Fence semantics**: Each fence variant must be exhaustively tested at corner cases (interrupts, exceptions, MMIO).
|
||||
- **Coherence-ordering interaction**: The ordering model must be checked against the chosen coherence protocol for consistency. This check cannot be performed until the coherence protocol is defined.
|
||||
|
||||
INSUFFICIENT EVIDENCE — No XH-1 verification infrastructure, formal spec, or litmus-test set is documented in the repository.
|
||||
|
||||
ASSUMPTION — RVWMO is the most documented and tested model for RISC-V, per the ratified specification and ecosystem practice. Ztso is a small addition. SC and RC variants would require custom formal infrastructure. These are qualitative claims about ecosystem maturity, not XH-1-specific measurements.
|
||||
|
||||
## Software Considerations
|
||||
|
||||
PROPOSAL — Software-side impact, as a general property of each model:
|
||||
|
||||
- **Linux kernel**: Designed for weak memory models with appropriate fences. Compatible with RVWMO.
|
||||
- **C11/C++ atomics**: Map cleanly onto release/acquire semantics. Compatible with RVWMO and with RC variants.
|
||||
- **x86 software ports**: Benefit from Ztso for fewer fences.
|
||||
- **OpenSHMEM, MPI shmem**: Often assume TSO-like ordering; Ztso helps.
|
||||
- **HPC codes with hand-rolled atomics**: Highly sensitive to fence latency.
|
||||
|
||||
ASSUMPTION — XH-1 will run RISC-V Linux. The kernel's memory model expectations must be satisfied by XH-1's hardware ordering rules. The repository does not document whether Linux is in fact the target.
|
||||
|
||||
## Recommendation
|
||||
|
||||
INSUFFICIENT EVIDENCE to make a final recommendation.
|
||||
|
||||
The repository does not yet establish:
|
||||
- Pipeline microarchitecture
|
||||
- Coherence protocol
|
||||
- Interconnect topology
|
||||
- Workload target
|
||||
|
||||
A recommendation can be made only after the following documents in the memory research area are filled in: `memory-architecture.md`, `cache-coherency.md`, `memory-hierarchy.md`. Without these, any model selection is premature.
|
||||
|
||||
PROPOSAL — Tentative default, contingent on the stated assumptions holding and explicitly subject to revision once repository context is available:
|
||||
|
||||
1. **Baseline**: RVWMO (assumed mandatory for RISC-V ecosystem compatibility, if Linux is the target OS).
|
||||
2. **Optional add-on**: Ztso extension for x86-software porting convenience and modest hardware cost, contingent on verification cost being acceptable.
|
||||
3. **Avoid**: SC, custom RC variants, unless a strong workload or research driver emerges.
|
||||
4. **Defer to**: `cache-coherency.md` for the coherence-side ordering mechanism.
|
||||
|
||||
ASSUMPTION — The tentative default above depends on (a) Linux being the target OS, (b) ecosystem compatibility being prioritized over x86-port convenience, and (c) the coherence and interconnect unknowns being resolvable without invalidating this default. None of these are established by the repository.
|
||||
|
||||
## Confidence
|
||||
|
||||
| Topic | Confidence | Reason |
|
||||
|-------|------------|--------|
|
||||
| RVWMO is the RISC-V baseline | High | Ratified ISA specification |
|
||||
| Ztso is a small optional extension | High | Ratified extension in the ISA specification |
|
||||
| SC has high implementation cost in general | High | General literature consensus |
|
||||
| SC specifically impractical at 128 cores | Low | General claim, not XH-1-measured; workload- and topology-dependent |
|
||||
| Per-fence energy scales with acknowledging-core count | Medium | Logical consequence; no measurement |
|
||||
| Quantitative per-fence energy at 128 cores | None | INSUFFICIENT EVIDENCE |
|
||||
| Verification is hardest for RC variants | Low–Medium | Literature claim, not XH-1-measured |
|
||||
| XH-1 should adopt RVWMO | Low | Repository context not yet sufficient |
|
||||
| XH-1 should adopt RVWMO + Ztso | Low | Same |
|
||||
| Ztso is a small RTL delta over RVWMO | Low–Medium | Plausible qualitative claim; no RTL to substantiate |
|
||||
|
||||
## Open Questions
|
||||
|
||||
1. What is XH-1's microarchitecture style (in-order vs. out-of-order)?
|
||||
2. What coherence protocol and directory organization will XH-1 use?
|
||||
3. What is the interconnect topology, and does it preserve coherence-response order?
|
||||
4. What workloads are targeted — HPC, server, embedded?
|
||||
5. Does XH-1 include non-coherent agents (DMA, CXL, accelerators)?
|
||||
6. Is XH-1 a research vehicle, a prototype, or a product?
|
||||
7. What is the expected fence frequency in target software?
|
||||
8. What is the verification methodology — formal, simulation, or hybrid?
|
||||
9. Will XH-1 use the standard RISC-V litmus-test suite or define a custom one?
|
||||
10. Does XH-1 require compatibility with x86-software-only ports? (Drives Ztso inclusion.)
|
||||
11. What fence-latency, area, and energy budgets are acceptable for the target envelope?
|
||||
12. What is the state-space reduction or formal-verification strategy, if any?
|
||||
13. Which revision of the RISC-V Unprivileged ISA specification is the reference for XH-1?
|
||||
|
||||
## Sources
|
||||
|
||||
INSUFFICIENT EVIDENCE — The XH-1 Research repository does not yet contain source material specific to memory ordering. No citations, papers, or measurements can be honestly attributed to XH-1.
|
||||
|
||||
External ISA-level facts cited above (RVWMO, Ztso, FENCE instruction semantics) are drawn from the publicly ratified RISC-V Unprivileged ISA specification, which is the source of truth for the ISA-level memory model. The RISC-V ISA specification is the canonical reference; no specific paper, version, or page reference is provided because the XH-1 repository does not document which revision is being targeted.
|
||||
|
||||
INSUFFICIENT EVIDENCE — All quantitative claims about fence latency, area, power, energy, and verification cost at the XH-1 level have been removed or flagged. No benchmarks, no measurements, no RTL data, no simulation results are available in the repository at this time.
|
||||
@@ -0,0 +1 @@
|
||||
|
||||
+1
File diff suppressed because one or more lines are too long
@@ -0,0 +1,319 @@
|
||||
# Memory Ordering
|
||||
|
||||
## Status
|
||||
|
||||
Stub document. XH-1 memory ordering model has not yet been defined. The repository contains no source artifacts, RTL, ISA extension proposals, or reference implementations that establish a memory consistency model for XH-1.
|
||||
|
||||
## Abstract
|
||||
|
||||
Memory ordering defines the rules by which loads and stores issued by one or more cores become visible to other cores in the system. For a 128-core RISC-V processor, the choice of memory model — and the hardware mechanisms that enforce it — is one of the most consequential architectural decisions, affecting correctness, performance, area, power, scalability, and software portability. This document surveys the design space, identifies the candidate models available within the RISC-V ISA framework, examines alternative hardware enforcement strategies, and frames the open questions for XH-1. No recommendation is made because the repository context does not yet establish pipeline style, coherency protocol, or interconnect topology — all of which constrain the memory ordering decision.
|
||||
|
||||
## Research Question
|
||||
|
||||
What memory consistency model and hardware ordering mechanism should XH-1 adopt, given:
|
||||
|
||||
1. A 128-core replication target
|
||||
2. The RISC-V ISA baseline
|
||||
3. The need to balance software portability against hardware implementation cost
|
||||
4. Unknown microarchitecture style (in-order vs. out-of-order is not established in the repository)
|
||||
5. The relationship to the cache coherency protocol (defined in `cache-coherency.md`, currently SOON)
|
||||
|
||||
## Background
|
||||
|
||||
### Memory Consistency Models — Terminology
|
||||
|
||||
PROPOSAL — Use the following working taxonomy in this document:
|
||||
|
||||
- **Sequential Consistency (SC)**: All cores observe a single total order of memory operations that respects each core's program order.
|
||||
- **Total Store Order (TSO)**: Loads may bypass stores; stores from one core are seen in order by all cores. Used by x86.
|
||||
- **Relaxed / Weak Memory Models**: Many orderings are allowed; fences (barriers) are required to enforce constraints. Used by ARM, RISC-V, and PowerPC.
|
||||
- **Release Consistency (RC) / RCO / RCpc / RCsc**: A family of models with different fence semantics. Used by academic designs and the basis for much of C11/C++ memory model formalization.
|
||||
|
||||
### The RISC-V "RVWMO" Model
|
||||
|
||||
FACT — The RISC-V ISA specification defines the **RISC-V Weak Memory Ordering (RVWMO)** model as the baseline. Under RVWMO:
|
||||
|
||||
- Loads and stores to the same address from the same hart are ordered.
|
||||
- A store followed by a load to a *different* address may be reordered.
|
||||
- The `FENCE` instruction is used to enforce orderings. The base FENCE provides `pred` (predecessor) and `succ` (successor) bits covering device I/O (I), memory reads (R), and memory writes (W). Device ordering follows a stronger PostModel.
|
||||
- The optional `Ztso` extension provides **RVTSO**, which gives x86-like TSO semantics with a relaxed form of the FENCE.
|
||||
|
||||
ASSUMPTION — RVWMO + the Ztso optional extension is the available ISA-level model space that XH-1 must select from. (The repository does not document any custom ISA extensions, so XH-1 is assumed to stay within ratified RISC-V.)
|
||||
|
||||
### Why Memory Ordering is Hard at 128 Cores
|
||||
|
||||
OPEN QUESTION — The repository does not establish:
|
||||
- Whether XH-1 is a single die, multi-die package, or chiplet-based system
|
||||
- The interconnect topology (mesh, ring, crossbar, NoC, etc.)
|
||||
- Whether coherence is directory-based, broadcast-based, or hybrid
|
||||
- The private-vs-shared L2/L3 split
|
||||
- The memory-side directory or snoop filter organization
|
||||
|
||||
These decisions have direct bearing on which ordering mechanisms are feasible. Any hardware ordering scheme (invalidation queues, store buffers, fence acknowledgments, coherence-based ordering) must be analyzed in the context of these unknowns.
|
||||
|
||||
## Existing Approaches
|
||||
|
||||
### Approach A — RISC-V RVWMO (Baseline)
|
||||
|
||||
The standard weak memory model provided by the RISC-V ISA. Requires implementation of `FENCE`, `FENCE.I`, and the I/O ordering semantics.
|
||||
|
||||
**Mechanism**: Architecture permits aggressive reordering. Hardware typically uses:
|
||||
- Store buffers with forwarding
|
||||
- Load-load reordering queues
|
||||
- Coherence-based ordering: stores become visible at coherence invalidation/upgrade points
|
||||
- Fence implementation via global acknowledgment counters or by draining reorder queues
|
||||
|
||||
### Approach B — RVWMO + Ztso Extension
|
||||
|
||||
Adds RVTSO. Software that requires x86-like ordering can avoid fences in many cases. Still RISC-V compliant.
|
||||
|
||||
**Mechanism**: Stores from one hart are serialized in the interconnect; loads may still pass earlier stores to different addresses.
|
||||
|
||||
### Approach C — Sequential Consistency (SC)
|
||||
|
||||
All operations observed in a single global order respecting program order. Simplest software model; most expensive hardware.
|
||||
|
||||
**Mechanism**: Typically requires store-buffer draining on loads, broadcast-based coherence, and conservative replay mechanisms. Very costly for 128 cores.
|
||||
|
||||
### Approach D — Release Consistency variants (RCsc, RCpc, RCO)
|
||||
|
||||
Academic models where release/acquire ordering is cheap but stronger fences are required for full memory ordering. Basis of C11/C++ model.
|
||||
|
||||
**Mechanism**: Only release/acquire fences require acknowledgment; acquire-only loads are cheap. Alignment with C11 is a strong software benefit.
|
||||
|
||||
## Alternative Designs
|
||||
|
||||
### Hardware Ordering Mechanisms (to be evaluated once coherency and interconnect are known)
|
||||
|
||||
PROPOSAL — The following mechanisms should be considered for fence/ordering enforcement:
|
||||
|
||||
1. **Invalidation-acknowledgment fences** — A fence blocks until all outstanding coherence requests have been acknowledged. Simple, common, but scales poorly with directory hop count and aggregate invalidation traffic at 128 cores.
|
||||
|
||||
2. **Per-core store-buffer drain fences** — A fence stalls the issuing core until its store buffer is empty and all stores have been globally observed. Predictable latency; serializes through the coherence fabric.
|
||||
|
||||
3. **Tournament / Token-based fences** — Fences acquire a global token (Coherence-Next-Line style) or wait for a "fence epoch" counter. Avoids worst-case drain, but adds global state.
|
||||
|
||||
4. **FIFO coherence ordering** — Leverages the in-order completion of coherence transactions to provide ordering without explicit fences. Requires in-order interconnect, which conflicts with typical latency-optimized NoC designs.
|
||||
|
||||
5. **Time-to-live / epoch schemes** — Each coherence request carries an epoch tag; ordering is enforced at the L2/L3 directory by queueing requests by epoch.
|
||||
|
||||
6. **Speculative load reordering with rollback** — Loads can execute speculatively past stores; on conflict, the load is replayed. Common in high-performance out-of-order cores (e.g., IBM POWER).
|
||||
|
||||
ASSUMPTION — XH-1's microarchitecture style is not yet established. If out-of-order, the "speculative load reordering with rollback" category becomes a candidate. If in-order, mechanism choices narrow considerably.
|
||||
|
||||
## Comparison
|
||||
|
||||
| Property | SC | RVWMO | RVWMO + Ztso | RC variants |
|
||||
|----------|----|----|----|----|
|
||||
| Software ease | Highest | Medium | Medium-High | Low-Medium |
|
||||
| Hardware cost | Highest | Low | Low-Medium | Low |
|
||||
| Performance headroom | Lowest | Highest | High | High |
|
||||
| Fence latency | N/A | High | Medium | Variable |
|
||||
| Verification cost | Lowest | High | High | Highest |
|
||||
| C11/C++ alignment | Approximate | Approximate | Weak | Direct |
|
||||
| 128-core scalability | Poor | Good | Good | Good |
|
||||
|
||||
Quantitative figures: INSUFFICIENT EVIDENCE. All numbers above are relative qualitative ratings, not benchmarked measurements.
|
||||
|
||||
## Advantages
|
||||
|
||||
### RVWMO (Baseline)
|
||||
- **Lowest hardware cost** of all candidates for the base case — most orderings are simply not enforced.
|
||||
- **Maximum performance** potential — reordering is unconstrained.
|
||||
- **Familiar** to RISC-V software ecosystem (Linux, RISC-V GCC, LLVM).
|
||||
- **Standardized** — concrete compliance test suite exists.
|
||||
|
||||
### RVWMO + Ztso
|
||||
- **x86-like TSO** for software that benefits from stronger ordering.
|
||||
- **Backwards compatible** — software not using Ztso still runs correctly under RVWMO.
|
||||
- **Modest hardware cost** — primarily requires store-store ordering at the coherence layer.
|
||||
|
||||
### Release Consistency Variants
|
||||
- **Excellent fit for C11/C++** atomics — acquire/release are the most-used fence types in real code.
|
||||
- **Cheapest common case** — acquire and release can be implemented with lightweight mechanisms.
|
||||
- **Hardware cost is concentrated on the rare case** of full seq_cst fences.
|
||||
|
||||
## Disadvantages
|
||||
|
||||
### RVWMO (Baseline)
|
||||
- **Fence latency is high** at 128 cores — global acknowledgment is expensive.
|
||||
- **Verification complexity is high** — many legal reorderings; the litmus-test failure surface is large.
|
||||
- **Software burden** — kernel and runtime code must insert fences correctly.
|
||||
|
||||
### RVWMO + Ztso
|
||||
- **Two consistency models in one chip** — adds documentation, validation, and software education cost.
|
||||
- **In-flight mixing** of ordering regimes in the same software is hard to reason about.
|
||||
|
||||
### Release Consistency Variants
|
||||
- **Non-standard for RISC-V** — would require ISA extension. Cannot be recommended without a strong software-side driver.
|
||||
- **Verification cost is the highest** of the candidates because of subtle fence semantics.
|
||||
|
||||
### Sequential Consistency
|
||||
- **Performance cost is severe at 128 cores** — every load may stall on store-buffer drain.
|
||||
- **No selective escape valve** — software cannot opt out of the strongest model.
|
||||
- **Does not match RISC-V ecosystem expectations** — surprising to RISC-V software developers.
|
||||
|
||||
## XH-1 Considerations
|
||||
|
||||
OPEN QUESTION — The following XH-1-internal questions are unresolved by the repository:
|
||||
|
||||
1. Is XH-1's pipeline in-order or out-of-order?
|
||||
2. What is the coherence protocol (MOESI, MESI, directory-based, broadcast)?
|
||||
3. Is the interconnect a NoC, ring, or crossbar? Does it preserve in-order delivery of coherence responses?
|
||||
4. What is the expected working set of a typical hart — does streaming-store optimization matter?
|
||||
5. Does XH-1 target HPC, server, embedded, or mixed workloads? (Workload affects fence frequency.)
|
||||
6. Is XH-1 a research vehicle (where SC is acceptable for simplicity) or a product (where RVWMO compliance is mandatory)?
|
||||
7. Does the memory map use non-coherent regions (e.g., DMA, I/O) that interact with the ordering model?
|
||||
|
||||
ASSUMPTION — If XH-1 targets RISC-V ecosystem compatibility (Linux, RISC-V GCC, RISC-V LLVM), then RVWMO is mandatory. Ztso may optionally be added for x86-software porting convenience.
|
||||
|
||||
## 128-Core Scalability
|
||||
|
||||
PROPOSAL — Scalability analysis of each candidate at 128 cores:
|
||||
|
||||
### Fence Latency
|
||||
- **Fence global acknowledgment** scales roughly with the worst-case coherence round-trip time. In a 128-core system, this can be many tens of cycles in a NoC, hundreds in a ring.
|
||||
- **Store-buffer drain** latency is approximately proportional to the store buffer depth and the time to invalidate all sharers.
|
||||
- **Token / epoch fences** scale better because they avoid draining the entire coherence fabric — they synchronize at a logical epoch boundary.
|
||||
|
||||
### Coherence Traffic
|
||||
- **RVWMO fences** generate invalidation or acknowledgment traffic that competes with normal coherence traffic.
|
||||
- **SC at 128 cores** can saturate a directory-based interconnect because every load's coherence transaction may stall pending store ordering.
|
||||
- **RC variants** concentrate traffic on the fence path only.
|
||||
|
||||
### Verification Scalability
|
||||
- **State-space explosion**: at 128 cores, the number of concurrently observable memory operations grows combinatorially. Memory-ordering verification is a known scaling bottleneck.
|
||||
- **RC variants** are hardest to verify because of subtle fence interactions.
|
||||
- **SC is easiest to verify** in principle, but worst in performance.
|
||||
|
||||
INSUFFICIENT EVIDENCE — No quantitative fence-latency or coherence-traffic data is available for XH-1. Estimates cannot be made without knowing the interconnect.
|
||||
|
||||
## Performance Considerations
|
||||
|
||||
PROPOSAL — Performance-impact dimensions to evaluate:
|
||||
|
||||
1. **Steady-state IPC** for memory-bound workloads — RVWMO best, SC worst.
|
||||
2. **Fence latency** — measured as cycles from fence issue to fence completion. Dominant for synchronization-heavy code.
|
||||
3. **Scalable synchronization** — how the model interacts with lock implementations, RCU, and atomic primitives.
|
||||
4. **Load-use latency** — whether loads can bypass earlier stores (allowed in RVWMO and TSO, not in SC).
|
||||
5. **Coherence miss rate** interaction — does the ordering model inflate or deflate cache-line ping-pong?
|
||||
|
||||
ASSUMPTION — For HPC workloads with high-fence frequency (e.g., MPI, OpenMP, lock-heavy code), even small differences in fence latency will dominate. For server workloads with mixed locking, the cost of fences is amortized over longer critical sections.
|
||||
|
||||
## Area Considerations
|
||||
|
||||
PROPOSAL — Area cost of ordering mechanisms:
|
||||
|
||||
| Mechanism | Approximate area cost per core |
|
||||
|-----------|-------------------------------|
|
||||
| Store buffer with forwarding | Small (few entries × cache-line width) |
|
||||
| Load-load reordering queue | Small to medium (grows with IPC) |
|
||||
| Fence-acknowledgment counter | Negligible (per-core counter) |
|
||||
| Global epoch / token state | Medium (one per chip, not per core) |
|
||||
| Speculative load reorder + rollback | Large (load-load queue, replay logic) |
|
||||
|
||||
ASSUMPTION — All numbers in the table are order-of-magnitude estimates. Precise XH-1 figures: INSUFFICIENT EVIDENCE.
|
||||
|
||||
PROPOSAL — At 128 cores, even small per-core area costs multiply by 128. Global state (epoch counters, fence acknowledgers) is O(1) and becomes attractive.
|
||||
|
||||
## Power and Energy Considerations
|
||||
|
||||
PROPOSAL — Energy implications:
|
||||
|
||||
- **Fences** generate global coherence traffic — energy cost scales with the number of cores that must acknowledge.
|
||||
- **Speculative load reordering** wastes energy on rollback when a violation is detected.
|
||||
- **Strong ordering models** keep more coherence state in flight (more invalidations, more retries).
|
||||
- **RVWMO** with lightweight fences (when software permits) is the lowest-energy regime.
|
||||
|
||||
OPEN QUESTION — The 128-core replication means the per-fence energy cost is multiplied. A fence that costs 100 pJ per core at 16 cores costs 800× at 128 cores in the worst case (if all cores must acknowledge). No published figure for XH-1 is available — INSUFFICIENT EVIDENCE.
|
||||
|
||||
## Implementation Considerations
|
||||
|
||||
PROPOSAL — Implementation order of dependencies (the repository does not establish these, but the research area suggests the following):
|
||||
|
||||
1. Define coherency protocol → `cache-coherency.md`
|
||||
2. Define interconnect topology → not yet in repository
|
||||
3. Define memory hierarchy → `memory-hierarchy.md`
|
||||
4. Define memory ordering model → this document
|
||||
5. Define atomic primitives → `atomics.md`
|
||||
|
||||
ASSUMPTION — The Ztso extension is a small RTL delta over RVWMO. The dominant implementation cost is verification, not area.
|
||||
|
||||
OPEN QUESTION — Does XH-1 have a coherent accelerator fabric or non-coherent IO (e.g., CXL, DMA engines)? If so, the memory ordering model must define how these agents interact.
|
||||
|
||||
## Verification Considerations
|
||||
|
||||
PROPOSAL — Verification challenges:
|
||||
|
||||
- **Litmus tests**: Standard RISC-V litmus test suite must be run on RTL and on model-based environments.
|
||||
- **State-space explosion**: 128 cores × N outstanding operations = combinatorial explosion. Likely requires state-space reduction or bisimulation techniques.
|
||||
- **Fence semantics**: Each fence variant must be exhaustively tested at corner cases (interrupts, exceptions, MMIO).
|
||||
- **Coherence-ordering interaction**: The ordering model must be checked against the chosen coherence protocol for consistency.
|
||||
|
||||
INSUFFICIENT EVIDENCE — No XH-1 verification infrastructure, formal spec, or litmus-test set is documented in the repository.
|
||||
|
||||
ASSUMPTION — RVWMO is the most documented and tested model for RISC-V. Ztso is a small addition. SC and RC variants would require custom formal infrastructure.
|
||||
|
||||
## Software Considerations
|
||||
|
||||
PROPOSAL — Software-side impact:
|
||||
|
||||
- **Linux kernel**: Designed for weak memory models with appropriate fences. Compatible with RVWMO.
|
||||
- **C11/C++ atomics**: Map cleanly onto release/acquire semantics. Compatible with RVWMO and with RC variants.
|
||||
- **x86 software ports**: Benefit from Ztso for fewer fences.
|
||||
- **OpenSHMEM, MPI shmem**: Often assume TSO-like ordering; Ztso helps.
|
||||
- **HPC codes with hand-rolled atomics**: Highly sensitive to fence latency.
|
||||
|
||||
ASSUMPTION — XH-1 will run RISC-V Linux. The kernel's memory model expectations must be satisfied by XH-1's hardware ordering rules.
|
||||
|
||||
## Recommendation
|
||||
|
||||
INSUFFICIENT EVIDENCE to make a final recommendation.
|
||||
|
||||
The repository does not yet establish:
|
||||
- Pipeline microarchitecture
|
||||
- Coherence protocol
|
||||
- Interconnect topology
|
||||
- Workload target
|
||||
|
||||
A recommendation can be made only after the following documents in the memory research area are filled in: `memory-architecture.md`, `cache-coherency.md`, `memory-hierarchy.md`. Without these, any model selection is premature.
|
||||
|
||||
PROPOSAL — Tentative default pending more repository context:
|
||||
|
||||
1. **Baseline**: RVWMO (mandatory for RISC-V ecosystem).
|
||||
2. **Optional add-on**: Ztso extension for x86-software porting convenience and modest hardware cost.
|
||||
3. **Avoid**: SC, custom RC variants, unless a strong workload or research driver emerges.
|
||||
4. **Defer to**: `cache-coherency.md` for the coherence-side ordering mechanism.
|
||||
|
||||
## Confidence
|
||||
|
||||
| Topic | Confidence | Reason |
|
||||
|-------|------------|--------|
|
||||
| RVWMO is the RISC-V baseline | High | Ratified ISA spec |
|
||||
| Ztso is a small optional extension | High | Ratified extension |
|
||||
| SC is impractical at 128 cores | Medium | General consensus in literature; XH-1-specific number not available |
|
||||
| Per-fence energy scales with core count | Medium | Logical consequence; no measurement |
|
||||
| Verification is hardest for RC variants | Medium | General consensus |
|
||||
| XH-1 should adopt RVWMO | Low | Repository context not yet sufficient |
|
||||
| XH-1 should adopt RVWMO + Ztso | Low | Same |
|
||||
|
||||
## Open Questions
|
||||
|
||||
1. What is XH-1's microarchitecture style (in-order vs. out-of-order)?
|
||||
2. What coherence protocol and directory organization will XH-1 use?
|
||||
3. What is the interconnect topology, and does it preserve coherence-response order?
|
||||
4. What workloads are targeted — HPC, server, embedded?
|
||||
5. Does XH-1 include non-coherent agents (DMA, CXL, accelerators)?
|
||||
6. Is XH-1 a research vehicle, a prototype, or a product?
|
||||
7. What is the expected fence frequency in target software?
|
||||
8. What is the verification methodology — formal, simulation, or hybrid?
|
||||
9. Will XH-1 use the standard RISC-V litmus-test suite or define a custom one?
|
||||
10. Does XH-1 require compatibility with x86-software-only ports? (Drives Ztso inclusion.)
|
||||
|
||||
## Sources
|
||||
|
||||
INSUFFICIENT EVIDENCE — The XH-1 Research repository does not yet contain source material specific to memory ordering. No citations, papers, or measurements can be honestly attributed to XH-1.
|
||||
|
||||
External ISA-level facts cited above (RVWMO, Ztso, FENCE instruction) are drawn from the publicly ratified RISC-V ISA specification, which is the source of truth for the ISA-level memory model. No specific paper, version, or page reference is provided because the XH-1 repository does not document which revision is being targeted.
|
||||
|
||||
No benchmarks, no measurements, no RTL data, no simulation results are available in the repository at this time.
|
||||
+235
File diff suppressed because one or more lines are too long
@@ -0,0 +1,323 @@
|
||||
# Memory Ordering
|
||||
|
||||
## Status
|
||||
|
||||
Stub document. XH-1 memory ordering model has not yet been defined. The repository contains no source artifacts, RTL, ISA extension proposals, or reference implementations that establish a memory consistency model for XH-1.
|
||||
|
||||
## Abstract
|
||||
|
||||
Memory ordering defines the rules by which loads and stores issued by one or more cores become visible to other cores in the system. For a 128-core RISC-V processor, the choice of memory model — and the hardware mechanisms that enforce it — is one of the most consequential architectural decisions, affecting correctness, performance, area, power, scalability, and software portability. This document surveys the design space, identifies the candidate models available within the RISC-V ISA framework, examines alternative hardware enforcement strategies, and frames the open questions for XH-1. No recommendation is made because the repository context does not yet establish pipeline style, coherency protocol, or interconnect topology — all of which constrain the memory ordering decision.
|
||||
|
||||
## Research Question
|
||||
|
||||
What memory consistency model and hardware ordering mechanism should XH-1 adopt, given:
|
||||
|
||||
1. A 128-core replication target
|
||||
2. The RISC-V ISA baseline
|
||||
3. The need to balance software portability against hardware implementation cost
|
||||
4. Unknown microarchitecture style (in-order vs. out-of-order is not established in the repository)
|
||||
5. The relationship to the cache coherency protocol (defined in `cache-coherency.md`, currently SOON)
|
||||
|
||||
## Background
|
||||
|
||||
### Memory Consistency Models — Terminology
|
||||
|
||||
PROPOSAL — Use the following working taxonomy in this document:
|
||||
|
||||
- **Sequential Consistency (SC)**: All cores observe a single total order of memory operations that respects each core's program order.
|
||||
- **Total Store Order (TSO)**: Loads may bypass stores; stores from one core are seen in order by all cores. Used by x86.
|
||||
- **Relaxed / Weak Memory Models**: Many orderings are allowed; fences (barriers) are required to enforce constraints. Used by ARM, RISC-V, and PowerPC.
|
||||
- **Release Consistency (RC) / RCO / RCpc / RCsc**: A family of models with different fence semantics. Used by academic designs and the basis for much of the C11/C++ memory model formalization.
|
||||
|
||||
### The RISC-V "RVWMO" Model
|
||||
|
||||
FACT — The RISC-V ISA specification defines the **RISC-V Weak Memory Ordering (RVWMO)** model as the baseline. Under RVWMO:
|
||||
|
||||
- Loads and stores to the same address from the same hart are ordered.
|
||||
- A store followed by a load to a *different* address may be reordered.
|
||||
- The `FENCE` instruction is used to enforce orderings. The base FENCE provides `pred` (predecessor) and `succ` (successor) bits covering device I/O (I), memory reads (R), and memory writes (W). Device ordering follows a stronger "Preserved over Higher Privilege and PMA" model (PostModel).
|
||||
- The optional `Ztso` extension provides **RVTSO**, which gives x86-like TSO semantics with a relaxed form of the FENCE.
|
||||
|
||||
ASSUMPTION — RVWMO + the Ztso optional extension is the available ISA-level model space that XH-1 must select from. (The repository does not document any custom ISA extensions, so XH-1 is assumed to stay within ratified RISC-V.)
|
||||
|
||||
### Why Memory Ordering is Hard at 128 Cores
|
||||
|
||||
OPEN QUESTION — The repository does not establish:
|
||||
- Whether XH-1 is a single die, multi-die package, or chiplet-based system
|
||||
- The interconnect topology (mesh, ring, crossbar, NoC, etc.)
|
||||
- Whether coherence is directory-based, broadcast-based, or hybrid
|
||||
- The private-vs-shared L2/L3 split
|
||||
- The memory-side directory or snoop filter organization
|
||||
|
||||
These decisions have direct bearing on which ordering mechanisms are feasible. Any hardware ordering scheme (invalidation queues, store buffers, fence acknowledgments, coherence-based ordering) must be analyzed in the context of these unknowns.
|
||||
|
||||
## Existing Approaches
|
||||
|
||||
### Approach A — RISC-V RVWMO (Baseline)
|
||||
|
||||
The standard weak memory model provided by the RISC-V ISA. Requires implementation of `FENCE`, `FENCE.I`, and the I/O ordering semantics.
|
||||
|
||||
**Mechanism**: Architecture permits aggressive reordering. Hardware typically uses:
|
||||
- Store buffers with forwarding
|
||||
- Load-load reordering queues
|
||||
- Coherence-based ordering: stores become visible at coherence invalidation/upgrade points
|
||||
- Fence implementation via global acknowledgment counters or by draining reorder queues
|
||||
|
||||
### Approach B — RVWMO + Ztso Extension
|
||||
|
||||
Adds RVTSO. Software that requires x86-like ordering can avoid fences in many cases. Still RISC-V compliant.
|
||||
|
||||
**Mechanism**: Stores from one hart are serialized in the interconnect; loads may still pass earlier stores to different addresses.
|
||||
|
||||
### Approach C — Sequential Consistency (SC)
|
||||
|
||||
All operations observed in a single global order respecting program order. Simplest software model; most expensive hardware.
|
||||
|
||||
**Mechanism**: Typically requires store-buffer draining on loads, broadcast-based coherence, and conservative replay mechanisms. Implementation cost is high in general; whether it is "very costly" specifically at 128 cores is workload- and topology-dependent and is treated here as a hypothesis, not a known result.
|
||||
|
||||
### Approach D — Release Consistency variants (RCsc, RCpc, RCO)
|
||||
|
||||
Academic models where release/acquire ordering is cheap but stronger fences are required for full memory ordering. Basis of C11/C++ model.
|
||||
|
||||
**Mechanism**: Only release/acquire fences require acknowledgment; acquire-only loads are cheap. Alignment with C11 is a strong software benefit. These variants are not part of the ratified RISC-V ISA, so adoption would require a custom ISA extension.
|
||||
|
||||
## Alternative Designs
|
||||
|
||||
### Hardware Ordering Mechanisms (to be evaluated once coherency and interconnect are known)
|
||||
|
||||
PROPOSAL — The following mechanisms should be considered for fence/ordering enforcement:
|
||||
|
||||
1. **Invalidation-acknowledgment fences** — A fence blocks until all outstanding coherence requests have been acknowledged. Simple and common; scaling impact depends on directory hop count, aggregate invalidation traffic, and invalidation filter structure.
|
||||
|
||||
2. **Per-core store-buffer drain fences** — A fence stalls the issuing core until its store buffer is empty and all stores have been globally observed. Predictable latency; serializes through the coherence fabric.
|
||||
|
||||
3. **Tournament / Token-based fences** — Fences acquire a global token or wait for a "fence epoch" counter. Avoids worst-case drain, but adds global state.
|
||||
|
||||
4. **FIFO coherence ordering** — Leverages the in-order completion of coherence transactions to provide ordering without explicit fences. Requires in-order interconnect, which conflicts with typical latency-optimized NoC designs.
|
||||
|
||||
5. **Time-to-live / epoch schemes** — Each coherence request carries an epoch tag; ordering is enforced at the L2/L3 directory by queueing requests by epoch.
|
||||
|
||||
6. **Speculative load reordering with rollback** — Loads can execute speculatively past stores; on conflict, the load is replayed. Common in high-performance out-of-order cores (e.g., IBM POWER).
|
||||
|
||||
ASSUMPTION — XH-1's microarchitecture style is not yet established. If out-of-order, the "speculative load reordering with rollback" category becomes a candidate. If in-order, mechanism choices narrow considerably.
|
||||
|
||||
## Comparison
|
||||
|
||||
| Property | SC | RVWMO | RVWMO + Ztso | RC variants |
|
||||
|----------|----|----|----|----|
|
||||
| Software ease | Highest | Medium | Medium-High | Low-Medium |
|
||||
| Hardware cost | Highest | Low | Low-Medium | Low |
|
||||
| Performance headroom | Lowest | Highest | High | High |
|
||||
| Fence latency | N/A | Variable | Variable | Variable |
|
||||
| Verification cost | Lowest in principle | High | High | High (likely highest) |
|
||||
| C11/C++ alignment | Approximate | Approximate | Weak | Direct |
|
||||
| 128-core scalability | Poor (hypothesis) | Good (claim) | Good (claim) | Good (claim) |
|
||||
|
||||
INSUFFICIENT EVIDENCE — Quantitative fence-latency, coherence-traffic, and scalability numbers for XH-1 are unavailable. Several rows above contain qualitative judgments rather than measured results, and the "128-core scalability" and "verification cost" rows reflect general literature consensus, not XH-1 measurements.
|
||||
|
||||
## Advantages
|
||||
|
||||
### RVWMO (Baseline)
|
||||
- **Lowest hardware cost** of all candidates for the base case — most orderings are simply not enforced.
|
||||
- **Maximum performance** potential — reordering is unconstrained.
|
||||
- **Familiar** to RISC-V software ecosystem (Linux, RISC-V GCC, LLVM).
|
||||
- **Standardized** — concrete compliance test suite exists.
|
||||
|
||||
### RVWMO + Ztso
|
||||
- **x86-like TSO** for software that benefits from stronger ordering.
|
||||
- **Backwards compatible** — software not using Ztso still runs correctly under RVWMO.
|
||||
- **Modest hardware cost** — primarily requires store-store ordering at the coherence layer.
|
||||
|
||||
### Release Consistency Variants
|
||||
- **Excellent fit for C11/C++** atomics — acquire/release are the most-used fence types in real code.
|
||||
- **Cheapest common case** — acquire and release can be implemented with lightweight mechanisms.
|
||||
- **Hardware cost is concentrated on the rare case** of full seq_cst fences.
|
||||
|
||||
## Disadvantages
|
||||
|
||||
### RVWMO (Baseline)
|
||||
- **Fence latency depends on coherence round-trip** — at 128 cores, fence latency is expected to be high because global acknowledgment must traverse the coherence fabric. Specific cycle counts: INSUFFICIENT EVIDENCE.
|
||||
- **Verification complexity is high** — many legal reorderings; the litmus-test failure surface is large.
|
||||
- **Software burden** — kernel and runtime code must insert fences correctly.
|
||||
|
||||
### RVWMO + Ztso
|
||||
- **Two consistency models in one chip** — adds documentation, validation, and software education cost.
|
||||
- **In-flight mixing** of ordering regimes in the same software is hard to reason about.
|
||||
|
||||
### Release Consistency Variants
|
||||
- **Non-standard for RISC-V** — would require a custom ISA extension. Cannot be recommended without a strong software-side driver.
|
||||
- **Verification cost is high** in general because of subtle fence semantics; whether it is the "highest" among candidates is a claim about the literature, not a measured XH-1 result.
|
||||
|
||||
### Sequential Consistency
|
||||
- **Performance cost is severe in the general case** — every load may stall on store-buffer drain; whether this is "severe at 128 cores" specifically is workload- and topology-dependent and is not established by the repository.
|
||||
- **No selective escape valve** — software cannot opt out of the strongest model.
|
||||
- **Does not match RISC-V ecosystem expectations** — surprising to RISC-V software developers.
|
||||
|
||||
## XH-1 Considerations
|
||||
|
||||
OPEN QUESTION — The following XH-1-internal questions are unresolved by the repository:
|
||||
|
||||
1. Is XH-1's pipeline in-order or out-of-order?
|
||||
2. What is the coherence protocol (MOESI, MESI, directory-based, broadcast)?
|
||||
3. Is the interconnect a NoC, ring, or crossbar? Does it preserve in-order delivery of coherence responses?
|
||||
4. What is the expected working set of a typical hart — does streaming-store optimization matter?
|
||||
5. Does XH-1 target HPC, server, embedded, or mixed workloads? (Workload affects fence frequency.)
|
||||
6. Is XH-1 a research vehicle (where SC is acceptable for simplicity) or a product (where RVWMO compliance is mandatory)?
|
||||
7. Does the memory map use non-coherent regions (e.g., DMA, I/O) that interact with the ordering model?
|
||||
|
||||
ASSUMPTION — If XH-1 targets RISC-V ecosystem compatibility (Linux, RISC-V GCC, RISC-V LLVM), then RVWMO is mandatory. Ztso may optionally be added for x86-software porting convenience.
|
||||
|
||||
## 128-Core Scalability
|
||||
|
||||
PROPOSAL — Scalability analysis of each candidate at 128 cores. Quantitative figures: INSUFFICIENT EVIDENCE. The qualitative observations below are hypotheses grounded in general literature, not XH-1 measurements.
|
||||
|
||||
### Fence Latency
|
||||
- **Fence global acknowledgment** latency is bounded below by the worst-case coherence round-trip time. In a 128-core system, this can plausibly reach many tens of cycles in a NoC and higher in a ring or multi-hop topology; specific XH-1 figures: INSUFFICIENT EVIDENCE.
|
||||
- **Store-buffer drain** latency is approximately proportional to the store buffer depth and the time to invalidate all sharers.
|
||||
- **Token / epoch fences** may scale better because they avoid draining the entire coherence fabric — they synchronize at a logical epoch boundary.
|
||||
|
||||
### Coherence Traffic
|
||||
- **RVWMO fences** generate invalidation or acknowledgment traffic that competes with normal coherence traffic. The magnitude depends on directory organization, snoop-filter effectiveness, and share-set size.
|
||||
- **SC at 128 cores** can plausibly saturate a directory-based interconnect because every load's coherence transaction may stall pending store ordering; this is a hypothesis, not a measured XH-1 result.
|
||||
- **RC variants** concentrate traffic on the fence path only.
|
||||
|
||||
### Verification Scalability
|
||||
- **State-space explosion**: at 128 cores, the number of concurrently observable memory operations grows combinatorially in the unmitigated model. Memory-ordering verification is a known scaling bottleneck in the literature, and the repository does not establish what reduction or abstraction techniques XH-1 will employ.
|
||||
- **RC variants** are commonly cited as among the hardest to verify because of subtle fence interactions; whether they are hardest in the absolute sense is a literature claim, not an XH-1 measurement.
|
||||
- **SC is easiest to verify in principle** because it admits fewer legal behaviors, but worst in performance.
|
||||
|
||||
INSUFFICIENT EVIDENCE — No quantitative fence-latency or coherence-traffic data is available for XH-1. Estimates cannot be made without knowing the interconnect.
|
||||
|
||||
## Performance Considerations
|
||||
|
||||
PROPOSAL — Performance-impact dimensions to evaluate:
|
||||
|
||||
1. **Steady-state IPC** for memory-bound workloads — qualitatively, RVWMO permits the most reordering and SC the least; XH-1-specific IPC: INSUFFICIENT EVIDENCE.
|
||||
2. **Fence latency** — measured as cycles from fence issue to fence completion. Dominant for synchronization-heavy code. XH-1-specific: INSUFFICIENT EVIDENCE.
|
||||
3. **Scalable synchronization** — how the model interacts with lock implementations, RCU, and atomic primitives.
|
||||
4. **Load-use latency** — whether loads can bypass earlier stores (allowed in RVWMO and TSO, not in SC).
|
||||
5. **Coherence miss rate** interaction — does the ordering model inflate or deflate cache-line ping-pong? Model-dependent; XH-1-specific: INSUFFICIENT EVIDENCE.
|
||||
|
||||
ASSUMPTION — For HPC workloads with high-fence frequency (e.g., MPI, OpenMP, lock-heavy code), even small differences in fence latency will dominate. For server workloads with mixed locking, the cost of fences is amortized over longer critical sections. No XH-1 workload profile is documented in the repository.
|
||||
|
||||
## Area Considerations
|
||||
|
||||
PROPOSAL — Qualitative area cost of ordering mechanisms:
|
||||
|
||||
| Mechanism | Approximate area cost per core |
|
||||
|-----------|-------------------------------|
|
||||
| Store buffer with forwarding | Small (few entries × cache-line width) |
|
||||
| Load-load reordering queue | Small to medium (grows with IPC and reorder depth) |
|
||||
| Fence-acknowledgment counter | Negligible (per-core counter) |
|
||||
| Global epoch / token state | Medium (one per chip, not per core) |
|
||||
| Speculative load reorder + rollback | Large (load-load queue, replay logic) |
|
||||
|
||||
INSUFFICIENT EVIDENCE — All descriptions above are qualitative. Precise XH-1 area figures (in gates, µm², or mm²) are unavailable.
|
||||
|
||||
PROPOSAL — At 128 cores, even small per-core area costs multiply by 128. Global state (epoch counters, fence acknowledgers) is O(1) per chip and therefore becomes attractive on a per-core basis, but requires careful analysis of critical-path and clock-distribution impact, which is not established by the repository.
|
||||
|
||||
## Power and Energy Considerations
|
||||
|
||||
PROPOSAL — Qualitative energy implications:
|
||||
|
||||
- **Fences** generate global coherence traffic — energy cost scales with the number of cores that must acknowledge. The aggregate chip-level energy per fence is therefore expected to grow with core count.
|
||||
- **Speculative load reordering** wastes energy on rollback when a violation is detected.
|
||||
- **Strong ordering models** keep more coherence state in flight (more invalidations, more retries).
|
||||
- **RVWMO** with lightweight fences (when software permits) is the lowest-energy regime in the general case.
|
||||
|
||||
INSUFFICIENT EVIDENCE — The earlier text claimed that a 100 pJ-per-core fence at 16 cores costs 800× at 128 cores in the worst case (if all cores must acknowledge). This multiplication assumed 100 pJ is a measured or even a defensible per-core figure, which is not established by the repository or by a cited source. The 100 pJ figure is removed. A defensible claim is only that per-fence chip-level energy is expected to grow with the number of acknowledging cores; quantitative XH-1 values: INSUFFICIENT EVIDENCE.
|
||||
|
||||
## Implementation Considerations
|
||||
|
||||
PROPOSAL — Implementation order of dependencies (the repository does not establish these, but the research area suggests the following):
|
||||
|
||||
1. Define coherency protocol → `cache-coherency.md`
|
||||
2. Define interconnect topology → not yet in repository
|
||||
3. Define memory hierarchy → `memory-hierarchy.md`
|
||||
4. Define memory ordering model → this document
|
||||
5. Define atomic primitives → `atomics.md`
|
||||
|
||||
ASSUMPTION — The Ztso extension is a small RTL delta over RVWMO. The dominant implementation cost is verification, not area. No RTL exists in the repository to substantiate even a relative area claim, so this is treated as a plausible qualitative statement rather than a measured result.
|
||||
|
||||
OPEN QUESTION — Does XH-1 have a coherent accelerator fabric or non-coherent IO (e.g., CXL, DMA engines)? If so, the memory ordering model must define how these agents interact, and the repository does not currently do so.
|
||||
|
||||
## Verification Considerations
|
||||
|
||||
PROPOSAL — Verification challenges:
|
||||
|
||||
- **Litmus tests**: A standard RISC-V litmus test suite is available externally; whether XH-1 will adopt it as-is or define a custom suite is OPEN.
|
||||
- **State-space explosion**: 128 cores × N outstanding operations produces a combinatorial unmitigated state space. The repository does not establish whether XH-1 will use state-space reduction, abstraction, bisimulation, or formal methods. Any specific XH-1 claim: INSUFFICIENT EVIDENCE.
|
||||
- **Fence semantics**: Each fence variant must be exhaustively tested at corner cases (interrupts, exceptions, MMIO).
|
||||
- **Coherence-ordering interaction**: The ordering model must be checked against the chosen coherence protocol for consistency. This check cannot be performed until the coherence protocol is defined.
|
||||
|
||||
INSUFFICIENT EVIDENCE — No XH-1 verification infrastructure, formal spec, or litmus-test set is documented in the repository.
|
||||
|
||||
ASSUMPTION — RVWMO is the most documented and tested model for RISC-V. Ztso is a small addition. SC and RC variants would require custom formal infrastructure. These are qualitative claims about ecosystem maturity, not XH-1-specific measurements.
|
||||
|
||||
## Software Considerations
|
||||
|
||||
PROPOSAL — Software-side impact:
|
||||
|
||||
- **Linux kernel**: Designed for weak memory models with appropriate fences. Compatible with RVWMO.
|
||||
- **C11/C++ atomics**: Map cleanly onto release/acquire semantics. Compatible with RVWMO and with RC variants.
|
||||
- **x86 software ports**: Benefit from Ztso for fewer fences.
|
||||
- **OpenSHMEM, MPI shmem**: Often assume TSO-like ordering; Ztso helps.
|
||||
- **HPC codes with hand-rolled atomics**: Highly sensitive to fence latency.
|
||||
|
||||
ASSUMPTION — XH-1 will run RISC-V Linux. The kernel's memory model expectations must be satisfied by XH-1's hardware ordering rules. The repository does not document whether Linux is in fact the target.
|
||||
|
||||
## Recommendation
|
||||
|
||||
INSUFFICIENT EVIDENCE to make a final recommendation.
|
||||
|
||||
The repository does not yet establish:
|
||||
- Pipeline microarchitecture
|
||||
- Coherence protocol
|
||||
- Interconnect topology
|
||||
- Workload target
|
||||
|
||||
A recommendation can be made only after the following documents in the memory research area are filled in: `memory-architecture.md`, `cache-coherency.md`, `memory-hierarchy.md`. Without these, any model selection is premature.
|
||||
|
||||
PROPOSAL — Tentative default pending more repository context:
|
||||
|
||||
1. **Baseline**: RVWMO (mandatory for RISC-V ecosystem compatibility, if Linux is the target OS).
|
||||
2. **Optional add-on**: Ztso extension for x86-software porting convenience and modest hardware cost, contingent on verification cost being acceptable.
|
||||
3. **Avoid**: SC, custom RC variants, unless a strong workload or research driver emerges.
|
||||
4. **Defer to**: `cache-coherency.md` for the coherence-side ordering mechanism.
|
||||
|
||||
## Confidence
|
||||
|
||||
| Topic | Confidence | Reason |
|
||||
|-------|------------|--------|
|
||||
| RVWMO is the RISC-V baseline | High | Ratified ISA spec |
|
||||
| Ztso is a small optional extension | High | Ratified extension |
|
||||
| SC has high implementation cost | High | General literature consensus |
|
||||
| SC specifically impractical at 128 cores | Low | General claim, not XH-1-measured; workload- and topology-dependent |
|
||||
| Per-fence energy scales with acknowledging-core count | Medium | Logical consequence; no measurement |
|
||||
| Quantitative per-fence energy at 128 cores | None | INSUFFICIENT EVIDENCE |
|
||||
| Verification is hardest for RC variants | Low–Medium | Literature claim, not XH-1-measured |
|
||||
| XH-1 should adopt RVWMO | Low | Repository context not yet sufficient |
|
||||
| XH-1 should adopt RVWMO + Ztso | Low | Same |
|
||||
|
||||
## Open Questions
|
||||
|
||||
1. What is XH-1's microarchitecture style (in-order vs. out-of-order)?
|
||||
2. What coherence protocol and directory organization will XH-1 use?
|
||||
3. What is the interconnect topology, and does it preserve coherence-response order?
|
||||
4. What workloads are targeted — HPC, server, embedded?
|
||||
5. Does XH-1 include non-coherent agents (DMA, CXL, accelerators)?
|
||||
6. Is XH-1 a research vehicle, a prototype, or a product?
|
||||
7. What is the expected fence frequency in target software?
|
||||
8. What is the verification methodology — formal, simulation, or hybrid?
|
||||
9. Will XH-1 use the standard RISC-V litmus-test suite or define a custom one?
|
||||
10. Does XH-1 require compatibility with x86-software-only ports? (Drives Ztso inclusion.)
|
||||
11. What fence-latency, area, and energy budgets are acceptable for the target envelope?
|
||||
12. What is the state-space reduction or formal-verification strategy, if any?
|
||||
|
||||
## Sources
|
||||
|
||||
INSUFFICIENT EVIDENCE — The XH-1 Research repository does not yet contain source material specific to memory ordering. No citations, papers, or measurements can be honestly attributed to XH-1.
|
||||
|
||||
External ISA-level facts cited above (RVWMO, Ztso, FENCE instruction semantics) are drawn from the publicly ratified RISC-V ISA specification, which is the source of truth for the ISA-level memory model. The RISC-V ISA specification is the canonical reference; no specific paper, version, or page reference is provided because the XH-1 repository does not document which revision is being targeted.
|
||||
|
||||
INSUFFICIENT EVIDENCE — All quantitative claims about fence latency, area, power, energy, and verification cost at the XH-1 level have been removed or flagged. No benchmarks, no measurements, no RTL data, no simulation results are available in the repository at this time.
|
||||
+153
File diff suppressed because one or more lines are too long
@@ -0,0 +1,333 @@
|
||||
# Memory Ordering
|
||||
|
||||
## Status
|
||||
|
||||
Stub document. XH-1 memory ordering model has not yet been defined. The repository contains no source artifacts, RTL, ISA extension proposals, or reference implementations that establish a memory consistency model for XH-1. All quantitative or XH-1-specific claims in this document are flagged INSUFFICIENT EVIDENCE where applicable.
|
||||
|
||||
## Abstract
|
||||
|
||||
Memory ordering defines the rules by which loads and stores issued by one or more cores become visible to other cores in the system. For a 128-core RISC-V processor, the choice of memory model — and the hardware mechanisms that enforce it — is one of the most consequential architectural decisions, affecting correctness, performance, area, power, scalability, and software portability. This document surveys the design space, identifies the candidate models available within the RISC-V ISA framework, examines alternative hardware enforcement strategies, and frames the open questions for XH-1. No final recommendation is made because the repository does not establish pipeline style, coherency protocol, or interconnect topology, all of which constrain the memory ordering decision. A conditional tentative default is stated, contingent on the assumed constraints being correct.
|
||||
|
||||
## Research Question
|
||||
|
||||
What memory consistency model and hardware ordering mechanism should XH-1 adopt, given:
|
||||
|
||||
1. A 128-core replication target
|
||||
2. The RISC-V ISA baseline
|
||||
3. The need to balance software portability against hardware implementation cost
|
||||
4. Unknown microarchitecture style (in-order vs. out-of-order is not established in the repository)
|
||||
5. The relationship to the cache coherency protocol (defined in `cache-coherency.md`, currently SOON)
|
||||
|
||||
## Background
|
||||
|
||||
### Memory Consistency Models — Terminology
|
||||
|
||||
PROPOSAL — Use the following working taxonomy in this document:
|
||||
|
||||
- **Sequential Consistency (SC)**: All cores observe a single total order of memory operations that respects each core's program order.
|
||||
- **Total Store Order (TSO)**: Loads may bypass stores from the same core; stores from one core are seen in order by all cores. Used by x86.
|
||||
- **Relaxed / Weak Memory Models**: Many orderings are allowed; fences (barriers) are required to enforce constraints. Used by ARM, RISC-V, and PowerPC.
|
||||
- **Release Consistency (RC) / RCsc / RCpc / RCO**: A family of models with different fence semantics. Used by academic designs and the basis for much of the C11/C++ memory model formalization.
|
||||
|
||||
ASSUMPTION — This document uses "RC variants" as an umbrella term for RCsc, RCpc, and RCO. The specific semantics of each variant are not redefined here; readers are referred to the literature.
|
||||
|
||||
### The RISC-V "RVWMO" Model
|
||||
|
||||
FACT — The RISC-V ISA specification defines the **RISC-V Weak Memory Ordering (RVWMO)** model as the baseline. Under RVWMO:
|
||||
|
||||
- Loads and stores to the same address from the same hart are ordered.
|
||||
- A store followed by a load to a *different* address may be reordered.
|
||||
- The `FENCE` instruction is used to enforce orderings. The base FENCE provides `pred` (predecessor) and `succ` (successor) bits covering device I/O (I), memory reads (R), and memory writes (W). Device ordering follows a stronger "Preserved over Higher Privilege and PMA" model (PostModel).
|
||||
- The optional `Ztso` extension provides **RVTSO**, which gives x86-like TSO semantics with a relaxed form of the FENCE.
|
||||
|
||||
FACT — RVWMO and the `Ztso` extension are documented in the ratified RISC-V Unprivileged ISA specification. The repository does not state which revision of the specification is the reference for XH-1; this is an open question.
|
||||
|
||||
ASSUMPTION — RVWMO + the Ztso optional extension is the available ISA-level model space that XH-1 must select from. (The repository does not document any custom ISA extensions, so XH-1 is assumed to stay within ratified RISC-V.)
|
||||
|
||||
### Why Memory Ordering is Hard at 128 Cores
|
||||
|
||||
OPEN QUESTION — The repository does not establish:
|
||||
- Whether XH-1 is a single die, multi-die package, or chiplet-based system
|
||||
- The interconnect topology (mesh, ring, crossbar, NoC, etc.)
|
||||
- Whether coherence is directory-based, broadcast-based, or hybrid
|
||||
- The private-vs-shared L2/L3 split
|
||||
- The memory-side directory or snoop filter organization
|
||||
|
||||
These decisions have direct bearing on which ordering mechanisms are feasible. Any hardware ordering scheme (invalidation queues, store buffers, fence acknowledgments, coherence-based ordering) must be analyzed in the context of these unknowns.
|
||||
|
||||
## Existing Approaches
|
||||
|
||||
### Approach A — RISC-V RVWMO (Baseline)
|
||||
|
||||
The standard weak memory model provided by the RISC-V ISA. Requires implementation of `FENCE`, `FENCE.I`, and the I/O ordering semantics.
|
||||
|
||||
**Mechanism**: Architecture permits aggressive reordering. Hardware implementations commonly use:
|
||||
- Store buffers with forwarding
|
||||
- Load-load reordering queues
|
||||
- Coherence-based ordering: stores become visible at coherence invalidation/upgrade points
|
||||
- Fence implementation via global acknowledgment counters or by draining reorder queues
|
||||
|
||||
INSUFFICIENT EVIDENCE — Which of these mechanisms XH-1 will use is not established by the repository.
|
||||
|
||||
### Approach B — RVWMO + Ztso Extension
|
||||
|
||||
Adds RVTSO. Software that requires x86-like ordering can avoid fences in many cases. Still RISC-V compliant.
|
||||
|
||||
**Mechanism**: Stores from one hart are serialized in the interconnect; loads may still pass earlier stores to different addresses.
|
||||
|
||||
### Approach C — Sequential Consistency (SC)
|
||||
|
||||
All operations observed in a single global order respecting program order. Simplest software model; most expensive hardware in general.
|
||||
|
||||
**Mechanism**: Commonly implemented via store-buffer draining on loads, broadcast-based coherence, and conservative replay mechanisms. The qualitative characterization of "very costly" is a general literature claim, not an XH-1 measurement; whether SC is specifically costly at 128 cores is workload- and topology-dependent and is treated here as a hypothesis, not a known result.
|
||||
|
||||
### Approach D — Release Consistency variants (RCsc, RCpc, RCO)
|
||||
|
||||
Academic models where release/acquire ordering is cheap but stronger fences are required for full memory ordering. Basis of C11/C++ model.
|
||||
|
||||
**Mechanism**: Only release/acquire fences require acknowledgment; acquire-only loads are cheap. Alignment with C11 is a strong software benefit. These variants are not part of the ratified RISC-V ISA, so adoption would require a custom ISA extension.
|
||||
|
||||
## Alternative Designs
|
||||
|
||||
### Hardware Ordering Mechanisms
|
||||
|
||||
PROPOSAL — The following mechanisms should be considered for fence/ordering enforcement once coherency and interconnect are known:
|
||||
|
||||
1. **Invalidation-acknowledgment fences** — A fence blocks until all outstanding coherence requests have been acknowledged. Simple and common; scaling impact depends on directory hop count, aggregate invalidation traffic, and invalidation filter structure.
|
||||
|
||||
2. **Per-core store-buffer drain fences** — A fence stalls the issuing core until its store buffer is empty and all stores have been globally observed. Predictable latency; serializes through the coherence fabric.
|
||||
|
||||
3. **Tournament / Token-based fences** — Fences acquire a global token or wait for a "fence epoch" counter. Avoids worst-case drain, but adds global state.
|
||||
|
||||
4. **FIFO coherence ordering** — Leverages the in-order completion of coherence transactions to provide ordering without explicit fences. Requires in-order interconnect, which conflicts with typical latency-optimized NoC designs.
|
||||
|
||||
5. **Time-to-live / epoch schemes** — Each coherence request carries an epoch tag; ordering is enforced at the L2/L3 directory by queueing requests by epoch.
|
||||
|
||||
6. **Speculative load reordering with rollback** — Loads can execute speculatively past stores; on conflict, the load is replayed. Common in high-performance out-of-order cores (e.g., reported in the IBM POWER literature).
|
||||
|
||||
ASSUMPTION — XH-1's microarchitecture style is not yet established. If out-of-order, the "speculative load reordering with rollback" category becomes a candidate. If in-order, mechanism choices narrow considerably.
|
||||
|
||||
## Comparison
|
||||
|
||||
| Property | SC | RVWMO | RVWMO + Ztso | RC variants |
|
||||
|----------|----|----|----|----|
|
||||
| Software ease | Highest | Medium | Medium-High | Low-Medium |
|
||||
| Hardware cost (general) | Highest | Low | Low-Medium | Low |
|
||||
| Performance headroom (general) | Lowest | Highest | High | High |
|
||||
| Fence latency (general) | N/A | Variable | Variable | Variable |
|
||||
| Verification cost (general) | Lowest in principle | High | High | High |
|
||||
| C11/C++ alignment | Approximate | Approximate | Weak | Direct |
|
||||
| 128-core scalability | Poor (hypothesis) | Unmeasured | Unmeasured | Unmeasured |
|
||||
|
||||
INSUFFICIENT EVIDENCE — Quantitative fence-latency, coherence-traffic, and scalability numbers for XH-1 are unavailable. Several rows above contain qualitative judgments rather than measured results, and the "128-core scalability" and "verification cost" rows reflect general literature consensus, not XH-1 measurements.
|
||||
|
||||
## Advantages
|
||||
|
||||
### RVWMO (Baseline)
|
||||
- **Lowest hardware cost** of all candidates for the base case, as a general property of the model — most orderings are simply not enforced.
|
||||
- **Maximum performance** potential, as a general property — reordering is unconstrained.
|
||||
- **Familiar** to the RISC-V software ecosystem (Linux, RISC-V GCC, LLVM), subject to the open question of whether Linux is the target OS.
|
||||
- **Standardized** — concrete compliance test suites exist externally.
|
||||
|
||||
### RVWMO + Ztso
|
||||
- **x86-like TSO** for software that benefits from stronger ordering.
|
||||
- **Backwards compatible** — software not using Ztso still runs correctly under RVWMO.
|
||||
- **Modest hardware cost** — primarily requires store-store ordering at the coherence layer (qualitative claim, not measured for XH-1).
|
||||
|
||||
### Release Consistency Variants
|
||||
- **Excellent fit for C11/C++** atomics — acquire/release are the most-used fence types in real code, per the C/C++ standards.
|
||||
- **Cheapest common case** — acquire and release can be implemented with lightweight mechanisms, as a general property.
|
||||
- **Hardware cost is concentrated on the rare case** of full seq_cst fences, as a general property.
|
||||
|
||||
## Disadvantages
|
||||
|
||||
### RVWMO (Baseline)
|
||||
- **Fence latency depends on coherence round-trip** — at 128 cores, fence latency is expected to be high because global acknowledgment must traverse the coherence fabric. Specific cycle counts: INSUFFICIENT EVIDENCE.
|
||||
- **Verification complexity is high** — many legal reorderings; the litmus-test failure surface is large. This is a general property of weak models, not an XH-1 measurement.
|
||||
- **Software burden** — kernel and runtime code must insert fences correctly.
|
||||
|
||||
### RVWMO + Ztso
|
||||
- **Two consistency models in one chip** — adds documentation, validation, and software education cost.
|
||||
- **In-flight mixing** of ordering regimes in the same software is hard to reason about.
|
||||
|
||||
### Release Consistency Variants
|
||||
- **Non-standard for RISC-V** — would require a custom ISA extension. Cannot be recommended without a strong software-side driver.
|
||||
- **Verification cost is high** in general because of subtle fence semantics; whether it is the "highest" among candidates is a claim about the literature, not a measured XH-1 result.
|
||||
|
||||
### Sequential Consistency
|
||||
- **Performance cost is severe in the general case** — every load may stall on store-buffer drain; whether this is "severe at 128 cores" specifically is workload- and topology-dependent and is not established by the repository.
|
||||
- **No selective escape valve** — software cannot opt out of the strongest model.
|
||||
- **Does not match RISC-V ecosystem expectations** — surprising to RISC-V software developers, as a general observation.
|
||||
|
||||
## XH-1 Considerations
|
||||
|
||||
OPEN QUESTION — The following XH-1-internal questions are unresolved by the repository:
|
||||
|
||||
1. Is XH-1's pipeline in-order or out-of-order?
|
||||
2. What is the coherence protocol (MOESI, MESI, directory-based, broadcast)?
|
||||
3. Is the interconnect a NoC, ring, or crossbar? Does it preserve in-order delivery of coherence responses?
|
||||
4. What is the expected working set of a typical hart — does streaming-store optimization matter?
|
||||
5. Does XH-1 target HPC, server, embedded, or mixed workloads? (Workload affects fence frequency.)
|
||||
6. Is XH-1 a research vehicle (where SC is acceptable for simplicity) or a product (where RVWMO compliance is mandatory)?
|
||||
7. Does the memory map use non-coherent regions (e.g., DMA, I/O) that interact with the ordering model?
|
||||
|
||||
ASSUMPTION — If XH-1 targets RISC-V ecosystem compatibility (Linux, RISC-V GCC, RISC-V LLVM), then RVWMO is mandatory. Ztso may optionally be added for x86-software porting convenience.
|
||||
|
||||
## 128-Core Scaling Considerations
|
||||
|
||||
INSUFFICIENT EVIDENCE — Quantitative figures for XH-1 at 128 cores are unavailable. The qualitative observations below are hypotheses grounded in general literature, not XH-1 measurements.
|
||||
|
||||
### Fence Latency
|
||||
- **Fence global acknowledgment** latency is bounded below by the worst-case coherence round-trip time. In a 128-core system, this can plausibly reach many tens of cycles in a NoC and higher in a ring or multi-hop topology; specific XH-1 figures: INSUFFICIENT EVIDENCE.
|
||||
- **Store-buffer drain** latency is approximately proportional to the store buffer depth and the time to invalidate all sharers. Specific XH-1 figures: INSUFFICIENT EVIDENCE.
|
||||
- **Token / epoch fences** may scale better because they avoid draining the entire coherence fabric — they synchronize at a logical epoch boundary. This is a general property, not an XH-1 measurement.
|
||||
|
||||
### Coherence Traffic
|
||||
- **RVWMO fences** generate invalidation or acknowledgment traffic that competes with normal coherence traffic. The magnitude depends on directory organization, snoop-filter effectiveness, and share-set size.
|
||||
- **SC at 128 cores** can plausibly saturate a directory-based interconnect because every load's coherence transaction may stall pending store ordering; this is a hypothesis, not a measured XH-1 result.
|
||||
- **RC variants** concentrate traffic on the fence path only, as a general property.
|
||||
|
||||
### Verification Scalability
|
||||
- **State-space explosion**: at 128 cores, the number of concurrently observable memory operations grows combinatorially in the unmitigated model. Memory-ordering verification is a known scaling bottleneck in the literature, and the repository does not establish what reduction or abstraction techniques XH-1 will employ.
|
||||
- **RC variants** are commonly cited as among the hardest to verify because of subtle fence interactions; whether they are hardest in the absolute sense is a literature claim, not an XH-1 measurement.
|
||||
- **SC is easiest to verify in principle** because it admits fewer legal behaviors, but worst in performance.
|
||||
|
||||
INSUFFICIENT EVIDENCE — No quantitative fence-latency, coherence-traffic, or verification-scaling data is available for XH-1. Estimates cannot be made without knowing the interconnect, coherence protocol, and verification methodology.
|
||||
|
||||
## Performance Considerations
|
||||
|
||||
PROPOSAL — Performance-impact dimensions to evaluate once workload and microarchitecture are defined:
|
||||
|
||||
1. **Steady-state IPC** for memory-bound workloads — qualitatively, RVWMO permits the most reordering and SC the least; XH-1-specific IPC: INSUFFICIENT EVIDENCE.
|
||||
2. **Fence latency** — measured as cycles from fence issue to fence completion. Dominant for synchronization-heavy code. XH-1-specific: INSUFFICIENT EVIDENCE.
|
||||
3. **Scalable synchronization** — how the model interacts with lock implementations, RCU, and atomic primitives.
|
||||
4. **Load-use latency** — whether loads can bypass earlier stores (allowed in RVWMO and TSO, not in SC).
|
||||
5. **Coherence miss rate** interaction — does the ordering model inflate or deflate cache-line ping-pong? Model-dependent; XH-1-specific: INSUFFICIENT EVIDENCE.
|
||||
|
||||
ASSUMPTION — For HPC workloads with high-fence frequency (e.g., MPI, OpenMP, lock-heavy code), even small differences in fence latency will dominate. For server workloads with mixed locking, the cost of fences is amortized over longer critical sections. No XH-1 workload profile is documented in the repository.
|
||||
|
||||
## Area Considerations
|
||||
|
||||
INSUFFICIENT EVIDENCE — All area descriptions in this section are qualitative. Precise XH-1 area figures (in gates, µm², or mm²) are unavailable; no synthesis, layout, or RTL data exists in the repository.
|
||||
|
||||
PROPOSAL — Qualitative area cost of ordering mechanisms, as a function of microarchitectural choices:
|
||||
|
||||
| Mechanism | Approximate area cost per core (qualitative) |
|
||||
|-----------|----------------------------------------------|
|
||||
| Store buffer with forwarding | Small (few entries × cache-line width) |
|
||||
| Load-load reordering queue | Small to medium (grows with IPC and reorder depth) |
|
||||
| Fence-acknowledgment counter | Negligible (per-core counter) |
|
||||
| Global epoch / token state | Medium (one per chip, not per core) |
|
||||
| Speculative load reorder + rollback | Large (load-load queue, replay logic) |
|
||||
|
||||
PROPOSAL — At 128 cores, per-core area costs are replicated 128 times. Global state (epoch counters, fence acknowledgers) is O(1) per chip and therefore amortizes favorably on a per-core basis, but the critical-path and clock-distribution impact of global state is not established by the repository.
|
||||
|
||||
## Power and Energy Considerations
|
||||
|
||||
INSUFFICIENT EVIDENCE — No per-fence, per-core, or aggregate energy figures for XH-1 are available. The earlier version of this document included a "100 pJ-per-core fence" figure that was not substantiated by any repository artifact or external citation; that figure has been removed.
|
||||
|
||||
PROPOSAL — Qualitative energy implications, as general properties of the model classes:
|
||||
|
||||
- **Fences** generate global coherence traffic — energy cost scales with the number of cores that must acknowledge. The aggregate chip-level energy per fence is therefore expected to grow with core count. Quantitative XH-1 values: INSUFFICIENT EVIDENCE.
|
||||
- **Speculative load reordering** wastes energy on rollback when a violation is detected, as a general property.
|
||||
- **Strong ordering models** keep more coherence state in flight (more invalidations, more retries), as a general property.
|
||||
- **RVWMO** with lightweight fences (when software permits) is the lowest-energy regime in the general case.
|
||||
|
||||
## Implementation Considerations
|
||||
|
||||
PROPOSAL — Implementation order of dependencies (the repository does not establish these, but the research area suggests the following):
|
||||
|
||||
1. Define coherency protocol → `cache-coherency.md`
|
||||
2. Define interconnect topology → not yet in repository
|
||||
3. Define memory hierarchy → `memory-hierarchy.md`
|
||||
4. Define memory ordering model → this document
|
||||
5. Define atomic primitives → `atomics.md`
|
||||
|
||||
ASSUMPTION — The Ztso extension is a small RTL delta over RVWMO. The dominant implementation cost is verification, not area. No RTL exists in the repository to substantiate even a relative area claim, so this is treated as a plausible qualitative statement rather than a measured result.
|
||||
|
||||
OPEN QUESTION — Does XH-1 have a coherent accelerator fabric or non-coherent IO (e.g., CXL, DMA engines)? If so, the memory ordering model must define how these agents interact, and the repository does not currently do so.
|
||||
|
||||
## Verification Considerations
|
||||
|
||||
PROPOSAL — Verification challenges:
|
||||
|
||||
- **Litmus tests**: A standard RISC-V litmus test suite is available externally; whether XH-1 will adopt it as-is or define a custom suite is OPEN.
|
||||
- **State-space explosion**: 128 cores × N outstanding operations produces a combinatorial unmitigated state space. The repository does not establish whether XH-1 will use state-space reduction, abstraction, bisimulation, or formal methods. Any specific XH-1 claim: INSUFFICIENT EVIDENCE.
|
||||
- **Fence semantics**: Each fence variant must be exhaustively tested at corner cases (interrupts, exceptions, MMIO).
|
||||
- **Coherence-ordering interaction**: The ordering model must be checked against the chosen coherence protocol for consistency. This check cannot be performed until the coherence protocol is defined.
|
||||
|
||||
INSUFFICIENT EVIDENCE — No XH-1 verification infrastructure, formal spec, or litmus-test set is documented in the repository.
|
||||
|
||||
ASSUMPTION — RVWMO is the most documented and tested model for RISC-V, per the ratified specification and ecosystem practice. Ztso is a small addition. SC and RC variants would require custom formal infrastructure. These are qualitative claims about ecosystem maturity, not XH-1-specific measurements.
|
||||
|
||||
## Software Considerations
|
||||
|
||||
PROPOSAL — Software-side impact, as a general property of each model:
|
||||
|
||||
- **Linux kernel**: Designed for weak memory models with appropriate fences. Compatible with RVWMO.
|
||||
- **C11/C++ atomics**: Map cleanly onto release/acquire semantics. Compatible with RVWMO and with RC variants.
|
||||
- **x86 software ports**: Benefit from Ztso for fewer fences.
|
||||
- **OpenSHMEM, MPI shmem**: Often assume TSO-like ordering; Ztso helps.
|
||||
- **HPC codes with hand-rolled atomics**: Highly sensitive to fence latency.
|
||||
|
||||
ASSUMPTION — XH-1 will run RISC-V Linux. The kernel's memory model expectations must be satisfied by XH-1's hardware ordering rules. The repository does not document whether Linux is in fact the target.
|
||||
|
||||
## Recommendation
|
||||
|
||||
INSUFFICIENT EVIDENCE to make a final recommendation.
|
||||
|
||||
The repository does not yet establish:
|
||||
- Pipeline microarchitecture
|
||||
- Coherence protocol
|
||||
- Interconnect topology
|
||||
- Workload target
|
||||
|
||||
A recommendation can be made only after the following documents in the memory research area are filled in: `memory-architecture.md`, `cache-coherency.md`, `memory-hierarchy.md`. Without these, any model selection is premature.
|
||||
|
||||
PROPOSAL — Tentative default, contingent on the stated assumptions holding and explicitly subject to revision once repository context is available:
|
||||
|
||||
1. **Baseline**: RVWMO (assumed mandatory for RISC-V ecosystem compatibility, if Linux is the target OS).
|
||||
2. **Optional add-on**: Ztso extension for x86-software porting convenience and modest hardware cost, contingent on verification cost being acceptable.
|
||||
3. **Avoid**: SC, custom RC variants, unless a strong workload or research driver emerges.
|
||||
4. **Defer to**: `cache-coherency.md` for the coherence-side ordering mechanism.
|
||||
|
||||
ASSUMPTION — The tentative default above depends on (a) Linux being the target OS, (b) ecosystem compatibility being prioritized over x86-port convenience, and (c) the coherence and interconnect unknowns being resolvable without invalidating this default. None of these are established by the repository.
|
||||
|
||||
## Confidence
|
||||
|
||||
| Topic | Confidence | Reason |
|
||||
|-------|------------|--------|
|
||||
| RVWMO is the RISC-V baseline | High | Ratified ISA specification |
|
||||
| Ztso is a small optional extension | High | Ratified extension in the ISA specification |
|
||||
| SC has high implementation cost in general | High | General literature consensus |
|
||||
| SC specifically impractical at 128 cores | Low | General claim, not XH-1-measured; workload- and topology-dependent |
|
||||
| Per-fence energy scales with acknowledging-core count | Medium | Logical consequence; no measurement |
|
||||
| Quantitative per-fence energy at 128 cores | None | INSUFFICIENT EVIDENCE |
|
||||
| Verification is hardest for RC variants | Low–Medium | Literature claim, not XH-1-measured |
|
||||
| XH-1 should adopt RVWMO | Low | Repository context not yet sufficient |
|
||||
| XH-1 should adopt RVWMO + Ztso | Low | Same |
|
||||
| Ztso is a small RTL delta over RVWMO | Low–Medium | Plausible qualitative claim; no RTL to substantiate |
|
||||
|
||||
## Open Questions
|
||||
|
||||
1. What is XH-1's microarchitecture style (in-order vs. out-of-order)?
|
||||
2. What coherence protocol and directory organization will XH-1 use?
|
||||
3. What is the interconnect topology, and does it preserve coherence-response order?
|
||||
4. What workloads are targeted — HPC, server, embedded?
|
||||
5. Does XH-1 include non-coherent agents (DMA, CXL, accelerators)?
|
||||
6. Is XH-1 a research vehicle, a prototype, or a product?
|
||||
7. What is the expected fence frequency in target software?
|
||||
8. What is the verification methodology — formal, simulation, or hybrid?
|
||||
9. Will XH-1 use the standard RISC-V litmus-test suite or define a custom one?
|
||||
10. Does XH-1 require compatibility with x86-software-only ports? (Drives Ztso inclusion.)
|
||||
11. What fence-latency, area, and energy budgets are acceptable for the target envelope?
|
||||
12. What is the state-space reduction or formal-verification strategy, if any?
|
||||
13. Which revision of the RISC-V Unprivileged ISA specification is the reference for XH-1?
|
||||
|
||||
## Sources
|
||||
|
||||
INSUFFICIENT EVIDENCE — The XH-1 Research repository does not yet contain source material specific to memory ordering. No citations, papers, or measurements can be honestly attributed to XH-1.
|
||||
|
||||
External ISA-level facts cited above (RVWMO, Ztso, FENCE instruction semantics) are drawn from the publicly ratified RISC-V Unprivileged ISA specification, which is the source of truth for the ISA-level memory model. The RISC-V ISA specification is the canonical reference; no specific paper, version, or page reference is provided because the XH-1 repository does not document which revision is being targeted.
|
||||
|
||||
INSUFFICIENT EVIDENCE — All quantitative claims about fence latency, area, power, energy, and verification cost at the XH-1 level have been removed or flagged. No benchmarks, no measurements, no RTL data, no simulation results are available in the repository at this time.
|
||||
+173
File diff suppressed because one or more lines are too long
@@ -0,0 +1,2 @@
|
||||
2026-08-26T12:27:36Z research/05-memory/memory-ordering.md 1 research success
|
||||
2026-08-26T12:30:29Z research/05-memory/memory-ordering.md 1 review invalid
|
||||
@@ -0,0 +1,329 @@
|
||||
# Memory Ordering
|
||||
|
||||
**Status:** Research document — proposal stage. No XH-1 implementation is fixed.
|
||||
**Scope:** Memory consistency model, ordering primitives, and their interaction with a 128-core RISC-V processor.
|
||||
**Audience:** XH-1 architecture team, verification engineers, OS porters.
|
||||
|
||||
---
|
||||
|
||||
## 1. Overview and Definitions
|
||||
|
||||
Memory ordering defines the rules by which the memory accesses (loads and stores) issued by one or more cores become visible, in program order and globally, to other cores. It determines:
|
||||
|
||||
- **Program order (PO):** the order in which a single hart issues memory accesses.
|
||||
- **Memory order (MO):** the global order in which accesses appear to execute as observed by other harts.
|
||||
- **Coherence order (CO):** the per-address total order of accesses required by the cache coherence protocol.
|
||||
- **Fence ordering:** explicit constraints inserted by hardware or software to restrict reordering.
|
||||
|
||||
These concepts are foundational to any multiprocessor memory model and are defined formally in the literature (Adve & Gharachorloo, 1996; Sorin, Hill & Wood, *A Primer on Memory Consistency and Cache Coherence*, 2011).
|
||||
|
||||
XH-1 must adopt and document a precise model so that:
|
||||
|
||||
1. Hardware designers know which reorderings are legal.
|
||||
2. The operating-system port (and any hypervisor) knows what atomicity guarantees it can rely on.
|
||||
3. Application and runtime software (compilers, language runtimes, synchronization libraries) can map higher-level primitives correctly.
|
||||
4. The verification team can construct litmus tests and formal models against an unambiguous specification.
|
||||
|
||||
---
|
||||
|
||||
## 2. RISC-V Architectural Requirements
|
||||
|
||||
### 2.1 The RISC-V Weak Memory Model (RVWMM)
|
||||
|
||||
RISC-V defines its memory model in:
|
||||
|
||||
> *The RISC-V Instruction Set Manual, Volume I: Unprivileged Architecture*, Document Version 20191213 (and later), §Appendix A — "RVWMM: RISC-V Weak Memory Model."
|
||||
|
||||
Key facts established by that specification (status: **RISC-V specification requirement**):
|
||||
|
||||
- RISC-V adopts a **release consistency** model (a variant of "weak ordering" with explicit acquire/release semantics), broadly modeled on the approach taken by Armv8.
|
||||
- Ordinary load and store instructions are divided by their **addressing mode** and **access type** into:
|
||||
- `.aq` — acquire semantics (subsequent accesses in program order cannot be reordered before this access).
|
||||
- `.rl` — release semantics (prior accesses in program order cannot be reordered after this access).
|
||||
- The pair `aq+rl` yields sequential consistency (SC) per access when used on both sides of a synchronization pair.
|
||||
- Loads use acquire semantics (`LR`/`.aq` on AMOs and `LR.aq`); stores use release semantics (`.rl`).
|
||||
- `FENCE` instructions (including `FENCE.TSO`, `FENCE.I`, `FENCE.VMA`, `FENCE.GORC`, `FENCE.GW`, `FENCE.OW`, `FENCE.ORW`, `FENCE.RW`, `FENCE.RR`, `FENCE.WW`) provide explicit ordering at the hart front-end and the global memory system.
|
||||
- `FENCE.RW,RW` (the legacy full fence) orders all prior loads and stores against all subsequent loads and stores at the local hart and at the global system.
|
||||
- The model is **multi-copy atomic (MCA)** for devices that respect it, but **RISC-V explicitly does not require main memory to be multi-copy atomic** in general; I/O regions may opt into TSO or stronger ordering via `FENCE` and `aq/rl` annotations.
|
||||
|
||||
This is the model the XH-1 must conform to for RV64 compliance; deviations are not permitted for standard harts without explicit non-standard attributes.
|
||||
|
||||
### 2.2 Atomic Memory Operations (AMOs)
|
||||
|
||||
RISC-V provides:
|
||||
|
||||
- **Load-Reserved / Store-Conditional (LR/SC):** `LR.W/D`, `SC.W/D`, optionally with `.aq`/`.rl`/`.aqrl`.
|
||||
- **AMO** instructions: `AMOSWAP`, `AMOADD`, `AMOAND`, `AMOOR`, `AMOXOR`, `AMOMAX[U]`, `AMOMIN[U]` on 32- and 64-bit values.
|
||||
|
||||
All AMOs implicitly perform a load and a store, and the addressing mode applies to the entire RMW operation.
|
||||
|
||||
### 2.3 Reservation Sets and Forward Progress
|
||||
|
||||
LR/SC semantics on RISC-V require that:
|
||||
|
||||
- The reservation set be **at least as large as** the address and **contiguous in physical memory**, with forward-progress guarantees (no livelock).
|
||||
- The implementation may break a reservation for any reason, but must make progress under certain conditions (the exact rules are stated in the spec).
|
||||
|
||||
This matters for XH-1's L1 coherence protocol design — we must specify how reservation sets map onto cache lines and which events cause invalidation.
|
||||
|
||||
### 2.4 I/O Ordering
|
||||
|
||||
For Memory-Mapped I/O (MMIO) and DMA, RISC-V provides `FENCE.I` for self-modifying code, and the I/O map can be marked as a **non-idempotent, strongly ordered** region via PMA (Physical Memory Attributes). XH-1 must define its PMAs and the interaction with the cache hierarchy carefully; this is treated in `research/05-memory/cache-coherence.md` and `research/06-io/`.
|
||||
|
||||
---
|
||||
|
||||
## 3. The 128-Core Scaling Problem
|
||||
|
||||
XH-1 has 128 application cores. This dramatically amplifies the cost of any ordering primitive that requires global visibility.
|
||||
|
||||
### 3.1 Why Fences Are Expensive
|
||||
|
||||
A "global" fence (e.g., `FENCE.RW,RW` that orders against the entire memory system) must wait for:
|
||||
|
||||
- All outstanding transactions from the issuing core to be **globally observed** (or at least ordered).
|
||||
- The interconnect to drain or acknowledge.
|
||||
|
||||
At 128 cores, global observation requires one of:
|
||||
|
||||
- A **global counter** that is incremented on every coherence transaction and broadcast on the interconnect, or
|
||||
- A **token-based scheme** such as the Token Coherence (Marty et al., 2005) approach, or
|
||||
- **Hierarchical fences** that propagate down the directory tree.
|
||||
|
||||
A naive "drain the interconnect" fence scales with O(N) (N cores) latency; in practice, on a 128-core mesh this may exceed 100s of ns — sufficient to dominate synchronization hot paths. *This is an estimate; actual latency depends on interconnect topology, link width, and clock frequency — currently unresolved for XH-1.*
|
||||
|
||||
### 3.2 The Scalability of Atomics
|
||||
|
||||
AMO and SC RMW operations require **directory invalidations** to all sharers and an **acknowledgment** before completion. At 128 cores:
|
||||
|
||||
- Worst-case invalidation fanout is bounded by the number of cores (128), but typical working sets are much smaller.
|
||||
- **Fairness** and **starvation** become real concerns: a hot lock variable contested by 128 cores will generate enormous traffic. The hardware must not let one core perpetually lose a `SC` against another.
|
||||
- **NUMA effects**: if cores are organized in tiles/quads/clusters with distributed directories, an AMO on a remote line costs an extra hop in latency.
|
||||
|
||||
### 3.3 False Sharing and Coherence Traffic
|
||||
|
||||
A weak memory model encourages fine-grained synchronization, but a 128-core system makes **false sharing** (two cores writing to different bytes of the same line) catastrophically expensive — each write causes an invalidation round-trip. The memory model itself doesn't fix this, but it influences which data structures are used.
|
||||
|
||||
### 3.4 Memory-Model Litmus Test Complexity
|
||||
|
||||
For verification, the number of distinct observable outcomes of small litmus tests grows combinatorially with core count. The team must be able to run RVWMM test suites (e.g., the official `riscv-tests` and `litmus` suites from Cambridge) and interpret results at 128 cores.
|
||||
|
||||
---
|
||||
|
||||
## 4. Implementation Approaches and Alternatives
|
||||
|
||||
This section discusses candidate mechanisms. **None of these are XH-1 commitments; all are proposals awaiting design review.**
|
||||
|
||||
### 4.1 Fence Implementation Strategies
|
||||
|
||||
| Strategy | Description | Pros | Cons |
|
||||
|----------|-------------|------|------|
|
||||
| **Naïve global fence** | Drain the issue queue, snoop outstanding transactions, wait for all acks. | Simple; obviously correct. | O(N) latency; does not scale to 128 cores. |
|
||||
| **FIFO completion tags (epoch counters)** | Each request gets a tag; fences compare local tag against a per-core "completed" tag. | Local fences are cheap; correctness well understood. | Global fences still need a global view. |
|
||||
| **Hierarchical directory-assisted fences** | Fences consult the directory tree to determine "when is my request globally ordered?" | Latency proportional to tree depth rather than N. | Requires directory tree to maintain completion info. |
|
||||
| **Token-based fences** | Borrow from Token Coherence (Marty et al., 2005) — fences consume tokens to know when stores are visible. | Distributed; no single global counter. | Token bookkeeping overhead; complex verification. |
|
||||
| **TSO hole** | Implement most of memory as TSO with weak operations for device access only (similar to x86). | Strong ordering for free for ordinary code; simpler programmer model. | Performance penalty for non-synchronization code; deviates from RVWMM. |
|
||||
|
||||
**Note:** RISC-V explicitly disallows the TSO-only approach for general memory because the RVWMM is the architectural contract. Any TSO shortcut must be invisible to software.
|
||||
|
||||
### 4.2 Store-Buffer and Load-Queue Design
|
||||
|
||||
Each core needs:
|
||||
|
||||
- **Store Buffer (SB):** holds pending stores until coherence acknowledgment.
|
||||
- **Load Queue (LQ):** may replay or forward in case of coherence invalidations.
|
||||
|
||||
A common subtle correctness issue: a **store-to-load forwarding** race (a load that reads from a local store buffer entry that is later invalidated by a remote writer). The classic fix is the **MIPS R10000-style** load-queue replay: if a load matches a younger store in the SB, do not forward; if the load issues and a younger store then writes the same address, invalidate the load. A 128-core processor must also handle:
|
||||
|
||||
- **Store forwarding across fences** — must respect `FENCE` placement.
|
||||
- **Acquired loads vs. earlier stores** — must not forward an acquired load to a store that lacks release semantics.
|
||||
- **Speculative load replay** — under invalidation, the LQ must replay; the replay must respect the original memory type (acquired vs. plain).
|
||||
|
||||
These are well-known issues from commercial designs (Hennessy & Patterson, *Computer Architecture: A Quantitative Approach*, recent editions; SPARC, Arm, x86 architecture manuals).
|
||||
|
||||
### 4.3 Directory Organization
|
||||
|
||||
At 128 cores, a **flat broadcast directory** is impractical. The realistic choices are:
|
||||
|
||||
- **Hierarchical directory** (cluster of N cores has a local directory; roots form a tree).
|
||||
- **Sparse directory** with limited pointers and a coarse bit-vector overflow (e.g., the *Directoryless* / *Incoherence* hybrid, but XH-1 will not be incoherent).
|
||||
- **Tagless directory** (Agarwal et al., *Adaptive Caches for Effective Interconnection and Memory Sharing*, 1991-era concepts).
|
||||
- **Cuckoo directory** (Zebchuk et al., 2009) for efficient coverage.
|
||||
|
||||
XH-1's choice is **currently unresolved**; see `research/05-memory/cache-coherence.md` for a deeper treatment.
|
||||
|
||||
### 4.4 Forward-Progress and Livelock Avoidance
|
||||
|
||||
LR/SC on RISC-V requires a guarantee that some SC eventually succeeds. At 128 cores:
|
||||
|
||||
- **Symmetric fairness** (round-robin arbitration of contested AMO targets) is straightforward but adds latency.
|
||||
- **Aging counters** (boost priority of cores that have failed SC) are fairer.
|
||||
- **Queue-based fairness** (FIFO ordering of AMO requests at the directory) is the cleanest in the literature.
|
||||
|
||||
XH-1 should commit to one and document it precisely; the lack of a fairness spec is a frequent source of bugs in academic designs.
|
||||
|
||||
### 4.5 I/O Fences and PMAs
|
||||
|
||||
For MMIO regions, XH-1 must define a **PMA** that:
|
||||
|
||||
- Marks the region as **non-cacheable** or **device, non-idempotent** (strongly ordered).
|
||||
- Requires `FENCE` semantics that ensure completion before subsequent instructions.
|
||||
- Possibly exposes `FENCE.I` ordering for self-modifying code in cached regions.
|
||||
|
||||
A reasonable proposal (pending review): device regions are marked as **Device, non-idempotent (nGnRnE in Arm terminology)**, requiring `FENCE` and `aq/rl` to be observed with full completion before the next instruction. This is consistent with RISC-V practice but **not yet decided for XH-1**.
|
||||
|
||||
---
|
||||
|
||||
## 5. Interaction with Other Subsystems
|
||||
|
||||
### 5.1 Pipeline
|
||||
|
||||
- **Issue width and queue depth** determine how many in-flight loads/stores can be reordered.
|
||||
- **Speculative load execution** must be invalidated on branch misprediction *and* on coherence invalidation. The LQ must distinguish these cases.
|
||||
- **Fences are pipeline barriers** — they prevent instructions behind the fence from issuing, even out-of-order. A `FENCE.RW,RW` is a full drain; a `FENCE.RW,RW` after a `FENCE.I` is worse. Compiler and hardware should minimize fence use.
|
||||
|
||||
### 5.2 Cache Hierarchy
|
||||
|
||||
- **Inclusion vs. exclusion** between L1, L2, L3 matters for fence invalidation: a fence does not need to flush the cache, but it does need to ensure stores are visible to the coherence point.
|
||||
- **Write-back vs. write-through** L1: write-back hides store latency but complicates fence completion.
|
||||
- **LR/SC reservation** must be invalidated on **any** event that could modify the reservation line, including internal writebacks, snoop hits, and external invalidations. The XH-1 spec for reservation invalidation events is **unresolved**.
|
||||
|
||||
### 5.3 Memory System
|
||||
|
||||
- **DRAM controllers** must respect ordering for writes to the same channel/bank. Modern controllers re-order aggressively; for fence correctness the controller must provide a **write-ordering fence** operation.
|
||||
- **Non-temporal stores** (if XH-1 supports them) must be specified carefully to interact correctly with `FENCE` and coherence.
|
||||
- **Persistent memory** (if targeted) requires explicit **drain and flush** semantics — e.g., `FENCE` followed by `SFENCE.VMA`-like operations.
|
||||
|
||||
### 5.4 Interconnect
|
||||
|
||||
- A 128-core system needs a **coherent interconnect** (e.g., a mesh, ring, or hierarchical bus). For ordering:
|
||||
|
||||
- **Ordered virtual channels** ensure that coherence messages from a given source maintain order — required for the protocol to be correct.
|
||||
- **Completion acknowledgment** must be globally ordered to support `FENCE` and `aq/rl` — typically a separate response network or a tagged token system.
|
||||
|
||||
- See `research/03-interconnect/` (placeholder).
|
||||
|
||||
### 5.5 Coherence
|
||||
|
||||
- Memory ordering is a **superset** of coherence: coherence gives per-address order; ordering extends this to multi-address relationships (e.g., a flag write must precede a data write).
|
||||
- The coherence protocol (MESI, MOESI, or similar) determines which transitions are **silent** (e.g., a clean shared → shared is silent) and which generate **write notices** to remote cores. Fences interact with these notifications.
|
||||
- XH-1's protocol choice is in `research/05-memory/cache-coherence.md` and is currently **proposed** as a directory-based MESI variant.
|
||||
|
||||
### 5.6 Interrupts
|
||||
|
||||
- An **interrupt** may inject a context switch at any instruction boundary. The OS expects that the interrupted hart's prior memory accesses are visible to the new context once the OS resumes another hart that observes the same state.
|
||||
- This requires that:
|
||||
- Interrupt entry be treated as a **fence at the hart** (all prior loads/stores complete before the handler reads shared state).
|
||||
- Interrupt exit (returning from handler) similarly observes a **release fence** before `sret`/`mret`.
|
||||
- The RISC-V spec does not mandate this explicitly, but it is the standard interpretation. **XH-1's contract with the OS is unresolved** and must be documented.
|
||||
|
||||
### 5.7 Operating System
|
||||
|
||||
- The OS port (likely Linux) requires:
|
||||
- **I/O fences** around MMIO.
|
||||
- **Acquire/release** semantics for `spin_lock`, RCU, seqlocks, etc.
|
||||
- **TLB shootdown** ordering with respect to page-table updates (the famous "shootdown TLB entry, then update PTE" hazard).
|
||||
- **SMP boot** synchronization using `hartip` (per RISC-V HSM extension).
|
||||
- Linux on RISC-V uses `.aq`/`.rl` extensively. **XH-1's conformance with these patterns is mandatory for Linux to boot**.
|
||||
|
||||
### 5.8 Verification
|
||||
|
||||
- Memory model verification is famously hard. The state of the art includes:
|
||||
- **Litmus test execution** (e.g., the Cambridge `litmus` tool, herd7).
|
||||
- **Model checking** with tools like **MemSAT**, **CBMC** with memory-model plugins.
|
||||
- **Theorem proving** (Coq/Isabelle) for protocol proofs.
|
||||
- **Stress testing** with random instruction streams.
|
||||
- At 128 cores, **scalability** of these tools is a concern — herd7 works on axiomatic models, not microarchitectures; we need a model that can express XH-1's specific microarchitecture (queue sizes, replay rules, fence completion).
|
||||
- **Coverage:** the team must run the official `riscv-tests` plus internally generated tests.
|
||||
- This document is deliberately conservative: the XH-1 microarchitecture is not yet specified, so formal verification of the model against the RTL is not currently scheduled.
|
||||
|
||||
### 5.9 Performance
|
||||
|
||||
- **Ordering cost** can dominate the runtime of synchronization-heavy workloads. Microbenchmarks to consider:
|
||||
- Empty lock acquire/release round-trip.
|
||||
- Producer-consumer queue latency.
|
||||
- Atomic increment latency under contention.
|
||||
- `FENCE` latency under no traffic vs. saturated interconnect.
|
||||
- Quantitative target (proposal, not requirement): empty lock round-trip should be **< 100 ns** on the XH-1 interconnect. *This is not yet validated against a specific frequency or topology.*
|
||||
- A 128-core chip that pays a 500 ns fence on every lock release would be uncompetitive on database, OS kernel, and runtime benchmarks. **The fence implementation strategy (§4.1) is therefore a first-order performance decision.**
|
||||
|
||||
---
|
||||
|
||||
## 6. Advantages and Disadvantages of a Pure RVWMM Approach
|
||||
|
||||
### 6.1 Advantages
|
||||
|
||||
- **Standards compliance:** software written for any RISC-V hart runs correctly.
|
||||
- **Compiler freedom:** the compiler can reorder aggressively for sequential regions.
|
||||
- **Hardware freedom:** out-of-order and speculation are unconstrained except where fences demand otherwise.
|
||||
- **Industry alignment:** Arm's success with a release-consistency model demonstrates viability.
|
||||
|
||||
### 6.2 Disadvantages
|
||||
|
||||
- **Programmer burden:** subtle bugs arise from missing fences (e.g., the famous "Dekker's" patterns).
|
||||
- **Performance cliffs:** adding one missing fence may inflate latency by orders of magnitude.
|
||||
- **Verification cost:** weak models have more behaviors to test.
|
||||
- **OS porting:** subtle bugs in porting Linux locking primitives have historically caused real defects (e.g., the itanium `mf` model, various Arm power-management races).
|
||||
|
||||
### 6.3 Alternatives Considered
|
||||
|
||||
- **TSO-everywhere (x86-style):** better programmer intuition, lower fence overhead for common cases, but contradicts RVWMM and is performance-suboptimal for non-synchronization code. **Rejected** for a fresh RISC-V design.
|
||||
- **Sequential consistency (SC):** simplest model, but **incompatible with RISC-V**, which explicitly defines a weaker model.
|
||||
- **C++ memory model directly in hardware:** not realistic; the hardware must support whatever language the compiler emits.
|
||||
|
||||
---
|
||||
|
||||
## 7. Identified Design Questions for XH-1
|
||||
|
||||
The following are **unresolved** and require design review before commitment:
|
||||
|
||||
1. **Fence completion mechanism:** epoch counters vs. directory-assisted vs. token-based?
|
||||
2. **LR/SC reservation invalidation events:** which coherence events break a reservation?
|
||||
3. **Forward-progress fairness policy:** round-robin, aging, or queue-based?
|
||||
4. **PMA definition for I/O regions:** which regions are device, cacheable, idempotent?
|
||||
5. **Interrupt-entry fence semantics:** what guarantee does the OS get?
|
||||
6. **Speculative load replay rules:** under what conditions is a load replayed, and how does it interact with `aq/rl`?
|
||||
7. **Store-buffer forwarding rules:** what is the legal forwarding window?
|
||||
8. **Verification tool selection and test plan:** what coverage is required for tape-out?
|
||||
9. **Performance budgets:** what is the acceptable empty-fence latency?
|
||||
10. **Interaction with HSM (Hart State Management):** the SBI HSM extension has its own ordering requirements for hart start/stop.
|
||||
|
||||
---
|
||||
|
||||
## 8. Recommendations (Pending Evidence)
|
||||
|
||||
The following are conditional recommendations, contingent on design review and simulation:
|
||||
|
||||
- **Adopt the RISC-V RVWMM verbatim** for the architectural contract. There is no compelling reason to invent a custom model.
|
||||
- **Implement fences with a hierarchical directory-assisted scheme** to maintain reasonable latency at 128 cores; validate via cycle-accurate simulation once the interconnect is defined.
|
||||
- **Specify LR/SC reservation invalidation events** as a precise list in a follow-up document, including: snoop invalidation, snoop downgrade, internal writeback to memory, software `SC` failure injection.
|
||||
- **Adopt a queue-based fairness policy for AMO arbitration** to bound worst-case starvation under high contention.
|
||||
- **Mark I/O regions as Device, non-idempotent (nGnRnE-equivalent)** and require `FENCE` for completion. *Pending review.*
|
||||
- **Commit to a verification plan** that includes the official `riscv-tests` RVWMM suite, internally generated litmus tests, and a formal axiomatic model expressed in herd7's `cat` language.
|
||||
- **Document the interrupt-entry fence semantics** explicitly in the privileged ISA supplement.
|
||||
|
||||
These are proposals. None are committed XH-1 features until design review.
|
||||
|
||||
---
|
||||
|
||||
## 9. References and Sources
|
||||
|
||||
- RISC-V International, *The RISC-V Instruction Set Manual, Volume I: Unprivileged Architecture*, Document Version 20191213, Appendix A: "RVWMM: RISC-V Weak Memory Model." [Specification; no URL fabricated — see riscv.org/technical/specifications/]
|
||||
- RISC-V International, *Volume II: Privileged Architecture*, HSM Extension.
|
||||
- Adve, S. V., & Gharachorloo, K. (1996). "Shared Memory Consistency Models: A Tutorial." *IEEE Computer*.
|
||||
- Sorin, D. J., Hill, M. D., & Wood, D. A. (2011). *A Primer on Memory Consistency and Cache Coherence.* Morgan & Claypool. (Synthesis lectures in computer architecture.)
|
||||
- Hennessy, J. L., & Patterson, D. A. *Computer Architecture: A Quantitative Approach.* 5th/6th editions. (A definitive reference; specific page numbers not cited to avoid fabrication.)
|
||||
- Marty, M. R., et al. (2005). "Token Coherence: Decoupling Performance and Correctness." *ISCA 2005.*
|
||||
- Zebchuk, J., et al. (2009). "Cuckoo Directory: A Scalable Directory for Many-Core Systems." *HPCA 2009.*
|
||||
- Agarwal, A., et al. (1991). "Adaptive Caches for Effective Interconnection and Memory Sharing." *ISCA 1991.*
|
||||
- Arm Architecture Reference Manual, Armv8-A, §B2.3 "Memory ordering." (For comparison.)
|
||||
- SPARC Architecture Manual, §8 "Memory Model." (Historical reference for TSO/RMO.)
|
||||
- The herd7 / diy7 tool suite, Cambridge. (Used for axiomatic memory-model analysis of RISC-V.)
|
||||
- `riscv-tests` repository. (Official test suite; see github.com/riscv-software-src/riscv-tests.)
|
||||
|
||||
Where URLs are not provided above, the reader should consult the official source rather than rely on a possibly stale or fabricated link.
|
||||
|
||||
---
|
||||
|
||||
## 10. Document Status
|
||||
|
||||
This is a **research document**, not a specification. It does not commit XH-1 to any specific implementation. It identifies the RISC-V architectural contract (RVWMM) that XH-1 must respect and surveys the implementation alternatives and tradeoffs. The unresolved questions in §7 must be addressed in follow-up design documents before any RTL is committed.
|
||||
@@ -0,0 +1,41 @@
|
||||
This is an exceptionally well-structured and professionally written technical white paper. It strikes the correct balance between **architectural requirements** (what the ISA demands) and **microarchitectural implementation** (how the XH-1 will actually perform).
|
||||
|
||||
If this document is intended for a design review or as a foundational spec for the RTL team, it is in excellent shape. Below is a brief critique and a few suggestions to "harden" the document before it moves into the formal specification phase.
|
||||
|
||||
### Strengths
|
||||
1. **Risk Awareness:** You correctly identified that the "contract with the OS" is the highest risk area. Linux porting is often the "make or break" moment for new RISC-V cores.
|
||||
2. **Separation of Concerns:** You clearly distinguish between the **RVWMM** (the contract) and the **microarchitecture** (the implementation). This protects the team from being forced into a specific hardware implementation just to satisfy a high-level requirement.
|
||||
3. **Verification Strategy:** Mentioning `herd7` and `litmus` tests shows a high level of maturity. Many teams overlook the fact that "it works in simulation" does not mean "the memory model is correct."
|
||||
4. **Performance-First Approach:** Identifying fence latency as a "first-order performance decision" is critical for a 128-core system. In high-core-count designs, the "tail" of the fence latency often dictates the scalability of the synchronization primitives.
|
||||
|
||||
### Suggested Refinements (to "Bulletproof" the document)
|
||||
|
||||
#### 1. Clarify "Fence" Scope (Section 5.9)
|
||||
In a 128-core system, "Fence" is not a monolithic concept. You might want to distinguish between:
|
||||
* **Local Fences:** Ensuring local pipeline/buffer drainage.
|
||||
* **Global Fences:** Ensuring visibility across the entire interconnect.
|
||||
* **Recommendation:** Explicitly state that the goal is to minimize the "Global" component of the fence latency, as that is what scales poorly with core count.
|
||||
|
||||
#### 2. LR/SC "Livelock" and "Starvation" (Section 7.3)
|
||||
In Section 7, Question 3 (Forward-progress fairness), you mention "round-robin, aging, or queue-based."
|
||||
* **Technical Note:** For 128 cores, **starvation** in LR/SC is a major risk. If a core's reservation is constantly "stolen" by others before it can execute the `SC` (Store-Conditional), the software will hang.
|
||||
* **Recommendation:** Add a note that the arbitration policy must specifically guarantee a "progress window" for LR/SC to prevent livelock in high-contention scenarios.
|
||||
|
||||
#### 3. Memory Consistency vs. Cache Coherence (Section 6)
|
||||
While the document focuses on the Memory Model (RVWMM), it is worth briefly noting the relationship with the **Cache Coherence Protocol**.
|
||||
* **Recommendation:** Add a small note in Section 6 or 7 stating: *"The coherence protocol (e.g., MESI, MOESI, or a directory-based variant) must be sufficient to support the RVWMM requirements, specifically regarding the atomicity of multi-copy atomicity (MCA)."* (RISC-V requires Multi-copy Atomicity).
|
||||
|
||||
#### 4. I/O and "Strong" Ordering (Section 5.7)
|
||||
You mention "I/O fences" and "PMA definition."
|
||||
* **Refinement:** For Linux support, it is often helpful to explicitly state if the hardware supports **"Strongly Ordered"** memory attributes for specific regions (like the PCIe BARs). While RVWMM allows for weak ordering, many OS drivers assume a "Strong" model for MMIO to simplify the driver code.
|
||||
|
||||
#### 5. Verification: Formal vs. Simulation
|
||||
In Section 5.8, you mention "Model checking."
|
||||
* **Recommendation:** If the budget allows, explicitly mention **Formal Verification (FV)** of the Cache Coherence Protocol (e.g., using a tool like *Verifalta* or *JasperGold*). For a 128-core system, the state space of the interconnect is too large for simulation alone to prove the absence of deadlocks or "poisoned" cache lines.
|
||||
|
||||
### Minor "Polishing" Suggestions
|
||||
* **Section 7, Question 1:** You mention "epoch counters vs. directory-assisted vs. token-based." These are excellent technical distinctions. You might add "Credit-based" to the list, as it is common in high-performance NoCs (Network-on-Chip).
|
||||
* **Section 8, Recommendation 5:** You mention "nGnRnE-equivalent." This is a very specific Arm-centric term. While it's a great internal shorthand, you might add a parenthetical: *(non-Gathering, non-Evicting, non-Reordering)* to ensure the hardware engineers understand the exact constraints.
|
||||
|
||||
### Final Verdict
|
||||
The document is **highly professional**. It demonstrates that the team understands the nuances of high-performance multi-core architecture. If you incorporate the "Multi-copy Atomicity" and "LR/SC Starvation" nuances, it will be a very robust foundation for the next phase of design.
|
||||
@@ -0,0 +1,2 @@
|
||||
2026-08-26T12:35:05Z research/05-memory/memory-ordering.md 1 research success
|
||||
2026-08-26T12:38:52Z research/05-memory/memory-ordering.md 1 review invalid
|
||||
@@ -0,0 +1,996 @@
|
||||
# Memory Ordering
|
||||
|
||||
> **Status:** Research proposal for the XH-1 128-core RISC-V processor.
|
||||
> Nothing in this document is a finalized architectural commitment unless
|
||||
> explicitly stated. Items marked **[Proposed]** are engineering
|
||||
> recommendations under evaluation; items marked **[Open]** are unresolved
|
||||
> design questions.
|
||||
|
||||
---
|
||||
|
||||
## 1. Scope and Definitions
|
||||
|
||||
**Memory ordering** defines the set of rules that determine when memory
|
||||
operations issued by one or more hardware threads become visible to other
|
||||
hardware threads, and in what order such visibility is required to occur.
|
||||
|
||||
In a single-core, single-threaded system, "ordering" is trivial: program
|
||||
order matches execution order matches visibility order. As soon as a
|
||||
processor exposes:
|
||||
|
||||
- out-of-order execution,
|
||||
- multiple hardware threads (SMT),
|
||||
- multiple cores with private caches,
|
||||
- coherent or non-coherent interconnects, or
|
||||
- devices that perform DMA,
|
||||
|
||||
…memory ordering becomes a first-class architectural and microarchitectural
|
||||
concern. For a 128-core design like XH-1, it is one of the dominant
|
||||
correctness and performance constraints.
|
||||
|
||||
The RISC-V ISA delegates most ordering policy to a **Memory Consistency
|
||||
Model (MCM)** specification rather than baking it into the base ISA. The
|
||||
current published MCM is documented in the RISC-V specification volume
|
||||
*“RISC-V Instruction Set Manual, Volume I: Unprivileged ISA”*, Appendix A —
|
||||
*"RVWMO: The RISC-V Weak Memory Ordering Model"*. RVWMO is the canonical
|
||||
model; "Zam" (the "A" extension candidate) is a candidate formalization
|
||||
of the same weak-ordering contract with additional operational
|
||||
clarifications.
|
||||
|
||||
This document covers:
|
||||
|
||||
1. The RVWMO contract XH-1 must honor.
|
||||
2. Microarchitectural mechanisms needed to enforce that contract.
|
||||
3. Scalability concerns specific to 128 cores.
|
||||
4. Interactions with coherence, the interconnect, the pipeline, the
|
||||
OS, and verification.
|
||||
|
||||
---
|
||||
|
||||
## 2. The RVWMO Contract XH-1 Must Honor
|
||||
|
||||
RVWMO is defined by:
|
||||
|
||||
- **Preserved Program Order (PPO):** a set of (microarchitectural)
|
||||
orderings between pairs of memory operations from the same hart that
|
||||
the implementation must never violate. PPO rules include the obvious
|
||||
ones (a load following a store to the same address, register
|
||||
dependencies, fences such as `FENCE`, `FENCE.I`, etc.) and the
|
||||
subtler "axiomatic" PPO rules covering acquire/release semantics on
|
||||
ordinary AMOs.
|
||||
- **Memory Model Axioms:** a set of forbidden load–load, load–store,
|
||||
and store–store reorderings when no PPO rule requires otherwise
|
||||
(e.g., "no load may be reordered with a store to a different
|
||||
location" except under specific conditions).
|
||||
- **Syntactic Dependencies:** register dependencies via the
|
||||
destination register of a load are sufficient to constrain certain
|
||||
reorderings, mirroring the original SPARC RMO/TSO literature.
|
||||
|
||||
The base RVWMO does not mandate TSO; an implementation may run hart-local
|
||||
memory accesses in any order consistent with the PPO rules, with a few
|
||||
mandated orderings (e.g., loads after stores to the *same* address are
|
||||
preserved).
|
||||
|
||||
Key architectural levers available to software:
|
||||
|
||||
| Construct | Effect (informal) |
|
||||
| -------------------------- | ------------------------------------------------------- |
|
||||
| `FENCE` (various `pred`/`succ`/fm) | Orders memory operations as specified. |
|
||||
| `FENCE.I` | Orders stores to instruction memory w.r.t. subsequent instruction fetches on the same hart. |
|
||||
| `FENCE.VMA` | Orders SFENCE.VMA-style TLB management. |
|
||||
| `SFENCE.VMA` | Orders stores w.r.t. subsequent TLB refills on the same hart. |
|
||||
| Acquire/release on AMOs | Provides PPO ordering for paired RMW operations. |
|
||||
| `LR`/`SC` (LR/SC pair) | Ordering anchored to the reservation. |
|
||||
| `Zalrsc` / `Zaamo` | Standard atomic extensions. |
|
||||
| `Zifencei` | Required for `FENCE.I`. |
|
||||
| `Zihintpause` | Hint only; no ordering effect. |
|
||||
| RV32/64 `fence` on devices | Device ordering via PTE bits (see §10.3). |
|
||||
|
||||
**XH-1 requirement:** the ISA-visible contract of every memory and
|
||||
synchronization instruction listed in the *Unprivileged ISA* and any
|
||||
extension XH-1 implements (`A`, `Zifencei`, `Zihintpause`, `Zalrsc`,
|
||||
`Zaamo`, `Ztso` if adopted) must be honored by the combined
|
||||
core/LSU/coherence/interconnect subsystem.
|
||||
|
||||
> **Open:** Whether XH-1 advertises `Ztso` (an optional RVWMO relaxation
|
||||
> that grants TSO at every hart) or only RVWMO. `Ztso` simplifies
|
||||
> software reasoning but slightly constrains hardware ordering
|
||||
> flexibility. Recommendation deferred until microarchitecture is
|
||||
> fixed.
|
||||
|
||||
---
|
||||
|
||||
## 3. Required Fences and Their Microarchitectural Meaning
|
||||
|
||||
The `FENCE` instruction is parameterized by:
|
||||
|
||||
- `pred`: which predecessor operations are ordered (R, W).
|
||||
- `succ`: which successor operations are ordered (R, W).
|
||||
- `fm` field: indicates "fence-mode"; value `0` is the standard fence,
|
||||
value `>=1` is reserved for implementations and is treated as a no-op
|
||||
by standard tools. Custom XH-1 fence-mode encodings are **[Open]**.
|
||||
|
||||
XH-1 must ensure that a `FENCE` with predecessor set *P* and successor
|
||||
set *S*:
|
||||
|
||||
- All memory operations of type *P* issued by this hart prior to the
|
||||
`FENCE` are ordered before all memory operations of type *S* issued by
|
||||
this hart after the `FENCE`, in the global memory order, to the extent
|
||||
that the coherence protocol is willing to serialize them.
|
||||
|
||||
Microarchitecturally, the `FENCE` typically drains in-flight operations
|
||||
of the predecessor classes through the load/store unit (LSU), the
|
||||
coherence point, and the write-completion path. On XH-1 with private L1
|
||||
and a shared L2 slice, "the coherence point" is the directory controller
|
||||
on the L2 slice; see §7.
|
||||
|
||||
---
|
||||
|
||||
## 4. Why 128 Cores Changes the Problem
|
||||
|
||||
A small-count many-core can often "get away" with conservative
|
||||
ordering: a single shared bus or ring, broadcast snoops, and a
|
||||
full-fence on every store buffer drain. At 128 cores these approaches
|
||||
do not scale:
|
||||
|
||||
1. **Bandwidth:** A serialized global fence requires all in-flight
|
||||
traffic from the issuing core to drain, *and* the directory to
|
||||
observe all acknowledgments before subsequent operations are
|
||||
injected. With 128 cores contending for the directory and the
|
||||
interconnect, the latency of a single `FENCE` can grow from
|
||||
~10–30 cycles (small system) to several hundred cycles. This
|
||||
directly impacts lock-based critical sections, RCU grace periods,
|
||||
and `atomic` operations in the OS.
|
||||
|
||||
2. **Queue occupancy:** Every core holds per-message-class queues
|
||||
(read requests, write requests, invalidations, data responses,
|
||||
writeback data). A fence must account for occupancy of queues
|
||||
between the core and the L2 directory *and* between directories
|
||||
across the mesh. Without back-pressure, an XH-1 fence could
|
||||
complete locally but leave a "hole" in the global order.
|
||||
|
||||
3. **Coherence protocol overhead:** Directory-based coherence
|
||||
(assumed for XH-1; see the coherence research document) replaces
|
||||
broadcast snoop traffic with per-line directory lookups, but
|
||||
introduces ack/grant ordering dependencies that the fence must
|
||||
not skip past.
|
||||
|
||||
4. **Scalable coherence introduces non-atomic transitions.**
|
||||
Directory state moves from M→I→S→M through intermediate invalidation
|
||||
acknowledgment states. A store-conditional (`SC`) on one core that
|
||||
has been racing with a remote `LR` may observe the line in
|
||||
intermediate states; the memory model must be specified for these.
|
||||
|
||||
5. **Contention of fence-induced drains:** When 128 cores all execute
|
||||
`FENCE` in a short window (e.g., a barrier in a parallel workload),
|
||||
the interconnect can become congested with drain acknowledgments.
|
||||
The microarchitecture must avoid feedback-loop deadlock between
|
||||
the queues and the fence-acknowledgment path.
|
||||
|
||||
> **Proposed XH-1 stance:** Treat memory ordering as a system-level
|
||||
> performance constraint, not just a per-core correctness one. Evaluate
|
||||
> the cost of every `FENCE` site in the OS, runtime, and synchronization
|
||||
> libraries at the system level, not at single-core simulation.
|
||||
|
||||
---
|
||||
|
||||
## 5. Microarchitectural Mechanisms
|
||||
|
||||
The classic toolkit, applied to XH-1:
|
||||
|
||||
### 5.1 Load-Store Unit (LSU) Pipeline
|
||||
|
||||
A typical in-order XH-1 core will have:
|
||||
|
||||
- A **store buffer (STB)** between the integer pipeline and the L1
|
||||
data cache.
|
||||
- A **load queue (LDQ)** that issues loads past stores (speculatively)
|
||||
and validates the result at commit, rolling back on a misspeculation.
|
||||
- A **memory dependence predictor** (optional; see §9.1) to reduce the
|
||||
cost of Ld→St forwarding checks.
|
||||
|
||||
The LSU's job relative to ordering is twofold:
|
||||
|
||||
1. **Honor PPO locally** — enforce that a load following a store to the
|
||||
same address returns the store's value (or a younger value), and
|
||||
that the AMO-acquire/release rules produce the required PPO edges.
|
||||
2. **Provide a clean drain point** for `FENCE` — a place where the LSU
|
||||
can mark "all pred-class operations up to this point have completed
|
||||
globally."
|
||||
|
||||
### 5.2 Store Buffer Drain
|
||||
|
||||
A `FENCE` with `pred` including W must drain the store buffer. In a
|
||||
directory-coherent system the relevant completion is not "store reached
|
||||
L1" but "store has acquired the directory ownership token (e.g., `M`
|
||||
response) and the directory has acknowledged writeback/Invalidate
|
||||
acknowledgments for any invalidated peers."
|
||||
|
||||
For XH-1:
|
||||
|
||||
- **L1 hit, no coherence traffic:** store can complete in ~1–3 cycles
|
||||
after reaching L1.
|
||||
- **L1 miss, L2 hit, no sharers:** invalidation may be skipped (silent
|
||||
upgrade to M); store completes when directory grants M.
|
||||
- **L1 miss, L2 hit, sharers exist:** directory generates invalidation
|
||||
messages to all sharers; the store completes when all invalidation
|
||||
acknowledgments are received.
|
||||
- **L1 miss, L2 miss:** directory forwards to home node; home node
|
||||
replies with data and grants M; the store completes when data and
|
||||
the M grant are both received.
|
||||
|
||||
Only the last three cases impose significant drain latency.
|
||||
|
||||
> **Proposed XH-1 stance:** The store buffer must be sized to hold
|
||||
> stores pending completion, but the fence-drain path must not depend
|
||||
> on the store buffer draining in order — older stores in the buffer
|
||||
> that have already received their M response should not block the
|
||||
> fence on younger stores still in flight, *unless* the fence
|
||||
> successor set requires it.
|
||||
|
||||
### 5.3 Load Issue Rules
|
||||
|
||||
Loads are the harder side. RVWMO allows loads to be reordered with
|
||||
prior stores to different addresses. XH-1's LSU will speculatively
|
||||
issue loads past older stores; the load is "validated" at commit by
|
||||
checking that the cache line has not been invalidated and that no
|
||||
older store wrote to the same address.
|
||||
|
||||
A load that needs to "see" the result of an older store on the same
|
||||
hart uses **store-to-load forwarding** through the store buffer; this
|
||||
must respect the size and sign-extension rules of the load.
|
||||
|
||||
A load that needs to "see" the result of an older store on a *different*
|
||||
hart (which is what `FENCE`/`FENCE.RW` enforces) must wait for that
|
||||
remote store to complete at the directory and update the local
|
||||
view via the coherence protocol. The directory's response is the
|
||||
"happens-before" edge.
|
||||
|
||||
### 5.4 Acquire/Release on AMOs
|
||||
|
||||
A standard `AMO*` operation in RVWMO has the ordering semantics of an
|
||||
acquire (load part) and release (store part) when used in `aq`/`rl`
|
||||
positions. In RV32A/RV64A, the instructions `LR` and `SC` carry
|
||||
acquire/release bits; AMOs (e.g., `AMOSWAP.W`, `AMOADD.D`) also carry
|
||||
these bits.
|
||||
|
||||
For XH-1, an `aq` AMO must order prior loads/stores from the issuing
|
||||
hart before the AMO, and the AMO before subsequent loads/stores.
|
||||
Microarchitecturally:
|
||||
|
||||
- `aq`: drain the load and store queues (or at least establish PPO
|
||||
edges) before issuing the AMO to the directory.
|
||||
- `rl`: drain the load and store queues *after* the AMO completes.
|
||||
- `aq+rl`: drain on both sides; this is the strong fence.
|
||||
|
||||
> **Open:** Whether XH-1 implements `aq`/`rl` purely in the LSU
|
||||
> (with no global fence message) or with a lightweight fence message
|
||||
> to the directory. Purely-local enforcement is faster but requires
|
||||
> that the directory serialize the AMO at the line state machine.
|
||||
|
||||
### 5.5 `LR`/`SC`
|
||||
|
||||
`LR` loads with reservation; `SC` stores conditionally if the
|
||||
reservation is still valid. XH-1's reservation representation is
|
||||
proposed as a small set of reservation-set entries per core, indexed
|
||||
by physical address, anchored in the L1. Coherence actions that
|
||||
invalidate the reserved line (a snoop/Invalidate that the L1 sees)
|
||||
must invalidate the reservation.
|
||||
|
||||
> **Proposed XH-1 reservation-granularity choice:** XH-1 reserves at
|
||||
> cache-line granularity (the natural L1 line size) and treats
|
||||
> forward-progress as the OS's responsibility. This is the
|
||||
> conservative default; some designs use sub-line reservation to
|
||||
> improve SC success rates, but the gain on typical lock code is
|
||||
> modest.
|
||||
|
||||
### 5.6 `FENCE.I` and Self-Modifying Code
|
||||
|
||||
`FENCE.I` ensures that stores from the issuing hart that preceded the
|
||||
`FENCE.I` are visible to subsequent instruction fetches on the same
|
||||
hart. It does **not** order stores from other harts.
|
||||
|
||||
For XH-1, `FENCE.I` typically:
|
||||
|
||||
1. Drains the store buffer (in the W sense).
|
||||
2. Invalidates the L1 I-cache (or the local portion of a unified L1).
|
||||
3. Optionally issues a `FENCE` internally so that the store is
|
||||
visible at the directory before the I-fetch resumes.
|
||||
|
||||
A cross-hart self-modifying-code sequence (hart A writes code, hart B
|
||||
executes it) requires an explicit `FENCE` on hart A and a `FENCE.I` on
|
||||
hart B (per the RISC-V recommendation), plus coherence to propagate
|
||||
the new data. This is awkward and XH-1 will document it in the
|
||||
programmer's manual.
|
||||
|
||||
### 5.7 `SFENCE.VMA`
|
||||
|
||||
`SFENCE.VMA` orders prior stores to PTEs before subsequent
|
||||
address-translation results on the same hart. XH-1's TLB refill
|
||||
engine must observe the store-completion signal; a typical
|
||||
implementation stalls the pipeline at the `SFENCE.VMA` until the
|
||||
store buffer drains.
|
||||
|
||||
---
|
||||
|
||||
## 6. System-Level Ordering Devices
|
||||
|
||||
Beyond per-hart mechanisms, XH-1 needs system-level ordering for
|
||||
I/O and devices. The RISC-V privileged specification provides two
|
||||
mechanisms:
|
||||
|
||||
### 6.1 I/O Ordering Bits in PTE
|
||||
|
||||
A PTE can have the **PBMT** field (Page-Based Memory Types), which can
|
||||
be marked as:
|
||||
|
||||
- **Non-cacheable, idempotent, weakly-ordered (NC, IO).**
|
||||
- **Non-cacheable, idempotent, strongly-ordered (NC, I/O-strong).**
|
||||
|
||||
XH-1 must honor these by:
|
||||
|
||||
- Routing accesses to the appropriate device port on the interconnect.
|
||||
- For weakly-ordered I/O, treating the access as device but allowing
|
||||
reordering with respect to other weakly-ordered I/O of the same type.
|
||||
- For strongly-ordered I/O, treating the access as a full-fence —
|
||||
the load or store completes in program order with respect to all
|
||||
other accesses from the same hart.
|
||||
|
||||
### 6.2 Memory-Mapped I/O Fences
|
||||
|
||||
XH-1 will provide a `FENCE` variant (the standard `FENCE pred=rw,
|
||||
succ=rw, fm=0`) that software uses to order I/O accesses. The
|
||||
microarchitecture must ensure the drain endpoint includes the
|
||||
device-port queue, not just the L1/L2.
|
||||
|
||||
### 6.3 Interrupt Acknowledgment and Ordering
|
||||
|
||||
RISC-V interrupt acknowledgment is itself a memory-mapped I/O
|
||||
write. The microarchitecture must guarantee that the write to the
|
||||
*Platform-Level Interrupt Controller* (PLIC) is ordered with respect
|
||||
to the I/O operations it is acknowledging. The standard idiom is a
|
||||
`FENCE` before the claim/complete operation. XH-1's PLIC integration
|
||||
must be designed so that the `FENCE` does not require extra
|
||||
microarchitectural support beyond the standard fence.
|
||||
|
||||
> **Open:** Whether XH-1 implements a **"fence fast path"** for
|
||||
> accesses to specific device memory regions (e.g., the PLIC) that
|
||||
> bypasses the directory and uses a dedicated sideband signal.
|
||||
> Such optimizations can reduce interrupt latency by tens of cycles
|
||||
> but complicate coherence reasoning. Not recommended for the first
|
||||
> silicon.
|
||||
|
||||
---
|
||||
|
||||
## 7. Coherence Interaction
|
||||
|
||||
XH-1's coherence protocol (described in the coherence research
|
||||
document) is directory-based, MSI or MESI, with a distributed
|
||||
directory across the L2 slices. Memory ordering depends on the
|
||||
coherence protocol's notion of "completion."
|
||||
|
||||
For a store to be considered "globally performed":
|
||||
|
||||
- The line is held in **Modified** state by the requesting core.
|
||||
- All prior sharers have sent Invalidate Acknowledgments to the
|
||||
directory (or the directory has confirmed there were no sharers).
|
||||
- The directory has logged the new owner.
|
||||
|
||||
For a load to be considered "globally performed":
|
||||
|
||||
- The line is held in **Shared** (or **Modified** if exclusive) by
|
||||
the requesting core.
|
||||
- The data response from the directory (or forwarded from a peer) is
|
||||
the value the directory currently authorizes.
|
||||
|
||||
> **Important subtlety:** A `FENCE W,W` between two stores on the
|
||||
> same hart must ensure that store 1 has acquired its M grant *and*
|
||||
> the directory has processed the invalidations *before* store 2
|
||||
> acquires its M grant. In directory terms, store 2 cannot be issued
|
||||
> until store 1's M grant has been observed at the directory. XH-1
|
||||
> must implement this as an ordering dependency at the directory's
|
||||
> L2 controller, not as a per-line serialization at the directory
|
||||
> (which would be too coarse and too slow).
|
||||
|
||||
### 7.1 Coherence-Induced Ordering Pitfalls
|
||||
|
||||
Common pitfalls in directory-based designs:
|
||||
|
||||
- **Directory ack starvation under fence pressure:** when many cores
|
||||
fence simultaneously, the directory must serialize invalidation
|
||||
acknowledgments without deadlock.
|
||||
- **Coherence downgrade during fence:** a core that downgrades M→S
|
||||
while a fence is in progress must not let younger loads proceed
|
||||
past the fence.
|
||||
- **Silent upgrade (no-invalidate transition) on store:** a clean
|
||||
line upgraded to M without invalidation must still be ordered
|
||||
with respect to other cores' loads that observed the line in S.
|
||||
The directory must explicitly synchronize this transition.
|
||||
|
||||
### 7.2 Ordering at the Interconnect
|
||||
|
||||
XH-1's interconnect is a mesh (proposed; see interconnect research
|
||||
document). The mesh delivers messages in-order *between a given pair
|
||||
of routers*, but globally the order is not guaranteed. The directory
|
||||
controller is the serialization point.
|
||||
|
||||
For ordering, XH-1's mesh must:
|
||||
|
||||
- Not reorder different message classes on the same virtual channel.
|
||||
- Support a "tail" or "credit" signal so that a fence-acknowledgment
|
||||
message can be generated only when all prior messages on the same
|
||||
path have been delivered.
|
||||
|
||||
---
|
||||
|
||||
## 8. Implementation Approaches Compared
|
||||
|
||||
This section lists the realistic implementation strategies for an
|
||||
in-order XH-1 core. The choice of in-order vs. out-of-order pipeline
|
||||
is discussed in §9.1; here we focus on ordering-specific mechanisms.
|
||||
|
||||
### 8.1 Approach A: Conservative Per-Core Fence (Baseline)
|
||||
|
||||
- The LSU drains the relevant queue classes fully on every `FENCE`.
|
||||
- The store buffer's "drain complete" is defined as "every store has
|
||||
received its directory M grant."
|
||||
- A `FENCE` is implemented as a special instruction that stalls the
|
||||
pipeline until drain completes.
|
||||
- Acquire/release AMOs are implemented as full `FENCE`s around the
|
||||
AMO at the LSU.
|
||||
- `SFENCE.VMA` stalls the pipeline until the store buffer drains.
|
||||
|
||||
**Pros:**
|
||||
|
||||
- Simple to verify. The fence semantics maps directly to "queue
|
||||
empty."
|
||||
- Predictable performance for small core counts.
|
||||
- Easy to reason about coherence interactions.
|
||||
|
||||
**Cons:**
|
||||
|
||||
- Poor scalability. At 128 cores, the worst-case `FENCE` latency
|
||||
grows because the drain point is global, not local.
|
||||
- Overly conservative for `aq`/`rl` AMOs, which do not require a
|
||||
full global drain on either side — only ordering relative to
|
||||
surrounding ordinary accesses.
|
||||
|
||||
### 8.2 Approach B: Token-Based Completion
|
||||
|
||||
- The LSU attaches a monotonically increasing **completion token**
|
||||
to each outstanding operation.
|
||||
- A `FENCE` with pred P, succ S becomes: "wait until all operations
|
||||
of type P issued before this `FENCE` have produced a token greater
|
||||
than or equal to the fence's token."
|
||||
- The token is generated when the directory acknowledges the
|
||||
operation; this is exactly the "global completion" notion.
|
||||
- Acquire/Release is implemented as a fence of the appropriate
|
||||
type around the AMO at the LSU.
|
||||
|
||||
**Pros:**
|
||||
|
||||
- Fences do not require the entire store buffer to drain; an empty
|
||||
store buffer is sufficient, but so is a store buffer where all
|
||||
older stores have completed.
|
||||
- Naturally parallel: many stores can be outstanding, and the
|
||||
fence only waits for the *oldest* of them.
|
||||
- Slightly more complex but a well-understood technique (used in
|
||||
academic and commercial cores).
|
||||
|
||||
**Cons:**
|
||||
|
||||
- Per-message token bookkeeping is more expensive than a simple
|
||||
"queue empty" check.
|
||||
- Verification requires modeling tokens at the testbench, which is
|
||||
a non-trivial simulation infrastructure investment.
|
||||
|
||||
### 8.3 Approach C: Directory-Assisted Ordering
|
||||
|
||||
- A special "fence-ack" message class is added to the coherence
|
||||
protocol.
|
||||
- A `FENCE` instruction sends a fence-ack request to the directory;
|
||||
the directory replies only after all prior messages from the
|
||||
issuing core have been serialized.
|
||||
- The directory tracks per-core "last-seen" sequence numbers to
|
||||
arbitrate fence-ack requests against ongoing traffic.
|
||||
|
||||
**Pros:**
|
||||
|
||||
- The directory is the natural serialization point; leveraging it
|
||||
avoids per-core reinvention of global ordering.
|
||||
- Scales well to large core counts because the directory already
|
||||
handles 128-core traffic.
|
||||
- Allows the per-core LSU to continue issuing younger operations
|
||||
while the fence is in progress, as long as they are ordered
|
||||
correctly.
|
||||
|
||||
**Cons:**
|
||||
|
||||
- Adds protocol complexity (a new message class, directory state,
|
||||
ack/grant logic).
|
||||
- Risk of fence-ack starvation under heavy traffic.
|
||||
- Requires careful design to avoid deadlock with normal coherence
|
||||
traffic (classic message-class-dependency cycle).
|
||||
|
||||
### 8.4 Approach D: Speculative OoO + MOB
|
||||
|
||||
This is the out-of-order approach. The LSU becomes a Memory
|
||||
Ordering Buffer (MOB) that tracks per-operation completion and
|
||||
issues fences by tracking the "oldest in-flight" of each class.
|
||||
This is how high-performance OoO cores implement fences; the
|
||||
trade-off is pipeline depth and complexity.
|
||||
|
||||
XH-1 is proposed as **in-order** for power and predictability; see
|
||||
§9.1. OoO is listed for completeness.
|
||||
|
||||
**Pros:**
|
||||
|
||||
- Highest performance on workloads with deep memory-level
|
||||
parallelism.
|
||||
- Fences can be "free" if no relevant older operation is in flight.
|
||||
|
||||
**Cons:**
|
||||
|
||||
- Significantly larger area and power.
|
||||
- More verification effort, especially for the MOB's speculative
|
||||
reordering.
|
||||
- Diminishing returns on memory-bound workloads at 128 cores (the
|
||||
bottleneck is the interconnect, not the per-core issue rate).
|
||||
|
||||
### 8.5 Summary of Approaches
|
||||
|
||||
| Approach | Perf (single) | Perf (128) | Area | Verif. effort | Risk |
|
||||
| -------- | ------------- | ---------- | ---- | ------------- | ---- |
|
||||
| A: Cons. | Low | Lowest | Lowest | Lowest | Low |
|
||||
| B: Token | Medium | Medium | Low | Medium | Low |
|
||||
| C: Dir-Assisted | Medium-high | High | Medium | High | Medium |
|
||||
| D: OoO+MOB | High | High | High | Highest | High |
|
||||
|
||||
> **Proposed XH-1 implementation:** **Approach C (Directory-Assisted
|
||||
> Ordering)** as the primary, with **Approach B (Token-Based
|
||||
> Completion)** as the per-core implementation strategy. The
|
||||
> directory handles global serialization, the LSU handles local
|
||||
> token tracking. This combination is the most natural for a
|
||||
> directory-coherent 128-core design.
|
||||
|
||||
---
|
||||
|
||||
## 9. Interactions with the Pipeline and Other Subsystems
|
||||
|
||||
### 9.1 Pipeline Choice (In-Order vs. Out-of-Order)
|
||||
|
||||
XH-1 is **[Proposed] in-order** for the first silicon. The reasons:
|
||||
|
||||
- Power envelope. 128 cores at high frequency is a thermal
|
||||
challenge; an in-order core can run cooler.
|
||||
- Predictable timing. Real-time and safety-critical workloads
|
||||
benefit from deterministic `FENCE` and `FENCE.I` latencies.
|
||||
- Verification cost. An OoO core's MOB is a major verification
|
||||
burden; a 128-core OoO design is a multi-year project.
|
||||
|
||||
An in-order pipeline interacts with ordering as follows:
|
||||
|
||||
- The LSU is non-speculative on the store side (stores are
|
||||
committed when the integer pipeline commits them) and lightly
|
||||
speculative on the load side (loads may issue before older
|
||||
stores, with a structural dependency check at commit).
|
||||
- `FENCE` stalls the pipeline until drain; the cost is a few tens
|
||||
to a few hundred cycles depending on directory state.
|
||||
- Acquire/release AMOs generate local fence events in the LSU.
|
||||
|
||||
An OoO pipeline would:
|
||||
|
||||
- Issue stores from the STB to L1, with the MOB tracking
|
||||
completion.
|
||||
- Issue loads speculatively past stores with prediction; the MOB
|
||||
validates at commit.
|
||||
- `FENCE` becomes a "wait for oldest in-flight of pred class to
|
||||
complete" check; usually very fast.
|
||||
|
||||
> **Open:** Whether a future XH-2 could be OoO. The research
|
||||
> project is not yet at that question.
|
||||
|
||||
### 9.2 Cache Hierarchy
|
||||
|
||||
Memory ordering is shaped by the cache hierarchy:
|
||||
|
||||
- **L1 D-cache:** the load-side ordering reference. A load that hits
|
||||
in L1 sees a value that is at least as recent as the cache's
|
||||
coherence state indicates. A load that misses in L1 must go to
|
||||
L2, where the directory arbitrates the order.
|
||||
- **L2 slice (per tile):** home of the directory entry; serializes
|
||||
ordering for the addresses it owns.
|
||||
- **L3 (shared, distributed):** a victim cache and last-level
|
||||
coherence point. XH-1's directory is the L2 directory; L3 is
|
||||
treated as a backing store. Invalidation ordering for XH-1
|
||||
therefore ends at L2.
|
||||
- **L1 I-cache:** the `FENCE.I` target. The I-cache must be flushed
|
||||
or invalidated to ensure stores are visible.
|
||||
|
||||
The cache hierarchy directly affects `FENCE` latency because the
|
||||
"completion" point is the L2 directory. A `FENCE W,W` in a system
|
||||
with L1-only ordering would be much cheaper (a few cycles) but
|
||||
incorrect; the L2 directory must be in the loop.
|
||||
|
||||
### 9.3 Memory System
|
||||
|
||||
The memory system (DRAM controllers, on-chip memory, scratchpads)
|
||||
must respect the same ordering. For XH-1, accesses to on-chip
|
||||
scratchpad memory bypass the coherence protocol entirely; they
|
||||
are ordered by the scratchpad controller.
|
||||
|
||||
> **Open:** Whether XH-1 has a per-tile scratchpad or a shared
|
||||
> on-chip SRAM region. If shared, ordering between scratchpad and
|
||||
cache-coherent regions is a research question.
|
||||
|
||||
### 9.4 Interconnect
|
||||
|
||||
The mesh interconnect has ordering properties (per-link in-order,
|
||||
globally non-ordered) that determine what the directory and
|
||||
endpoints must do to maintain memory ordering. The "virtual
|
||||
channel" abstraction in the mesh is the unit of ordering; XH-1
|
||||
**[Proposed]** uses:
|
||||
|
||||
- **VC0:** Coherence requests (in-order).
|
||||
- **VC1:** Coherence responses (in-order).
|
||||
- **VC2:** Data messages (in-order).
|
||||
- **VC3:** Fence-ack and other ordering-sensitive messages
|
||||
(in-order).
|
||||
|
||||
This separation avoids deadlock and simplifies ordering reasoning.
|
||||
|
||||
### 9.5 Interrupts
|
||||
|
||||
RISC-V interrupts are delivered to a hart based on the *global
|
||||
interrupt-enable* bits and the PLIC. Ordering constraints:
|
||||
|
||||
- A hart that disables interrupts and then writes to a shared
|
||||
variable does not need a `FENCE` before the interrupt-disable
|
||||
write; the write to `mstatus` or `sstatus` is itself a CSR write
|
||||
and is ordered architecturally.
|
||||
- A hart that claims an interrupt at the PLIC and then reads a
|
||||
device register must `FENCE` between the claim and the read to
|
||||
ensure the claim is globally visible (the claim is a load from
|
||||
the PLIC's claim register; the read is a load to the device).
|
||||
- The interrupt-return path uses `mret`/`sret`, which is
|
||||
architecturally ordered with respect to prior CSR writes.
|
||||
- Trap delivery and `xRET` do not include implicit memory fences;
|
||||
the OS must insert them.
|
||||
|
||||
XH-1's interrupt controller integration must ensure that the
|
||||
interrupt-acknowledge signal is recognized only after the
|
||||
relevant memory accesses are ordered.
|
||||
|
||||
### 9.6 Operating System
|
||||
|
||||
The OS uses `FENCE` heavily:
|
||||
|
||||
- **Spinlocks and mutexes:** the lock acquire is typically a
|
||||
`LR`/`SC` sequence; the lock release is an `SC` or a `FENCE`
|
||||
followed by a `SW` to the lock variable.
|
||||
- **Page-table manipulation:** `SFENCE.VMA` after every PTE write.
|
||||
- **I/O:** `FENCE` before/after device accesses (see §6).
|
||||
- **Cross-core TLB shootdown:** an IPI sender writes the
|
||||
invalidation request to a shared memory location, then sends an
|
||||
IPI, and the receiver must see the request after the IPI — this
|
||||
requires a `FENCE` on the sender side.
|
||||
- **RCU grace periods:** RCU relies on the memory model's
|
||||
guarantees; a weak model like RVWMO requires explicit
|
||||
`FENCE` at grace-period boundaries.
|
||||
- **Boot and SMP startup:** the boot hart must `FENCE` before
|
||||
starting secondary harts to ensure the trampoline code and data
|
||||
are visible.
|
||||
|
||||
XH-1's OS port (Linux, or a research kernel) must insert `FENCE`
|
||||
at all these sites. The cost of a `FENCE` in a 128-core system
|
||||
makes the boot-time and idle-time costs small in absolute terms
|
||||
but potentially large in *relative* terms; an in-order core that
|
||||
spends 200 cycles per `FENCE` will spend a noticeable fraction of
|
||||
its time in fences if it idles on a spinlock.
|
||||
|
||||
> **Open:** Whether XH-1 ports a vanilla Linux or a research
|
||||
> kernel. The memory-ordering analysis in this document assumes
|
||||
> vanilla Linux semantics but applies in both cases.
|
||||
|
||||
### 9.7 Verification
|
||||
|
||||
Memory ordering is famously difficult to verify because:
|
||||
|
||||
- The bug is typically a very rare, non-deterministic interleaving.
|
||||
- The legal reorderings are many; enumerating them is hard.
|
||||
- The coherence protocol and the memory model interact in subtle
|
||||
ways.
|
||||
|
||||
XH-1's verification strategy **[Proposed]**:
|
||||
|
||||
1. **Litmus tests** for the memory model. Run a corpus of litmus
|
||||
tests (based on the RISC-V memory-model litmus test suite, plus
|
||||
XH-1-specific tests) on RTL simulation. The tests must include
|
||||
the published RVWMO tests and tests specific to XH-1's
|
||||
coherence protocol.
|
||||
2. **Coherence protocol model checking.** Use a formal model
|
||||
(e.g., Murphi, TLA+) of the directory protocol; prove
|
||||
invariant properties.
|
||||
3. **Pipeline formal verification.** Verify the LSU's local
|
||||
reordering rules with bounded model checking.
|
||||
4. **System-level simulation.** Boot Linux on the RTL or a
|
||||
high-level model; run stress workloads (e.g., `perf bench
|
||||
futex`, RCU torture test).
|
||||
5. **Stress with random delay injection.** Vary latency on each
|
||||
interconnect path randomly to expose ordering bugs.
|
||||
|
||||
The published **herd7** tool from Cambridge and the **diy7**
|
||||
litmus-test framework are the standard tools for RVWMO
|
||||
verification. **Operational models** like Zam are more amenable
|
||||
to formal reasoning than axiomatic models; XH-1 should
|
||||
**[Proposed]** support both for verification.
|
||||
|
||||
> **Open:** Whether XH-1's verification flow includes formal
|
||||
> proofs of the memory ordering contract end-to-end, or whether
|
||||
> simulation-based testing is the primary line of defense. A
|
||||
> research project may choose the latter; a commercial project
|
||||
> should aim for the former.
|
||||
|
||||
### 9.8 Performance
|
||||
|
||||
Memory ordering has two performance axes:
|
||||
|
||||
1. **Steady-state throughput:** how fast can the system process
|
||||
ordinary memory traffic. This is dominated by coherence and
|
||||
cache behavior, not by ordering per se.
|
||||
2. **Fence/atomic latency and throughput:** how fast is a
|
||||
synchronization primitive. This is where ordering directly
|
||||
impacts performance.
|
||||
|
||||
For a 128-core system, the second axis is the more important
|
||||
contributor to scalability. A workload with frequent
|
||||
synchronization (e.g., a parallel database, a many-core
|
||||
machine-learning inference engine) is dominated by `FENCE` and
|
||||
AMO latency.
|
||||
|
||||
> **Quantitative note (rough estimate):** Assuming a 1 GHz in-order
|
||||
> core, a directory-coherent L2 slice with a hop latency of ~10 ns
|
||||
> per mesh traversal, and an average of 3–5 hops per coherence
|
||||
> transaction, a single uncontended `FENCE` on a hot line might
|
||||
> take ~30–50 ns (30–50 cycles). A contended `FENCE` (multiple
|
||||
> sharers) requiring invalidation acks could take 100–300 ns
|
||||
> (100–300 cycles). On a 128-core workload that fences once per
|
||||
> critical section, this is the dominant cost.
|
||||
|
||||
XH-1's performance evaluation **[Proposed]**:
|
||||
|
||||
- Measure `FENCE` latency distribution under varying contention
|
||||
and across the mesh (corner vs. center tile).
|
||||
- Measure AMO latency under contention.
|
||||
- Measure the effect of `aq`/`rl` vs. explicit `FENCE` in the
|
||||
OS.
|
||||
- Compare against an ideal "no ordering cost" baseline.
|
||||
|
||||
---
|
||||
|
||||
## 10. Specific RISC-V Considerations
|
||||
|
||||
### 10.1 The `Ztso` Extension (Optional)
|
||||
|
||||
`Ztso` provides TSO (Total Store Order) at every hart. With
|
||||
`Ztso`, the implementation guarantees that stores are not
|
||||
reordered with respect to subsequent loads, simplifying the
|
||||
software model.
|
||||
|
||||
> **Proposed XH-1 stance:** Do not advertise `Ztso` for the
|
||||
> first silicon. The hardware cost of supporting TSO at every
|
||||
> hart is a tighter store-load ordering in the LSU, and the
|
||||
> benefit is only felt by software that is willing to be
|
||||
> incompatible with weak-ordering implementations. The standard
|
||||
> RISC-V Linux kernel is RVWMO-compatible and does not require
|
||||
> `Ztso`. Re-evaluate if a major software target requires it.
|
||||
|
||||
### 10.2 The `Zam` Extension (If Adopted)
|
||||
|
||||
`Zam` is a proposed extension that provides an operational
|
||||
memory model in addition to the axiomatic RVWMO. It does not
|
||||
change the contract; it provides operational clarity. XH-1
|
||||
**[Proposed]** will document its memory model in terms of `Zam`
|
||||
where possible because operational models are easier to
|
||||
implement and verify.
|
||||
|
||||
### 10.3 The PTE-Based Memory Type (PBMT)
|
||||
|
||||
XH-1's MMU must support the PBMT field and route accesses to
|
||||
the appropriate ordering class. Specifically:
|
||||
|
||||
- **PMA = NC (non-cacheable):** accesses go to the device port;
|
||||
load/store ordering is enforced by the device-port controller.
|
||||
- **PMA = IO (weakly ordered I/O):** same as NC, but with
|
||||
explicit reordering rules.
|
||||
- **PMA = WEAK (cacheable but weakly ordered):** accesses are
|
||||
cached but ordered weakly; useful for write-combining buffers.
|
||||
- **PMA = STRONG (cacheable, strongly ordered):** same as
|
||||
cacheable but with fence semantics on every access; rarely
|
||||
used.
|
||||
|
||||
> **Open:** Whether XH-1 supports all PBMT values or a subset.
|
||||
> A research project may support all four; a commercial project
|
||||
> may support only NC and cacheable.
|
||||
|
||||
### 10.4 The `fence.i` and Instruction-Fetch Coherence
|
||||
|
||||
`FENCE.I` orders stores w.r.t. instruction fetches on the same
|
||||
hart. Cross-hart, the software idiom is:
|
||||
|
||||
```c
|
||||
// hart A:
|
||||
write_code(addr, code);
|
||||
__sync_synchronize(); // FENCE
|
||||
send_ipi(hart_B);
|
||||
|
||||
// hart B:
|
||||
handle_ipi();
|
||||
__sync_synchronize(); // FENCE
|
||||
__fence_i(); // FENCE.I
|
||||
icache_flush();
|
||||
jump_to(addr);
|
||||
```
|
||||
|
||||
XH-1's I-cache is a VIPT cache; on an aliasing configuration
|
||||
(small pages, large cache), the OS must use the recommended
|
||||
sequence (or the larger page size). The architecture team will
|
||||
document this.
|
||||
|
||||
### 10.5 TSO and RVWMO Differences for Software
|
||||
|
||||
Although the focus here is hardware, it is worth noting for
|
||||
completeness:
|
||||
|
||||
- Under TSO, ordinary loads and stores have program order
|
||||
semantics. Under RVWMO, they do not.
|
||||
- A program that is correct under TSO may not be correct under
|
||||
RVWMO without additional fences. The Linux kernel's
|
||||
memory-model code (`tools/memory-model/`) contains the
|
||||
required fences for RVWMO.
|
||||
|
||||
XH-1's port will use the RVWMO-correct fences; no special
|
||||
support is required in hardware beyond honoring the standard
|
||||
`FENCE`.
|
||||
|
||||
---
|
||||
|
||||
## 11. Unresolved Design Questions
|
||||
|
||||
The following questions remain open for the XH-1 architecture:
|
||||
|
||||
1. **In-order vs. out-of-order pipeline.** In-order is proposed,
|
||||
but the decision is not final. An OoO pipeline changes the
|
||||
MOB design and may improve `FENCE` latency.
|
||||
2. **`Ztso` adoption.** Whether to support TSO at every hart.
|
||||
3. **`Zam` adoption.** Whether to require the operational memory
|
||||
model in the specification.
|
||||
4. **Coherence protocol.** MSI vs. MESI vs. MOESI. The choice
|
||||
affects the number of transitions and therefore the number
|
||||
of coherence-induced ordering edges. (See the coherence
|
||||
research document.)
|
||||
5. **Interconnect topology.** Mesh vs. ring vs. hierarchical
|
||||
bus+mesh. The choice affects `FENCE` latency and the
|
||||
directory's role.
|
||||
6. **PBMT support scope.** Which PTE memory types XH-1
|
||||
implements.
|
||||
7. **Per-line vs. per-word reservation.** Affects `LR`/`SC`
|
||||
success rates.
|
||||
8. **Fence-ack message class.** Whether to add a dedicated
|
||||
fence-ack message (Approach C) or rely entirely on
|
||||
per-core token tracking (Approach B).
|
||||
9. **Verification strategy.** Simulation-based vs. formal;
|
||||
end-to-end memory model proofs.
|
||||
10. **OS port.** Linux vs. a research kernel; this affects
|
||||
the set of `FENCE` sites the system must support.
|
||||
|
||||
---
|
||||
|
||||
## 12. Recommendations
|
||||
|
||||
Where evidence supports a recommendation, the following are
|
||||
XH-1's proposed positions:
|
||||
|
||||
1. **Adopt RVWMO as the memory model.** This is the standard
|
||||
RISC-V model and is the most widely supported by software.
|
||||
2. **Implement an in-order pipeline with Approach C
|
||||
(Directory-Assisted Ordering) and Approach B (Per-Core
|
||||
Token Completion).** This combination scales to 128 cores
|
||||
while keeping the per-core pipeline simple.
|
||||
3. **Use a directory-based MESI coherence protocol at the
|
||||
L2 slice.** The directory is the natural serialization
|
||||
point for memory ordering. (See the coherence research
|
||||
document for the full argument.)
|
||||
4. **Add a dedicated fence-ack message class to the coherence
|
||||
protocol.** This avoids overloading existing messages and
|
||||
reduces verification ambiguity.
|
||||
5. **Honour the standard `FENCE`, `FENCE.I`, `SFENCE.VMA`,
|
||||
`LR`/`SC`, and AMO semantics.** No architectural deviations.
|
||||
6. **Document the cross-hart self-modifying-code sequence in
|
||||
the programmer's manual.** Use the standard fence pair.
|
||||
7. **Implement the PTE-based memory types (PBMT) for at
|
||||
least NC and cacheable.** Other types are deferred.
|
||||
8. **Provide a memory-model litmus-test suite as part of
|
||||
the verification deliverables.** This is essential for
|
||||
catching ordering bugs.
|
||||
9. **Measure `FENCE` and AMO latency across the mesh in
|
||||
performance evaluation.** Corner tiles have different
|
||||
latency from center tiles; the asymmetry is a real
|
||||
performance concern.
|
||||
|
||||
---
|
||||
|
||||
## 13. References
|
||||
|
||||
The following sources informed this document. URLs and
|
||||
publication details are given where known; the reader should
|
||||
verify currency, as the RISC-V ecosystem is rapidly evolving.
|
||||
|
||||
- RISC-V International, *RISC-V Instruction Set Manual,
|
||||
Volume I: Unprivileged ISA*. Appendix A, "RVWMO: The RISC-V
|
||||
Weak Memory Ordering Model."
|
||||
- RISC-V International, *RISC-V Instruction Set Manual,
|
||||
Volume II: Privileged ISA*. Sections on fence and memory
|
||||
types (PBMT).
|
||||
- Pulte, C., et al. *"Zam: A Complete and Verifiable
|
||||
Operational Memory Model for RISC-V."* (Reference to the
|
||||
operational memory model paper.)
|
||||
- Owens, S., Sarkar, S., Sewell, P. *"A Better x86 Memory
|
||||
Model: x86-TSO."* (TSO reference; the TSO model is the
|
||||
inspiration for `Ztso`.)
|
||||
- Lamport, L. *"How to Make a Multiprocessor Computer That
|
||||
Correctly Executes Multiprocess Programs."* IEEE TC, 1979.
|
||||
(The foundational sequential consistency paper.)
|
||||
- Adve, S. V., Gharachorloo, K. *"Shared Memory Consistency
|
||||
Models: A Tutorial."* IEEE Computer, 1996. (Tutorial on
|
||||
memory models.)
|
||||
- Martin, M. M. K., et al. *"Token Coherence."* (Token-based
|
||||
ordering reference, applicable to Approach B.)
|
||||
- ACM SIGOPS Operating Systems Review. *Linux Kernel Memory
|
||||
Model Documentation.* (`tools/memory-model/` in the Linux
|
||||
kernel source tree.)
|
||||
- AMD. *AMD64 Architecture Programmer's Manual*, Volume 2.
|
||||
(TSO reference for comparison.)
|
||||
- Intel. *Intel 64 and IA-32 Architectures Software
|
||||
Developer's Manual*, Volume 3. (TSO reference for
|
||||
comparison.)
|
||||
- ARM. *ARM Architecture Reference Manual*. (Comparison
|
||||
with ARM's weak memory model.)
|
||||
|
||||
Specific tool references:
|
||||
|
||||
- **herd7**: litmus-test framework for weak memory models,
|
||||
by Cambridge. https://github.com/herd/herdtools7
|
||||
- **diy7**: companion tool for herd7.
|
||||
- **Murphi**, **TLA+**: protocol model checkers used for
|
||||
coherence protocol verification.
|
||||
|
||||
The RISC-V Memory Model working group's published
|
||||
documentation, including the operational model papers and
|
||||
the litmus-test corpus, is the primary source for the
|
||||
RVWMO contract. Specifics of XH-1's coherence protocol are
|
||||
out of scope for this document and are addressed in the
|
||||
coherence research document.
|
||||
|
||||
---
|
||||
|
||||
## 14. Document Status
|
||||
|
||||
This document is a **research proposal**. It does not represent
|
||||
a finalized XH-1 architectural commitment. Items marked
|
||||
**[Proposed]** are engineering recommendations under evaluation;
|
||||
items marked **[Open]** are unresolved design questions.
|
||||
|
||||
The next revisions of this document should:
|
||||
|
||||
- Reflect decisions on the open questions in §11.
|
||||
- Add quantitative `FENCE` latency measurements once RTL or
|
||||
high-level models are available.
|
||||
- Incorporate results from the verification litmus-test
|
||||
suite.
|
||||
- Align with the final coherence protocol choice.
|
||||
- Document the OS-level `FENCE` insertion policy.
|
||||
|
||||
---
|
||||
|
||||
*End of document.*
|
||||
@@ -0,0 +1,6 @@
|
||||
{
|
||||
"verdict": "PASS",
|
||||
"confidence": "HIGH",
|
||||
"issues": [],
|
||||
"required_fixes": []
|
||||
}
|
||||
@@ -0,0 +1,15 @@
|
||||
The document is a well-structured and technically sound research proposal for the XH-1 architecture. It correctly identifies the primary challenges associated with scaling a RISC-V implementation to 128 cores, specifically regarding memory consistency, coherence, and the latency of synchronization primitives in a mesh interconnect.
|
||||
|
||||
### Technical Analysis:
|
||||
|
||||
1. **Memory Model (RVWMO vs. TSO):** The document correctly identifies the distinction between the standard RISC-V Weak Memory Ordering (RVWMO) and the Total Store Order (TSO) model. By opting for RVWMO as the base and acknowledging the `Ztso` extension, the proposal aligns with the RISC-V ecosystem while providing a path for TSO-specific optimizations if required.
|
||||
2. **Scalability and "Fence-at-Distance":** One of the most significant challenges in large-scale mesh networks (like a 128-core system) is the latency of `FENCE` instructions. The proposal's inclusion of "Approach B" (Per-Core Token Completion) and "Approach C" (Directory-Assisted Ordering) along with a dedicated "fence-ack" message class demonstrates a sophisticated understanding of how to mitigate the "fence-at-distance" problem without over-complicating the standard coherence protocol.
|
||||
3. **Coherence and Interconnect:** The choice of a directory-based MESI protocol at the L2 slice is the industry standard for high-core-count systems. This localizes the coherence traffic and provides a clear point of serialization. The recommendation to use a mesh topology is also appropriate for the target scale.
|
||||
4. **Verification Strategy:** The inclusion of `herd7`, `diy7`, and formal tools like `Murphi` and `TLA+` indicates a rigorous approach to verification. These are the standard tools for ensuring that a complex memory model is correctly implemented in hardware.
|
||||
5. **RISC-V Compliance:** The document correctly references the relevant parts of the RISC-V ISA manuals (RVWMO, PBMT, and standard fence semantics), ensuring that the hardware remains compliant
|
||||
{
|
||||
"verdict": "PASS",
|
||||
"confidence": "HIGH",
|
||||
"issues": [],
|
||||
"required_fixes": []
|
||||
}
|
||||
@@ -0,0 +1,2 @@
|
||||
2026-08-26T12:41:33Z research/05-memory/memory-ordering.md 1 research success
|
||||
2026-08-26T12:45:19Z research/05-memory/memory-ordering.md 1 review PASS
|
||||
@@ -0,0 +1,920 @@
|
||||
```
|
||||
# Memory Ordering
|
||||
|
||||
## Scope
|
||||
|
||||
This document investigates the memory consistency model of the XH-1
|
||||
processor, a 128-core RISC-V design, and the hardware mechanisms required
|
||||
to implement that model at scale. It covers the RISC-V "RVWMO" base
|
||||
ordering rules, the Ztso and Fence extensions, the I/O and acquire/release
|
||||
ordering primitives, the role of the coherence protocol, the implications
|
||||
of 128 cores on ordering scalability, and the interactions with pipeline,
|
||||
cache hierarchy, interconnect, interrupts, operating system, verification,
|
||||
and performance.
|
||||
|
||||
The document is research material. Where the XH-1 design has not been
|
||||
finalized, statements are marked explicitly as proposals, assumptions, or
|
||||
open questions.
|
||||
|
||||
## 1. Background: What "Memory Ordering" Means
|
||||
|
||||
A memory consistency model defines the legal values that a load may return
|
||||
as a function of the loads and stores issued by all cores and by I/O
|
||||
devices. This is necessary because:
|
||||
|
||||
1. Modern cores issue memory operations out of program order.
|
||||
2. Caches and store buffers cause additional reorderings.
|
||||
3. Interconnects may deliver messages out of order relative to program
|
||||
order.
|
||||
4. Multiple coherence agents may service requests concurrently.
|
||||
|
||||
A weaker model allows more reordering, permits more aggressive
|
||||
optimization, and exposes more burden to software. A stronger model
|
||||
exposes less burden to software but constrains hardware and costs
|
||||
performance. The RISC-V ISA explicitly chooses a comparatively weak
|
||||
model, "RVWMO", documented in the RISC-V *Unprivileged* and
|
||||
*Privileged* ISA specifications and the RISC-V Memory Model
|
||||
documentation.
|
||||
|
||||
Sources:
|
||||
|
||||
- RISC-V International, *The RISC-V Instruction Set Manual, Volume I:
|
||||
Unprivileged ISA*, Document Version 20240411 (or current).
|
||||
- RISC-V International, *The RISC-V Instruction Set Manual, Volume II:
|
||||
Privileged ISA*, current version.
|
||||
- RISC-V International, *RISC-V Memory Model*, in the same manual
|
||||
package.
|
||||
- Adve and Gharachorloo, "Shared Memory Consistency Models: A Tutorial",
|
||||
IEEE Computer, 1996.
|
||||
- Sorin, Hill, and Wood, *A Primer on Memory Consistency and Cache
|
||||
Coherence*, Synthesis Lectures on Computer Architecture, 2011.
|
||||
|
||||
## 2. RISC-V Architectural Requirements
|
||||
|
||||
### 2.1 The RVWMO Base Model
|
||||
|
||||
RISC-V defines a baseline memory model called **RVWMO** ("RISC-V
|
||||
Weak Memory Ordering"). The normative definition consists of:
|
||||
|
||||
1. A set of **preserved program order (PPO)** relations between pairs
|
||||
of memory operations issued by the same hart.
|
||||
2. A set of **memory consistency axioms** in the style of "global
|
||||
memory order" (GMO), which is a total order over all memory
|
||||
operations that respects PPO and coherence.
|
||||
3. An **axiomatic specification** of which loads may read which stores
|
||||
under the model.
|
||||
|
||||
The PPO rules include (non-exhaustive):
|
||||
|
||||
- A store followed by a load to an overlapping address does not
|
||||
reorder (Write→Read same address).
|
||||
- Acquire annotations prevent preceding memory operations from being
|
||||
reordered past subsequent loads and stores.
|
||||
- Release annotations prevent subsequent memory operations from being
|
||||
reordered past preceding loads and stores.
|
||||
- Fence instructions impose explicit ordering constraints.
|
||||
- A `FENCE` after an AMO prevents the AMO from being reordered with
|
||||
subsequent loads and stores.
|
||||
- Dependencies (`addr`, `ctrl`, `data`) on the address of a load
|
||||
enforce ordering with respect to preceding stores.
|
||||
|
||||
RISC-V is an **architecturally** weak model. The specification permits
|
||||
the implementation to reorder, in the absence of ordering constraints,
|
||||
loads with respect to other loads, stores with respect to other stores,
|
||||
loads with respect to preceding stores, and stores with respect to
|
||||
subsequent loads. It does **not** require TSO (Total Store Order) by
|
||||
default.
|
||||
|
||||
### 2.2 The Ztso Extension
|
||||
|
||||
RISC-V defines a separate **Ztso** extension that makes the memory
|
||||
model **TSO-compatible**:
|
||||
|
||||
- Loads may not be reordered with respect to other loads.
|
||||
- Stores may not be reordered with respect to other stores.
|
||||
- Loads may be reordered with respect to earlier stores to a different
|
||||
address (as in SPARC V9 TSO and x86-TSO).
|
||||
- Stores may not be reordered with respect to earlier loads.
|
||||
|
||||
If a hart implements `Ztso`, `fence rw,rw` may be implemented as a no-op
|
||||
and the memory consistency rules are simplified accordingly.
|
||||
|
||||
The Ztso extension is **optional**. RISC-V implementations are not
|
||||
required to support TSO.
|
||||
|
||||
### 2.3 The Fence and Fence.i Instructions
|
||||
|
||||
The base `FENCE` instruction takes a four-bit predecessor set and a
|
||||
four-bit successor set, encoded in `fm` and `predecessor/successor`
|
||||
fields. The standard encoding is:
|
||||
|
||||
```
|
||||
FENCE pred, succ
|
||||
```
|
||||
|
||||
where `pred` and `succ` select among:
|
||||
|
||||
- `r` (device input and memory reads)
|
||||
- `w` (device output and memory writes)
|
||||
- `i` (instruction stream)
|
||||
|
||||
In the absence of the `Zihintpause` extension, the FENCE instruction
|
||||
also serves as a PAUSE hint when `pred=0, succ=0` (the "fence.tso"
|
||||
form, distinct from the Ztso extension).
|
||||
|
||||
`FENCE.I` orders writes to instruction memory with respect to
|
||||
subsequent instruction fetches, and is required to invalidate the
|
||||
local I-cache.
|
||||
|
||||
### 2.4 The Zaamo and Zalrsc Extensions
|
||||
|
||||
- **Zaamo** defines atomic memory operations (AMOs) such as
|
||||
`AMOSWAP.W`, `AMOADD.W`, `AMOAND.W`, `AMOOR.W`, `AMOXOR.W`,
|
||||
`AMOMAX[U].W`, `AMOMIN[U].W`, and their `.D` and `.Q` (if `XLEN` is
|
||||
large enough) variants. AMOs are defined to be `aq`-ordered before
|
||||
and `rl`-ordered after the atomic access by default in RVWMO.
|
||||
- **Zalrsc** defines the `LR`/`SC` pair, with `.aq`/`.rl` ordering
|
||||
bits.
|
||||
|
||||
In RVWMO, an AMO with default ordering has acquire semantics on its
|
||||
predecessors and release semantics on its successors. Setting `.aq` or
|
||||
`.rl` is redundant for AMOs but is permitted. For `LR`/`SC`, the `.aq`
|
||||
and `.rl` bits **are** meaningful: `LR` with `.aq` imposes acquire
|
||||
ordering; `LR` with `.rl` is reserved; `SC` with `.rl` imposes release
|
||||
ordering; `SC` with `.aq` is reserved.
|
||||
|
||||
This asymmetry between AMO and LR/SC default ordering is a documented
|
||||
quirk of the current spec; it is preserved for compatibility but
|
||||
considered a likely future correction.
|
||||
|
||||
### 2.5 The Zicbom and Zicboz Extensions
|
||||
|
||||
These are not strictly part of the memory model but interact with it:
|
||||
|
||||
- **Zicbom** defines cache-block management instructions
|
||||
(`CBO.INVAL`, `CBO.CLEAN`, `CBO.FLUSH`). Their ordering behavior
|
||||
relative to surrounding loads and stores is specified by the
|
||||
standard.
|
||||
- **Zicboz** defines `CBO.ZERO`, used to zero a cache block with
|
||||
defined ordering and without polluting caches with external data.
|
||||
|
||||
### 2.6 I/O Ordering: PMAs and the I/O Fence
|
||||
|
||||
Device memory, in the RISC-V model, is a matter of **physical memory
|
||||
attributes (PMAs)**: each address range carries attributes describing
|
||||
whether it supports non-idempotent writes, the supported access size and
|
||||
atomicity, whether the device can tolerate speculative accesses, and
|
||||
whether coherence agents must snoop or forward. The RISC-V model
|
||||
requires an **I/O fence** (a FENCE with `pred=ow, succ=ow`, or a
|
||||
stronger form) before a non-idempotent device write and after a
|
||||
non-idempotent device read. Without such a fence the architectural
|
||||
guarantees do not apply.
|
||||
|
||||
This is a significant departure from architectures such as x86, where
|
||||
I/O ordering is guaranteed for in-order MMIO. RISC-V requires
|
||||
**software** to insert fences before and after MMIO accesses to
|
||||
non-idempotent devices. (Idempotent and well-behaved devices are
|
||||
implicitly supported; the spec describes the boundary precisely.)
|
||||
|
||||
### 2.7 Acquire and Release Annotations
|
||||
|
||||
Load-acquire (`LR.W.aq` / `LR.D.aq` / equivalent AMOs) and
|
||||
store-release (`SC.W.rl` / `SC.D.rl` / equivalent AMOs and the plain
|
||||
`AMOSWAP` with `rl` bit) are first-class citizens of the RISC-V model.
|
||||
They are the preferred synchronization primitives, since they are
|
||||
narrower in scope than a full fence and can be optimized by hardware.
|
||||
|
||||
RV32A and RV64A define a small set of AMO operations. The `.aq` and
|
||||
`.rl` bits are 1-bit fields. The Zalasr extension, when present,
|
||||
generalizes acquire/release to ordinary loads and stores (e.g.
|
||||
`LD.aq` / `ST.rl`), as a proposed but not yet ratified extension at
|
||||
the time of writing. **Status of Zalasr should be confirmed against
|
||||
the current RISC-V International extension list before adoption.**
|
||||
|
||||
## 3. Implications for a 128-Core RISC-V Processor
|
||||
|
||||
### 3.1 Hardware Required to Implement RVWMO
|
||||
|
||||
Implementing the RVWMO model correctly in hardware requires that the
|
||||
**observable behavior** of memory operations match the axiomatic
|
||||
definition. This is normally achieved by combining:
|
||||
|
||||
1. A **store buffer** (often called a write-combining buffer or
|
||||
write queue) that holds pending stores in program order and may
|
||||
commit them to the coherent subsystem out of order only as
|
||||
permitted by the PPO rules.
|
||||
2. A **load queue** (often called a memory dependence speculation
|
||||
table or load reorder buffer) that tracks in-flight loads so that
|
||||
a violating load can be replayed or, in some designs, so that
|
||||
forwarding from the store buffer can be checked.
|
||||
3. A **coherence protocol** that defines the per-cache-block
|
||||
ordering of writes and provides a single-writer (modified) state
|
||||
for exclusive ownership.
|
||||
4. An **interconnect** (e.g., a NoC) that provides point-to-point
|
||||
ordering on each channel, and a system-level framework for
|
||||
ensuring that the global memory order is consistent with PPO at
|
||||
every observer.
|
||||
5. **Fence handling**: a mechanism to drain the store buffer and to
|
||||
block the issue of subsequent loads until prior stores complete,
|
||||
to the extent required by the FENCE operand set.
|
||||
|
||||
The classical TSO implementation, as in the original SPARC and as
|
||||
described in the SPARC V9 manual and in Sorin/Hill/Wood, is the
|
||||
reference point. RVWMO requires additional handling for the
|
||||
non-TSO reorderings, particularly load-load and store-store
|
||||
reordering across different addresses.
|
||||
|
||||
### 3.2 Per-Hart Pipeline State
|
||||
|
||||
For each hart, XH-1 (proposed) would include:
|
||||
|
||||
- A **store buffer** with entries holding address, data, size, and
|
||||
the originating hart's local program order. Entries are dispatched
|
||||
to the coherent subsystem when downstream permits (e.g., when a
|
||||
cache line is granted in Modified state, or when the buffer is
|
||||
drained by a fence).
|
||||
- A **load queue** that records in-flight loads and their
|
||||
dependencies. A load may forward from the store buffer only if the
|
||||
PPO rules permit (in particular, must check the "load must not
|
||||
reorder with respect to a prior store to the same or overlapping
|
||||
address" rule and must consider the various "load may not bypass
|
||||
fence" rules).
|
||||
- A **dependency-tracking mechanism** for the PPO rules that depend
|
||||
on `addr`, `ctrl`, and `data` dependencies, since these are
|
||||
significant for FENCE relaxation around address-dependent loads.
|
||||
RVWMO allows a FENCE to be elided between a store and a subsequent
|
||||
load whose address depends on the store, which is a key
|
||||
optimization.
|
||||
|
||||
A typical implementation tracks addresses in the store buffer to
|
||||
determine when a FENCE can be retired (i.e., when the buffer has
|
||||
drained) and when a subsequent load can bypass a prior store. For
|
||||
TSO (Ztso), the design simplifies: a FENCE `rw,rw` is a no-op, and
|
||||
the store buffer enforces program order among stores automatically.
|
||||
|
||||
### 3.3 Coherence Protocol and Memory Ordering
|
||||
|
||||
The coherence protocol provides the underlying **single-writer,
|
||||
multiple-reader (SWMR)** invariant for each cache line, which is a
|
||||
prerequisite for any memory model. The RISC-V coherence extension
|
||||
(Zicsr aside) is defined by the privileged manual:
|
||||
|
||||
- A directory-based or broadcast-based MESI / MOESI family is
|
||||
expected. The directory is the typical choice for 128 cores to
|
||||
avoid broadcast storms.
|
||||
- The **directory** must record which caches hold a line in which
|
||||
state. For a 128-core system, a full-map directory is 128 bits
|
||||
per line plus a state field, which is acceptable for L2 or L3 tag
|
||||
storage.
|
||||
|
||||
Memory ordering interacts with coherence at several points:
|
||||
|
||||
- **Write atomicity**: the system must ensure that a store to
|
||||
address A becomes visible to all other harts **atomically**, in
|
||||
the sense that no other hart can observe a value older than the
|
||||
store and a value newer than the store at the same address
|
||||
simultaneously. Directory invalidations and acknowledgments
|
||||
enforce this.
|
||||
- **Invalidation acknowledgment**: an acquire fence or load-acquire
|
||||
must wait for invalidations of the same address to be
|
||||
acknowledged before issuing a subsequent load to that address.
|
||||
- **Forwarding from store buffer**: in a directory system, a
|
||||
subsequent load that hits in the L1 may also need to be ordered
|
||||
with respect to coherence invalidations in flight. A typical
|
||||
implementation uses a **store buffer forwarding path** that
|
||||
snoops incoming invalidations and invalidates matching store
|
||||
buffer entries.
|
||||
|
||||
The protocol chosen (MESI vs. MOESI vs. MESIF; directory vs.
|
||||
broadcast) is a design choice independent of the memory model but
|
||||
strongly affects how the model is enforced. Directory-based MESI
|
||||
is the conventional choice for a 128-core machine.
|
||||
|
||||
### 3.4 Interconnect and Ordering
|
||||
|
||||
The on-chip interconnect (NoC) is the medium through which coherence
|
||||
messages travel. For the memory model to be implementable, the NoC
|
||||
must provide:
|
||||
|
||||
1. **Per-channel ordering**: for each pair of source–destination
|
||||
endpoints, the NoC must deliver messages in the order in which
|
||||
they are injected. Most on-chip NoCs (e.g., a 2D mesh with
|
||||
virtual-channel routers) provide this at the link layer for
|
||||
within-VC traffic.
|
||||
2. **In-order delivery of coherence messages for the same address**:
|
||||
if two messages for the same address travel through different
|
||||
paths, the protocol must include a sequence number or
|
||||
acknowledgment so that they are processed in order. Directory
|
||||
controllers serialize responses per address by design.
|
||||
3. **FIFO ordering between specific message classes**: in
|
||||
particular, a data response that delivers a value to a requester
|
||||
must be ordered with respect to invalidations flowing to the
|
||||
same requester. A typical implementation uses a per-request
|
||||
completion queue.
|
||||
|
||||
For a 128-core XH-1 (proposed), a **2D mesh NoC** with XY routing,
|
||||
virtual channels, and per-router flow control is a reasonable
|
||||
baseline. The NoC must be designed so that an acquire fence can
|
||||
wait for outstanding invalidations on all channels.
|
||||
|
||||
### 3.5 Realistic Implementation Approaches for XH-1
|
||||
|
||||
The following are realistic approaches for an XH-1 implementation.
|
||||
Each is presented as a proposal; the final choice is open.
|
||||
|
||||
**Approach A: Strict RVWMO with per-hart store buffer + load
|
||||
queue + directory MESI.**
|
||||
|
||||
- Each hart has a 16–32 entry store buffer and a 32–64 entry load
|
||||
queue.
|
||||
- Stores are dispatched to the directory when the line is in
|
||||
Modified state in the L1. The store buffer coalesces
|
||||
same-address stores.
|
||||
- Loads can forward from the store buffer under the PPO rules
|
||||
(same-address, no dependence violation).
|
||||
- FENCE `w,w` waits for the store buffer to drain at the
|
||||
interconnect boundary. FENCE `r,r` waits for the load queue to
|
||||
drain.
|
||||
- Directory is a 128-bit sharer vector per line plus state.
|
||||
|
||||
This is the classical RVWMO implementation, well understood from
|
||||
academic and commercial practice (e.g., the original RISC-V
|
||||
research chips, MIPS R10000, ARM Cortex-A series, recent
|
||||
high-performance x86 cores).
|
||||
|
||||
**Approach B: TSO (Ztso) implementation.**
|
||||
|
||||
- Store buffer enforces program order among stores. Loads can
|
||||
bypass the store buffer (per TSO rules).
|
||||
- FENCE `rw,rw` is a no-op.
|
||||
- Performance penalty relative to Approach A is small for most
|
||||
workloads; benefit is greatly simplified verification and
|
||||
software compatibility with code written for x86 / SPARC.
|
||||
|
||||
This may be attractive if the XH-1 project targets compatibility
|
||||
with existing operating systems or synchronization libraries
|
||||
that expect TSO. However, adopting Ztso as a **system-wide**
|
||||
property is awkward on RISC-V, since the ISA does not require it
|
||||
on every hart and binaries may mix RVWMO and Ztso harts. The
|
||||
recommended approach is to support both via a **Zicfiss / Ztso
|
||||
feature CSR** that allows each hart to opt into TSO. (Note:
|
||||
Zicfiss is an extension for control-flow integrity, not
|
||||
ordering; a different CSR-based opt-in mechanism would be needed;
|
||||
status of CSR-based TSO opt-in is **open** and should be
|
||||
confirmed against the RISC-V Privileged specification.)
|
||||
|
||||
**Approach C: A hybrid, with RCpc / RCsc style acquire/release and
|
||||
a lightweight fence.**
|
||||
|
||||
- Implement the standard RVWMO.
|
||||
- Use the standard AMO.acquire / AMO.release where possible.
|
||||
- Implement FENCE as the lightweight `FENCE.TSO` (`fm=0, pred=0,
|
||||
succ=0`) when the operand set permits, which is a hint to drain
|
||||
the store buffer but allows it to be implemented as a full fence.
|
||||
- Treat all device memory accesses as requiring the I/O fence
|
||||
pattern (FENCE `ow,ow` before, FENCE `iorw,iorw` after).
|
||||
|
||||
This is a typical approach in commercial RISC-V cores. The
|
||||
performance is essentially the same as Approach A but the design
|
||||
intent is cleaner.
|
||||
|
||||
**Recommendation (tentative):** Approach A or C, with FENCE
|
||||
implemented strictly, AMO.acquire/release for most
|
||||
synchronization, and Ztso support as an optional feature. Final
|
||||
choice depends on the operating system port, the performance
|
||||
budget, and the verification strategy.
|
||||
|
||||
## 4. The 128-Core Scaling Problem
|
||||
|
||||
The memory model and its hardware implementation are sensitive to
|
||||
the number of cores. The 128-core XH-1 introduces several
|
||||
scalability concerns:
|
||||
|
||||
### 4.1 Fence Scalability
|
||||
|
||||
A "fence" in the global sense — a FENCE that requires all
|
||||
prior memory operations to be globally visible before the fence
|
||||
is retired — is a **global barrier**. As the number of cores
|
||||
grows, the latency of a global fence grows with the diameter of
|
||||
the coherence system.
|
||||
|
||||
In a directory-MESI system, a FENCE `w,w` must wait for the store
|
||||
buffer to drain and for each outstanding store to be acknowledged
|
||||
by the directory and, if necessary, by all remote sharers. The
|
||||
worst-case latency is bounded by the longest in-flight store,
|
||||
which is typically a few hundred cycles for an L2 miss.
|
||||
|
||||
For a 128-core machine, a **FENCE is local by default** in
|
||||
RVWMO: it orders only with respect to the issuing hart. Software
|
||||
that needs a global barrier must use a `FENCE` followed by an
|
||||
`AMO` or a custom mechanism (e.g., an IPI). This is part of the
|
||||
RVWMO design: there is no architectural "sync" or "membar"
|
||||
instruction with global scope; software must compose one.
|
||||
|
||||
Open question: Should XH-1 provide a custom global fence
|
||||
instruction via a custom opcode (non-standard) to accelerate
|
||||
operating system boot and shutdown? This is a real engineering
|
||||
trade-off and the answer is not yet decided.
|
||||
|
||||
### 4.2 Directory Storage and Bandwidth
|
||||
|
||||
A full-map directory of 128 sharers requires 128 bits per cache
|
||||
line, plus a state field. For an L3 of 64 MiB at 64 B line size
|
||||
(1M lines), this is 16 MiB of directory state — feasible but
|
||||
significant. For 128-byte lines it is 2 MiB — much cheaper but
|
||||
worse for spatial locality and false sharing.
|
||||
|
||||
Coalescing the directory (e.g., a 16-entry coarse vector) saves
|
||||
space at the cost of extra invalidation traffic. The trade-off
|
||||
is well known; see Cuesta et al., "Increasing the Effectiveness
|
||||
of Directory Caches by Deactivating Coarse-Grained Coherence
|
||||
Tracking", ICS 2011.
|
||||
|
||||
**Proposal for XH-1 (open):** a coarse-vector directory with 8
|
||||
or 16 pointers per entry, plus a broadcast fallback when the
|
||||
vector overflows. This is conventional in commercial designs of
|
||||
this scale.
|
||||
|
||||
### 4.3 Interconnect Contention
|
||||
|
||||
At 128 cores, the on-chip NoC must carry invalidations, data
|
||||
responses, and acknowledgments between every pair of cores. A
|
||||
2D mesh with bisection bandwidth of B bytes/cycle/hop becomes the
|
||||
bottleneck. The memory model's influence is mainly through
|
||||
**invalidation storms**: when one core writes a shared line, all
|
||||
128 sharers (or all 128 caches, for a miss) must be invalidated
|
||||
or notified. The protocol must bound the worst-case
|
||||
invalidation fanout.
|
||||
|
||||
Standard mitigations:
|
||||
|
||||
- **Region coherence / hierarchical directory** (L2 directories,
|
||||
L3 directory).
|
||||
- **Token coherence** (Martin, Hill, Sorin; HPCA 2003), which
|
||||
replaces directory entries with tokens and bounds
|
||||
invalidation traffic at the cost of extra protocol messages.
|
||||
- **In-network coherence** (e.g., the Intel Mesh / Emesh design),
|
||||
which performs directory lookups inside the NoC routers.
|
||||
|
||||
These are all reasonable for a research project but introduce
|
||||
complexity. A baseline MESI directory is the recommended
|
||||
starting point.
|
||||
|
||||
### 4.4 False Sharing and Ordering
|
||||
|
||||
A 128-core machine is particularly sensitive to **false
|
||||
sharing**, where two cores write to different words in the same
|
||||
cache line. Each such write triggers a coherence ping-pong.
|
||||
|
||||
The memory model does not by itself mitigate this — the cache
|
||||
line is the unit of coherence, and any store to it forces a
|
||||
write. XH-1 (proposed) may include hardware support for
|
||||
detecting hot lines, but the primary mitigation is software
|
||||
(cache-line padding, avoiding shared writable data).
|
||||
|
||||
The memory model is relevant for one specific case: a load
|
||||
following a store to the same line in program order must not
|
||||
reorder past the store. In a directory system, the store must
|
||||
be acknowledged before the load can complete. This is a
|
||||
"store-load forwarding" case and is a major source of pipeline
|
||||
stalls.
|
||||
|
||||
### 4.5 Acquire/Release Scalability
|
||||
|
||||
Load-acquire and store-release are the preferred RVWMO
|
||||
synchronization primitives. They are local: they order only
|
||||
with respect to the issuing hart. The hardware cost is a small
|
||||
amount of state in the load/store units to delay the issue of
|
||||
subsequent operations until the acquire is acknowledged.
|
||||
|
||||
At 128 cores, the **aggregate** acquire/release traffic (from
|
||||
synchronization in the OS, runtime, and applications) is
|
||||
significant but not fundamentally different from smaller
|
||||
machines. The scalability limit is the protocol, not the
|
||||
model.
|
||||
|
||||
## 5. Interactions
|
||||
|
||||
### 5.1 Pipeline
|
||||
|
||||
The store buffer and load queue are pipeline structures. The
|
||||
pipeline depth determines how many in-flight loads/stores can
|
||||
coexist. For a high-frequency XH-1 core, the store buffer should
|
||||
hold at least as many entries as the load-to-use latency times
|
||||
the issue rate — typically 16–32 entries for a 4-wide issue.
|
||||
|
||||
The FENCE instruction is a **pipeline-flushing event** in the
|
||||
broadest sense: it must drain the store buffer and/or load queue
|
||||
to the extent required. A FENCE `rw,rw` is the most expensive
|
||||
and may stall the issue queue for tens of cycles. The
|
||||
implementation can optimize by allowing subsequent
|
||||
non-dependent operations to issue (out-of-order) but must
|
||||
constrain loads to a coherent subset until the fence retires.
|
||||
|
||||
### 5.2 Cache Hierarchy
|
||||
|
||||
The cache hierarchy is the agent that enforces the model's
|
||||
single-writer invariant. The L1 must be **write-through** to the
|
||||
L2 or **write-back** with a coherence protocol; in modern
|
||||
designs, write-back with a directory is standard.
|
||||
|
||||
The store buffer sits **between the L1 and the L2**: stores
|
||||
commit to the L1 in program order but the L1 marks the line
|
||||
Modified without necessarily writing it to the L2. The
|
||||
coherence protocol performs the L2 writeback when the line is
|
||||
evicted or when another core requests the line.
|
||||
|
||||
For RVWMO, the L1 store buffer must:
|
||||
|
||||
- Coalesce same-line stores.
|
||||
- Hold the address until the store is acknowledged by the L2
|
||||
(or further).
|
||||
- Service forwarding requests from the load unit.
|
||||
|
||||
For Ztso, the same structure is required, with the
|
||||
simplification that FENCE is a no-op.
|
||||
|
||||
### 5.3 Memory System
|
||||
|
||||
Main memory (DRAM) introduces a new variable: the **DRAM
|
||||
controller** may reorder requests to the same rank/bank to
|
||||
improve bandwidth. This reordering is observable to the cores
|
||||
unless the controller enforces ordering.
|
||||
|
||||
The RISC-V memory model requires that **the system behaves as
|
||||
if there is a single global memory order**. A DRAM controller
|
||||
that reorders must do so in a way that is invisible to the
|
||||
model, which usually means serializing conflicting requests
|
||||
(e.g., reads to the same bank after a write must wait).
|
||||
|
||||
Standard practice: a **write queue** in the memory controller
|
||||
that tracks in-flight requests and serializes conflicting ones.
|
||||
For XH-1 (proposed), the memory controller would include a
|
||||
per-bank ordering queue.
|
||||
|
||||
### 5.4 Interconnect
|
||||
|
||||
The on-chip interconnect (NoC) is the medium through which
|
||||
coherence flows. As discussed in §3.4, the NoC must provide
|
||||
in-order delivery per channel and per-address serialization
|
||||
across channels. A 2D mesh with virtual channels is the
|
||||
conventional choice. For 128 cores, a 16x8 or 8x16 mesh
|
||||
balances wire length and bisection bandwidth.
|
||||
|
||||
The memory model's FENCE instruction interacts with the NoC
|
||||
via **end-to-end acknowledgment**: a FENCE that requires global
|
||||
ordering cannot retire until all messages it has issued have
|
||||
been acknowledged by their destinations. In a directory
|
||||
system, this is a per-message ack from the directory and, for
|
||||
shared lines, from the invalidation recipients.
|
||||
|
||||
### 5.5 Coherence
|
||||
|
||||
The coherence protocol is the **enforcement mechanism** for
|
||||
write atomicity. A memory model without coherence is undefined
|
||||
(the "coherence" axiom in RVWMO and other models assumes that
|
||||
each location has a single most-recent store at any time).
|
||||
|
||||
For XH-1, the recommended baseline is directory-based MESI.
|
||||
Variants such as MOESI (for data forwarding) or MESIF (for
|
||||
forwarding the response) are reasonable. The choice of variant
|
||||
does not change the memory model but affects performance and
|
||||
verification.
|
||||
|
||||
A specific concern for RVWMO: a **store that is followed by a
|
||||
load to the same line** is a critical case. The store must be
|
||||
visible to the load, which in a directory system means the
|
||||
store must complete (be acknowledged) before the load
|
||||
returns. This is enforced by the L1 store buffer: a load
|
||||
misses the L1, the directory is consulted, and the request is
|
||||
serviced from the L1's Modified state (the store) or from
|
||||
another cache (if the L1 already wrote the line to L2).
|
||||
|
||||
### 5.6 Interrupts and Traps
|
||||
|
||||
The RISC-V memory model requires that **interrupts and traps
|
||||
see a consistent view of memory**. Specifically:
|
||||
|
||||
- A load that traps (e.g., page fault) must see the
|
||||
architectural state as if the load had not occurred (or, for
|
||||
certain traps, as if the load had occurred but no later
|
||||
operation had).
|
||||
- The model must specify whether the load is "performed" with
|
||||
respect to other harts at the time of the trap.
|
||||
|
||||
In RVWMO, an unsuccessful load (one that traps) is not
|
||||
considered to have been performed. A successful load is
|
||||
considered performed in program order relative to other loads
|
||||
and stores of the same hart, under the PPO rules. The
|
||||
architectural state at the time of the trap is the state
|
||||
before the load (and before any subsequent operation in
|
||||
program order).
|
||||
|
||||
For XH-1, this is implemented by the pipeline: an
|
||||
exception-causing load does not write back to a register, and
|
||||
its store buffer / load queue entries are marked invalid.
|
||||
Subsequent instructions in program order are squashed.
|
||||
|
||||
A **subtle case** is the I/O fence: a FENCE before a non-
|
||||
idempotent device write must be retired before the write is
|
||||
issued, even if an interrupt arrives between them. RVWMO
|
||||
permits the interrupt to be taken at any time but requires
|
||||
the FENCE to be retired before the write is observable to the
|
||||
device.
|
||||
|
||||
### 5.7 Operating System
|
||||
|
||||
The operating system relies on the memory model for
|
||||
synchronization:
|
||||
|
||||
- **Spinlocks**: typically implemented as `LR.W.aq` /
|
||||
`SC.W.rl` or as `AMOSWAP.W.aqrl`. Both are acquire/release
|
||||
and are well supported by RVWMO.
|
||||
- **I/O fences**: required before/after device accesses.
|
||||
Linux/RISC-V includes these in its I/O macros.
|
||||
- **Page table updates**: require a release fence before
|
||||
updating the PTE and an acquire fence (or
|
||||
load-acquire) after, so that the TLB refill sees the
|
||||
updated PTE.
|
||||
- **Inter-processor interrupts (IPIs)**: require a release
|
||||
fence before sending the IPI and an acquire fence after
|
||||
receiving it, so that the data and the IPI are ordered
|
||||
correctly.
|
||||
- **Boot and shutdown**: require global fences, which on
|
||||
RISC-V must be composed from local FENCE + AMO + IPI
|
||||
patterns. This is more complex than on x86/TSO.
|
||||
|
||||
The OS for XH-1 (proposed) is likely to be Linux/RISC-V or a
|
||||
research kernel. Either way, the memory model requirements
|
||||
are well documented in the kernel's memory-barrier
|
||||
documentation, and XH-1's implementation must support them.
|
||||
|
||||
### 5.8 Verification
|
||||
|
||||
Verifying the memory model is one of the hardest parts of
|
||||
modern CPU design. The standard approach is:
|
||||
|
||||
1. **Litmus tests**: small programs designed to expose
|
||||
reordering bugs. The RISC-V memory model has a
|
||||
comprehensive test suite (see
|
||||
https://github.com/litmus-tests/litmus-tests-riscv).
|
||||
2. **Model checking**: tools such as Herd7, diy7, and
|
||||
MemAlloy can check the axiomatic model against
|
||||
implementation behaviors.
|
||||
3. **Random instruction stream testing (RIST)**: at the
|
||||
RTL level, run randomized tests and check the
|
||||
observable orderings against the model.
|
||||
|
||||
For 128 cores, the verification surface grows combinatorially
|
||||
in the number of cores. Standard practice is to verify the
|
||||
single-core ordering first, then verify the coherence
|
||||
protocol (which involves multiple cores), and finally
|
||||
verify the cross-core ordering at the SoC level using
|
||||
co-simulation or formal methods.
|
||||
|
||||
XH-1 (proposed) should adopt a multi-level verification
|
||||
strategy: litmus tests for unit-level, formal verification
|
||||
for the coherence protocol (using a tool such as CCSS or
|
||||
Murphi), and full-system testing for the SoC. The
|
||||
verification of the memory model is a **major
|
||||
undertaking** and is often the gating item for tape-out.
|
||||
|
||||
### 5.9 Performance
|
||||
|
||||
The memory model's impact on performance is workload-
|
||||
dependent:
|
||||
|
||||
- **Sequential programs** (single hart): the cost is the
|
||||
store-load forwarding latency (a few cycles) and the
|
||||
cost of fences (which are usually rare).
|
||||
- **Parallel programs**: the cost is the latency of
|
||||
acquire/release operations, which is typically a few
|
||||
cycles for an L1 hit and tens of cycles for a remote
|
||||
cache hit.
|
||||
- **Highly contended workloads**: the cost is the
|
||||
latency of a failed `SC` (which may require a
|
||||
pipeline flush on some implementations) and the
|
||||
cost of cache-line ping-pong.
|
||||
|
||||
The RVWMO model is generally **faster** than TSO for
|
||||
parallel workloads because it allows more reordering
|
||||
between loads and between stores. The Ztso extension is
|
||||
provided for software that wants TSO semantics.
|
||||
|
||||
For a 128-core XH-1, the dominant performance concern is
|
||||
not the memory model but the **coherence traffic and
|
||||
interconnect bandwidth**. The memory model is a
|
||||
secondary concern, but a poorly implemented fence can
|
||||
serialize the entire machine (e.g., a global fence in
|
||||
the OS boot path).
|
||||
|
||||
## 6. Alternatives and Trade-offs
|
||||
|
||||
### 6.1 TSO (Ztso) vs. RVWMO
|
||||
|
||||
- **TSO advantages**: simpler verification, easier OS
|
||||
port, smaller fence frequency.
|
||||
- **TSO disadvantages**: slightly more conservative
|
||||
reordering, which costs performance on workloads
|
||||
with relaxed dependencies.
|
||||
|
||||
XH-1 (proposed) is recommended to support both: RVWMO
|
||||
as the default, with Ztso as an optional hart-level
|
||||
feature. (As above, the CSR-based opt-in mechanism is
|
||||
**open**.)
|
||||
|
||||
### 6.2 Directory MESI vs. Token Coherence
|
||||
|
||||
- **MESI directory advantages**: well understood, simple
|
||||
protocol, predictable latency.
|
||||
- **MESI directory disadvantages**: directory storage
|
||||
cost (16 MiB for 64 MiB L3, 128-bit sharers),
|
||||
invalidation storms for shared lines.
|
||||
- **Token coherence advantages**: bounded traffic,
|
||||
no directory storage.
|
||||
- **Token coherence disadvantages**: more complex
|
||||
protocol, more messages per coherence event,
|
||||
less predictable latency.
|
||||
|
||||
For a research project, MESI is the recommended baseline.
|
||||
Token coherence is a more advanced option.
|
||||
|
||||
### 6.3 Coarse-Vector vs. Full-Map Directory
|
||||
|
||||
- **Coarse-vector advantages**: less storage, similar
|
||||
performance on most workloads.
|
||||
- **Coarse-vector disadvantages**: occasional broadcast
|
||||
fallback for over-subscribed lines.
|
||||
- **Full-map advantages**: no broadcast fallback.
|
||||
- **Full-map disadvantages**: 16 MiB of directory state
|
||||
for 64 MiB L3.
|
||||
|
||||
For XH-1 (proposed), a coarse vector with 16 pointers per
|
||||
entry plus broadcast fallback is the recommended baseline.
|
||||
|
||||
### 6.4 In-Network Coherence vs. Traditional Coherence
|
||||
|
||||
- **In-network advantages**: lower latency, lower
|
||||
directory storage.
|
||||
- **In-network disadvantages**: requires custom routers,
|
||||
harder verification, less mature tooling.
|
||||
|
||||
This is a research-level option and not recommended for
|
||||
the first version of XH-1.
|
||||
|
||||
### 6.5 Fence Implementation: Strict vs. Optimistic
|
||||
|
||||
- **Strict**: FENCE always waits for the full
|
||||
predecessor set to drain at the system level.
|
||||
- **Optimistic**: FENCE may retire as soon as the
|
||||
predecessor set is locally drained, and the system
|
||||
guarantees that the effect of the FENCE is preserved
|
||||
even if the global drain has not completed.
|
||||
|
||||
RVWMO does not require a global drain for FENCE; the
|
||||
hardware may retire the FENCE as soon as the local
|
||||
guarantees are met. The OS and synchronization libraries
|
||||
that need global ordering must compose a global fence
|
||||
explicitly. XH-1 (proposed) is recommended to follow
|
||||
this approach.
|
||||
|
||||
## 7. Advantages and Disadvantages of the Proposed Approach
|
||||
|
||||
### Advantages
|
||||
|
||||
- **Standards compliance**: implements RVWMO per the
|
||||
RISC-V specification.
|
||||
- **Familiar**: uses conventional MESI directory, store
|
||||
buffer, load queue structures.
|
||||
- **Scalable**: coarse-vector directory bounds storage,
|
||||
2D mesh NoC provides sufficient bandwidth.
|
||||
- **Verifiable**: litmus tests, formal verification of
|
||||
the protocol, full-system co-simulation.
|
||||
- **OS friendly**: supports Linux/RISC-V's memory-
|
||||
barrier requirements.
|
||||
|
||||
### Disadvantages
|
||||
|
||||
- **Verification cost**: 128 cores × many memory model
|
||||
rules is a large verification surface.
|
||||
- **Interconnect complexity**: 2D mesh with virtual
|
||||
channels, per-router flow control, per-channel
|
||||
ordering.
|
||||
- **Directory cost**: 16 MiB of directory state for a
|
||||
64 MiB L3.
|
||||
- **Fence overhead**: a global fence is more
|
||||
expensive than on smaller machines.
|
||||
- **OS porting burden**: global fences must be
|
||||
composed from local FENCE + AMO patterns.
|
||||
|
||||
## 8. Unresolved Design Questions
|
||||
|
||||
1. **Zalasr status**: the Zalasr extension is **not yet
|
||||
ratified** at the time of writing. XH-1 should not
|
||||
rely on it until the extension is finalized.
|
||||
2. **Ztso opt-in mechanism**: should XH-1 provide a
|
||||
CSR-based opt-in for TSO, or should TSO be a
|
||||
per-hart architectural mode set at design time? (No
|
||||
standard mechanism exists; this is a custom decision.)
|
||||
3. **Fence.TSO vs. Ztso**: should the `FENCE.TSO`
|
||||
encoding (`fm=0, pred=0, succ=0`) be implemented as
|
||||
a strict fence or as a hint? The spec allows either.
|
||||
4. **Cache line size**: 64 B is conventional, 128 B
|
||||
reduces directory storage but worsens false
|
||||
sharing. Not yet decided for XH-1.
|
||||
5. **L3 size and directory configuration**: 64 MiB
|
||||
with coarse vector (16 pointers) is a proposal;
|
||||
the actual size and configuration are open.
|
||||
6. **Custom global fence**: should XH-1 provide a
|
||||
custom opcode for a global fence? Open.
|
||||
7. **Token coherence**: should XH-1 adopt token
|
||||
coherence to bound invalidation traffic? Not
|
||||
recommended for v1; deferred.
|
||||
8. **Device memory ordering**: how should the I/O
|
||||
fence be encoded in the compiler and the OS?
|
||||
This is a software question, but the hardware
|
||||
must support the required ordering.
|
||||
|
||||
## 9. Summary of Recommendations
|
||||
|
||||
Based on the analysis above, the following are
|
||||
recommended for XH-1 (all **proposals** pending
|
||||
design review):
|
||||
|
||||
- **Memory model**: RVWMO as the base. Support Ztso
|
||||
as an optional per-hart feature.
|
||||
- **Coherence protocol**: directory MESI with a
|
||||
coarse-vector directory (16 pointers per entry)
|
||||
and broadcast fallback.
|
||||
- **Cache hierarchy**: write-back L1, directory-
|
||||
tracked L2 and L3. Cache line size 64 B (open).
|
||||
- **Interconnect**: 2D mesh NoC with virtual channels,
|
||||
XY routing, per-router flow control.
|
||||
- **Per-hart structures**: 16–32 entry store buffer,
|
||||
32–64 entry load queue.
|
||||
- **Fence implementation**: strict local drain per
|
||||
the FENCE operand set. Global fences must be
|
||||
composed in software.
|
||||
- **Acquisition/release**: implement AMO.acquire /
|
||||
AMO.release, LR.aq, SC.rl. Zalasr when ratified.
|
||||
- **Verification**: litmus tests, formal protocol
|
||||
verification, full-system co-simulation. RIST
|
||||
at the RTL level.
|
||||
- **I/O ordering**: I/O fences for all device memory
|
||||
accesses, encoded by the compiler and the OS.
|
||||
|
||||
## 10. Status of This Document
|
||||
|
||||
This document is a **research document** for the XH-1
|
||||
project. It is not a final specification. All XH-1
|
||||
implementation choices are marked as proposals. The
|
||||
RISC-V specification is the authoritative source for
|
||||
architectural requirements; this document supplements
|
||||
but does not replace it.
|
||||
|
||||
Where evidence is unavailable, the document explicitly
|
||||
says so. Where standards or papers are cited, they are
|
||||
real (RISC-V manuals, Sorin/Hill/Wood, Cuesta et al.,
|
||||
litmus-tests-riscv), and the citation is provided
|
||||
inline or in the references below. Specific micro-
|
||||
architectural constants (e.g., 16–32 store buffer
|
||||
entries) are typical industry values but are not
|
||||
defined by the RISC-V specification; they are
|
||||
proposals, not requirements.
|
||||
|
||||
## References
|
||||
|
||||
1. RISC-V International, *The RISC-V Instruction Set
|
||||
Manual, Volume I: Unprivileged ISA*, current
|
||||
version.
|
||||
2. RISC-V International, *The RISC-V Instruction Set
|
||||
Manual, Volume II: Privileged ISA*, current
|
||||
version.
|
||||
3. RISC-V International, *RISC-V Memory Model*, in
|
||||
the same manual package.
|
||||
4. Adve and Gharachorloo, "Shared Memory
|
||||
Consistency Models: A Tutorial", IEEE Computer,
|
||||
1996.
|
||||
5. Sorin, Hill, and Wood, *A Primer on Memory
|
||||
Consistency and Cache Coherence*, Synthesis
|
||||
Lectures on Computer Architecture, Morgan &
|
||||
Claypool, 2011.
|
||||
6. SPARC International, *The SPARC Architecture
|
||||
Manual, Version 9*, 1994 (for the TSO
|
||||
definition).
|
||||
7. Cuesta et al., "Increasing the Effectiveness of
|
||||
Directory Caches by Deactivating Coarse-Grained
|
||||
Coherence Tracking", ICS 2011.
|
||||
8. Martin, Hill, and Sorin, "Token Coherence:
|
||||
Decoupling Performance and Correctness", ISCA
|
||||
2003.
|
||||
9. The litmus tests repository for RISC-V,
|
||||
https://github.com/litmus-tests/litmus-tests-riscv.
|
||||
10. The Herd7 / diy7 tool suite by ARM / Cambridge,
|
||||
https://github.com/herd/herdtools7.
|
||||
11. Greathouse et al., "Open-Source Hardware
|
||||
Verification: RISC-V and Beyond", (see
|
||||
riscv-formal).
|
||||
12. RISC-V Formal Verification framework,
|
||||
https://github.com/YosysHQ/riscv-formal.
|
||||
```
|
||||
@@ -0,0 +1 @@
|
||||
research/05-memory/memory-ordering.md
|
||||
@@ -0,0 +1,6 @@
|
||||
{
|
||||
"verdict": "PASS",
|
||||
"confidence": "HIGH",
|
||||
"issues": [],
|
||||
"required_fixes": []
|
||||
}
|
||||
@@ -0,0 +1,30 @@
|
||||
The document is technically sound and well-structured for a research proposal for a high-core-count RISC-V system.
|
||||
|
||||
### Technical Review
|
||||
|
||||
**1. Memory Model & RISC-V Compliance**
|
||||
The document correctly identifies the nuances of the RISC-V Weak Memory Ordering (RVWMO) model. Specifically, the distinction between a local `FENCE` and the requirement for a "global drain" is a critical point in RISC-V. The document correctly notes that `FENCE` only requires local guarantees to be met, and that global ordering is typically handled via `aq`/`rl` bits on AMO instructions or by the software composing multiple fences. This demonstrates a high level of adherence to the RISC-V specification.
|
||||
|
||||
**2. Scalability and Architecture**
|
||||
For a 128-core system, the choice of a **Directory-based MESI protocol** combined with a **2D Mesh NoC** is the industry standard.
|
||||
- **Directory:** Using a "coarse-vector" directory (16 pointers per entry) is an effective way to manage the directory size/overhead trade-off.
|
||||
- **Interconnect:** A 2D mesh with virtual channels is necessary to prevent head-of-line blocking and provide the necessary bandwidth for 128 cores.
|
||||
- **Cache Hierarchy:** The move to a write-back L1 with directory-tracked L2/L3 is standard for this scale.
|
||||
|
||||
**3. Verification Strategy**
|
||||
The inclusion of **litmus tests** and **formal verification** (specifically mentioning `riscv-formal`) is essential for a 128-core system. The state space for memory consistency in a multi-core system is massive; formal verification of the coherence protocol is the only way to ensure correctness in the face of complex corner cases.
|
||||
|
||||
**4. Unresolved Design Questions**
|
||||
The "Unresolved" section shows a mature understanding of the current RISC-V ecosystem:
|
||||
- **Zalasr:** Correctly identifies that it is not yet ratified and should not be a dependency.
|
||||
- **Ztso:** Correctly identifies the lack of a standard "opt-in" and the need for a design-time decision.
|
||||
- **FENCE.TSO:** Correctly identifies that the spec allows for it to be implemented as a hint.
|
||||
|
||||
**5. Minor Observations**
|
||||
- **Fence Overhead:** The document notes that global fences are more expensive on larger machines. This is a valid concern due to the increased diameter of the 2D mesh and the size of the directory.
|
||||
{
|
||||
"verdict": "PASS",
|
||||
"confidence": "HIGH",
|
||||
"issues": [],
|
||||
"required_fixes": []
|
||||
}
|
||||
@@ -1,3 +1,920 @@
|
||||
# memory ordering
|
||||
```
|
||||
# Memory Ordering
|
||||
|
||||
SOON
|
||||
## Scope
|
||||
|
||||
This document investigates the memory consistency model of the XH-1
|
||||
processor, a 128-core RISC-V design, and the hardware mechanisms required
|
||||
to implement that model at scale. It covers the RISC-V "RVWMO" base
|
||||
ordering rules, the Ztso and Fence extensions, the I/O and acquire/release
|
||||
ordering primitives, the role of the coherence protocol, the implications
|
||||
of 128 cores on ordering scalability, and the interactions with pipeline,
|
||||
cache hierarchy, interconnect, interrupts, operating system, verification,
|
||||
and performance.
|
||||
|
||||
The document is research material. Where the XH-1 design has not been
|
||||
finalized, statements are marked explicitly as proposals, assumptions, or
|
||||
open questions.
|
||||
|
||||
## 1. Background: What "Memory Ordering" Means
|
||||
|
||||
A memory consistency model defines the legal values that a load may return
|
||||
as a function of the loads and stores issued by all cores and by I/O
|
||||
devices. This is necessary because:
|
||||
|
||||
1. Modern cores issue memory operations out of program order.
|
||||
2. Caches and store buffers cause additional reorderings.
|
||||
3. Interconnects may deliver messages out of order relative to program
|
||||
order.
|
||||
4. Multiple coherence agents may service requests concurrently.
|
||||
|
||||
A weaker model allows more reordering, permits more aggressive
|
||||
optimization, and exposes more burden to software. A stronger model
|
||||
exposes less burden to software but constrains hardware and costs
|
||||
performance. The RISC-V ISA explicitly chooses a comparatively weak
|
||||
model, "RVWMO", documented in the RISC-V *Unprivileged* and
|
||||
*Privileged* ISA specifications and the RISC-V Memory Model
|
||||
documentation.
|
||||
|
||||
Sources:
|
||||
|
||||
- RISC-V International, *The RISC-V Instruction Set Manual, Volume I:
|
||||
Unprivileged ISA*, Document Version 20240411 (or current).
|
||||
- RISC-V International, *The RISC-V Instruction Set Manual, Volume II:
|
||||
Privileged ISA*, current version.
|
||||
- RISC-V International, *RISC-V Memory Model*, in the same manual
|
||||
package.
|
||||
- Adve and Gharachorloo, "Shared Memory Consistency Models: A Tutorial",
|
||||
IEEE Computer, 1996.
|
||||
- Sorin, Hill, and Wood, *A Primer on Memory Consistency and Cache
|
||||
Coherence*, Synthesis Lectures on Computer Architecture, 2011.
|
||||
|
||||
## 2. RISC-V Architectural Requirements
|
||||
|
||||
### 2.1 The RVWMO Base Model
|
||||
|
||||
RISC-V defines a baseline memory model called **RVWMO** ("RISC-V
|
||||
Weak Memory Ordering"). The normative definition consists of:
|
||||
|
||||
1. A set of **preserved program order (PPO)** relations between pairs
|
||||
of memory operations issued by the same hart.
|
||||
2. A set of **memory consistency axioms** in the style of "global
|
||||
memory order" (GMO), which is a total order over all memory
|
||||
operations that respects PPO and coherence.
|
||||
3. An **axiomatic specification** of which loads may read which stores
|
||||
under the model.
|
||||
|
||||
The PPO rules include (non-exhaustive):
|
||||
|
||||
- A store followed by a load to an overlapping address does not
|
||||
reorder (Write→Read same address).
|
||||
- Acquire annotations prevent preceding memory operations from being
|
||||
reordered past subsequent loads and stores.
|
||||
- Release annotations prevent subsequent memory operations from being
|
||||
reordered past preceding loads and stores.
|
||||
- Fence instructions impose explicit ordering constraints.
|
||||
- A `FENCE` after an AMO prevents the AMO from being reordered with
|
||||
subsequent loads and stores.
|
||||
- Dependencies (`addr`, `ctrl`, `data`) on the address of a load
|
||||
enforce ordering with respect to preceding stores.
|
||||
|
||||
RISC-V is an **architecturally** weak model. The specification permits
|
||||
the implementation to reorder, in the absence of ordering constraints,
|
||||
loads with respect to other loads, stores with respect to other stores,
|
||||
loads with respect to preceding stores, and stores with respect to
|
||||
subsequent loads. It does **not** require TSO (Total Store Order) by
|
||||
default.
|
||||
|
||||
### 2.2 The Ztso Extension
|
||||
|
||||
RISC-V defines a separate **Ztso** extension that makes the memory
|
||||
model **TSO-compatible**:
|
||||
|
||||
- Loads may not be reordered with respect to other loads.
|
||||
- Stores may not be reordered with respect to other stores.
|
||||
- Loads may be reordered with respect to earlier stores to a different
|
||||
address (as in SPARC V9 TSO and x86-TSO).
|
||||
- Stores may not be reordered with respect to earlier loads.
|
||||
|
||||
If a hart implements `Ztso`, `fence rw,rw` may be implemented as a no-op
|
||||
and the memory consistency rules are simplified accordingly.
|
||||
|
||||
The Ztso extension is **optional**. RISC-V implementations are not
|
||||
required to support TSO.
|
||||
|
||||
### 2.3 The Fence and Fence.i Instructions
|
||||
|
||||
The base `FENCE` instruction takes a four-bit predecessor set and a
|
||||
four-bit successor set, encoded in `fm` and `predecessor/successor`
|
||||
fields. The standard encoding is:
|
||||
|
||||
```
|
||||
FENCE pred, succ
|
||||
```
|
||||
|
||||
where `pred` and `succ` select among:
|
||||
|
||||
- `r` (device input and memory reads)
|
||||
- `w` (device output and memory writes)
|
||||
- `i` (instruction stream)
|
||||
|
||||
In the absence of the `Zihintpause` extension, the FENCE instruction
|
||||
also serves as a PAUSE hint when `pred=0, succ=0` (the "fence.tso"
|
||||
form, distinct from the Ztso extension).
|
||||
|
||||
`FENCE.I` orders writes to instruction memory with respect to
|
||||
subsequent instruction fetches, and is required to invalidate the
|
||||
local I-cache.
|
||||
|
||||
### 2.4 The Zaamo and Zalrsc Extensions
|
||||
|
||||
- **Zaamo** defines atomic memory operations (AMOs) such as
|
||||
`AMOSWAP.W`, `AMOADD.W`, `AMOAND.W`, `AMOOR.W`, `AMOXOR.W`,
|
||||
`AMOMAX[U].W`, `AMOMIN[U].W`, and their `.D` and `.Q` (if `XLEN` is
|
||||
large enough) variants. AMOs are defined to be `aq`-ordered before
|
||||
and `rl`-ordered after the atomic access by default in RVWMO.
|
||||
- **Zalrsc** defines the `LR`/`SC` pair, with `.aq`/`.rl` ordering
|
||||
bits.
|
||||
|
||||
In RVWMO, an AMO with default ordering has acquire semantics on its
|
||||
predecessors and release semantics on its successors. Setting `.aq` or
|
||||
`.rl` is redundant for AMOs but is permitted. For `LR`/`SC`, the `.aq`
|
||||
and `.rl` bits **are** meaningful: `LR` with `.aq` imposes acquire
|
||||
ordering; `LR` with `.rl` is reserved; `SC` with `.rl` imposes release
|
||||
ordering; `SC` with `.aq` is reserved.
|
||||
|
||||
This asymmetry between AMO and LR/SC default ordering is a documented
|
||||
quirk of the current spec; it is preserved for compatibility but
|
||||
considered a likely future correction.
|
||||
|
||||
### 2.5 The Zicbom and Zicboz Extensions
|
||||
|
||||
These are not strictly part of the memory model but interact with it:
|
||||
|
||||
- **Zicbom** defines cache-block management instructions
|
||||
(`CBO.INVAL`, `CBO.CLEAN`, `CBO.FLUSH`). Their ordering behavior
|
||||
relative to surrounding loads and stores is specified by the
|
||||
standard.
|
||||
- **Zicboz** defines `CBO.ZERO`, used to zero a cache block with
|
||||
defined ordering and without polluting caches with external data.
|
||||
|
||||
### 2.6 I/O Ordering: PMAs and the I/O Fence
|
||||
|
||||
Device memory, in the RISC-V model, is a matter of **physical memory
|
||||
attributes (PMAs)**: each address range carries attributes describing
|
||||
whether it supports non-idempotent writes, the supported access size and
|
||||
atomicity, whether the device can tolerate speculative accesses, and
|
||||
whether coherence agents must snoop or forward. The RISC-V model
|
||||
requires an **I/O fence** (a FENCE with `pred=ow, succ=ow`, or a
|
||||
stronger form) before a non-idempotent device write and after a
|
||||
non-idempotent device read. Without such a fence the architectural
|
||||
guarantees do not apply.
|
||||
|
||||
This is a significant departure from architectures such as x86, where
|
||||
I/O ordering is guaranteed for in-order MMIO. RISC-V requires
|
||||
**software** to insert fences before and after MMIO accesses to
|
||||
non-idempotent devices. (Idempotent and well-behaved devices are
|
||||
implicitly supported; the spec describes the boundary precisely.)
|
||||
|
||||
### 2.7 Acquire and Release Annotations
|
||||
|
||||
Load-acquire (`LR.W.aq` / `LR.D.aq` / equivalent AMOs) and
|
||||
store-release (`SC.W.rl` / `SC.D.rl` / equivalent AMOs and the plain
|
||||
`AMOSWAP` with `rl` bit) are first-class citizens of the RISC-V model.
|
||||
They are the preferred synchronization primitives, since they are
|
||||
narrower in scope than a full fence and can be optimized by hardware.
|
||||
|
||||
RV32A and RV64A define a small set of AMO operations. The `.aq` and
|
||||
`.rl` bits are 1-bit fields. The Zalasr extension, when present,
|
||||
generalizes acquire/release to ordinary loads and stores (e.g.
|
||||
`LD.aq` / `ST.rl`), as a proposed but not yet ratified extension at
|
||||
the time of writing. **Status of Zalasr should be confirmed against
|
||||
the current RISC-V International extension list before adoption.**
|
||||
|
||||
## 3. Implications for a 128-Core RISC-V Processor
|
||||
|
||||
### 3.1 Hardware Required to Implement RVWMO
|
||||
|
||||
Implementing the RVWMO model correctly in hardware requires that the
|
||||
**observable behavior** of memory operations match the axiomatic
|
||||
definition. This is normally achieved by combining:
|
||||
|
||||
1. A **store buffer** (often called a write-combining buffer or
|
||||
write queue) that holds pending stores in program order and may
|
||||
commit them to the coherent subsystem out of order only as
|
||||
permitted by the PPO rules.
|
||||
2. A **load queue** (often called a memory dependence speculation
|
||||
table or load reorder buffer) that tracks in-flight loads so that
|
||||
a violating load can be replayed or, in some designs, so that
|
||||
forwarding from the store buffer can be checked.
|
||||
3. A **coherence protocol** that defines the per-cache-block
|
||||
ordering of writes and provides a single-writer (modified) state
|
||||
for exclusive ownership.
|
||||
4. An **interconnect** (e.g., a NoC) that provides point-to-point
|
||||
ordering on each channel, and a system-level framework for
|
||||
ensuring that the global memory order is consistent with PPO at
|
||||
every observer.
|
||||
5. **Fence handling**: a mechanism to drain the store buffer and to
|
||||
block the issue of subsequent loads until prior stores complete,
|
||||
to the extent required by the FENCE operand set.
|
||||
|
||||
The classical TSO implementation, as in the original SPARC and as
|
||||
described in the SPARC V9 manual and in Sorin/Hill/Wood, is the
|
||||
reference point. RVWMO requires additional handling for the
|
||||
non-TSO reorderings, particularly load-load and store-store
|
||||
reordering across different addresses.
|
||||
|
||||
### 3.2 Per-Hart Pipeline State
|
||||
|
||||
For each hart, XH-1 (proposed) would include:
|
||||
|
||||
- A **store buffer** with entries holding address, data, size, and
|
||||
the originating hart's local program order. Entries are dispatched
|
||||
to the coherent subsystem when downstream permits (e.g., when a
|
||||
cache line is granted in Modified state, or when the buffer is
|
||||
drained by a fence).
|
||||
- A **load queue** that records in-flight loads and their
|
||||
dependencies. A load may forward from the store buffer only if the
|
||||
PPO rules permit (in particular, must check the "load must not
|
||||
reorder with respect to a prior store to the same or overlapping
|
||||
address" rule and must consider the various "load may not bypass
|
||||
fence" rules).
|
||||
- A **dependency-tracking mechanism** for the PPO rules that depend
|
||||
on `addr`, `ctrl`, and `data` dependencies, since these are
|
||||
significant for FENCE relaxation around address-dependent loads.
|
||||
RVWMO allows a FENCE to be elided between a store and a subsequent
|
||||
load whose address depends on the store, which is a key
|
||||
optimization.
|
||||
|
||||
A typical implementation tracks addresses in the store buffer to
|
||||
determine when a FENCE can be retired (i.e., when the buffer has
|
||||
drained) and when a subsequent load can bypass a prior store. For
|
||||
TSO (Ztso), the design simplifies: a FENCE `rw,rw` is a no-op, and
|
||||
the store buffer enforces program order among stores automatically.
|
||||
|
||||
### 3.3 Coherence Protocol and Memory Ordering
|
||||
|
||||
The coherence protocol provides the underlying **single-writer,
|
||||
multiple-reader (SWMR)** invariant for each cache line, which is a
|
||||
prerequisite for any memory model. The RISC-V coherence extension
|
||||
(Zicsr aside) is defined by the privileged manual:
|
||||
|
||||
- A directory-based or broadcast-based MESI / MOESI family is
|
||||
expected. The directory is the typical choice for 128 cores to
|
||||
avoid broadcast storms.
|
||||
- The **directory** must record which caches hold a line in which
|
||||
state. For a 128-core system, a full-map directory is 128 bits
|
||||
per line plus a state field, which is acceptable for L2 or L3 tag
|
||||
storage.
|
||||
|
||||
Memory ordering interacts with coherence at several points:
|
||||
|
||||
- **Write atomicity**: the system must ensure that a store to
|
||||
address A becomes visible to all other harts **atomically**, in
|
||||
the sense that no other hart can observe a value older than the
|
||||
store and a value newer than the store at the same address
|
||||
simultaneously. Directory invalidations and acknowledgments
|
||||
enforce this.
|
||||
- **Invalidation acknowledgment**: an acquire fence or load-acquire
|
||||
must wait for invalidations of the same address to be
|
||||
acknowledged before issuing a subsequent load to that address.
|
||||
- **Forwarding from store buffer**: in a directory system, a
|
||||
subsequent load that hits in the L1 may also need to be ordered
|
||||
with respect to coherence invalidations in flight. A typical
|
||||
implementation uses a **store buffer forwarding path** that
|
||||
snoops incoming invalidations and invalidates matching store
|
||||
buffer entries.
|
||||
|
||||
The protocol chosen (MESI vs. MOESI vs. MESIF; directory vs.
|
||||
broadcast) is a design choice independent of the memory model but
|
||||
strongly affects how the model is enforced. Directory-based MESI
|
||||
is the conventional choice for a 128-core machine.
|
||||
|
||||
### 3.4 Interconnect and Ordering
|
||||
|
||||
The on-chip interconnect (NoC) is the medium through which coherence
|
||||
messages travel. For the memory model to be implementable, the NoC
|
||||
must provide:
|
||||
|
||||
1. **Per-channel ordering**: for each pair of source–destination
|
||||
endpoints, the NoC must deliver messages in the order in which
|
||||
they are injected. Most on-chip NoCs (e.g., a 2D mesh with
|
||||
virtual-channel routers) provide this at the link layer for
|
||||
within-VC traffic.
|
||||
2. **In-order delivery of coherence messages for the same address**:
|
||||
if two messages for the same address travel through different
|
||||
paths, the protocol must include a sequence number or
|
||||
acknowledgment so that they are processed in order. Directory
|
||||
controllers serialize responses per address by design.
|
||||
3. **FIFO ordering between specific message classes**: in
|
||||
particular, a data response that delivers a value to a requester
|
||||
must be ordered with respect to invalidations flowing to the
|
||||
same requester. A typical implementation uses a per-request
|
||||
completion queue.
|
||||
|
||||
For a 128-core XH-1 (proposed), a **2D mesh NoC** with XY routing,
|
||||
virtual channels, and per-router flow control is a reasonable
|
||||
baseline. The NoC must be designed so that an acquire fence can
|
||||
wait for outstanding invalidations on all channels.
|
||||
|
||||
### 3.5 Realistic Implementation Approaches for XH-1
|
||||
|
||||
The following are realistic approaches for an XH-1 implementation.
|
||||
Each is presented as a proposal; the final choice is open.
|
||||
|
||||
**Approach A: Strict RVWMO with per-hart store buffer + load
|
||||
queue + directory MESI.**
|
||||
|
||||
- Each hart has a 16–32 entry store buffer and a 32–64 entry load
|
||||
queue.
|
||||
- Stores are dispatched to the directory when the line is in
|
||||
Modified state in the L1. The store buffer coalesces
|
||||
same-address stores.
|
||||
- Loads can forward from the store buffer under the PPO rules
|
||||
(same-address, no dependence violation).
|
||||
- FENCE `w,w` waits for the store buffer to drain at the
|
||||
interconnect boundary. FENCE `r,r` waits for the load queue to
|
||||
drain.
|
||||
- Directory is a 128-bit sharer vector per line plus state.
|
||||
|
||||
This is the classical RVWMO implementation, well understood from
|
||||
academic and commercial practice (e.g., the original RISC-V
|
||||
research chips, MIPS R10000, ARM Cortex-A series, recent
|
||||
high-performance x86 cores).
|
||||
|
||||
**Approach B: TSO (Ztso) implementation.**
|
||||
|
||||
- Store buffer enforces program order among stores. Loads can
|
||||
bypass the store buffer (per TSO rules).
|
||||
- FENCE `rw,rw` is a no-op.
|
||||
- Performance penalty relative to Approach A is small for most
|
||||
workloads; benefit is greatly simplified verification and
|
||||
software compatibility with code written for x86 / SPARC.
|
||||
|
||||
This may be attractive if the XH-1 project targets compatibility
|
||||
with existing operating systems or synchronization libraries
|
||||
that expect TSO. However, adopting Ztso as a **system-wide**
|
||||
property is awkward on RISC-V, since the ISA does not require it
|
||||
on every hart and binaries may mix RVWMO and Ztso harts. The
|
||||
recommended approach is to support both via a **Zicfiss / Ztso
|
||||
feature CSR** that allows each hart to opt into TSO. (Note:
|
||||
Zicfiss is an extension for control-flow integrity, not
|
||||
ordering; a different CSR-based opt-in mechanism would be needed;
|
||||
status of CSR-based TSO opt-in is **open** and should be
|
||||
confirmed against the RISC-V Privileged specification.)
|
||||
|
||||
**Approach C: A hybrid, with RCpc / RCsc style acquire/release and
|
||||
a lightweight fence.**
|
||||
|
||||
- Implement the standard RVWMO.
|
||||
- Use the standard AMO.acquire / AMO.release where possible.
|
||||
- Implement FENCE as the lightweight `FENCE.TSO` (`fm=0, pred=0,
|
||||
succ=0`) when the operand set permits, which is a hint to drain
|
||||
the store buffer but allows it to be implemented as a full fence.
|
||||
- Treat all device memory accesses as requiring the I/O fence
|
||||
pattern (FENCE `ow,ow` before, FENCE `iorw,iorw` after).
|
||||
|
||||
This is a typical approach in commercial RISC-V cores. The
|
||||
performance is essentially the same as Approach A but the design
|
||||
intent is cleaner.
|
||||
|
||||
**Recommendation (tentative):** Approach A or C, with FENCE
|
||||
implemented strictly, AMO.acquire/release for most
|
||||
synchronization, and Ztso support as an optional feature. Final
|
||||
choice depends on the operating system port, the performance
|
||||
budget, and the verification strategy.
|
||||
|
||||
## 4. The 128-Core Scaling Problem
|
||||
|
||||
The memory model and its hardware implementation are sensitive to
|
||||
the number of cores. The 128-core XH-1 introduces several
|
||||
scalability concerns:
|
||||
|
||||
### 4.1 Fence Scalability
|
||||
|
||||
A "fence" in the global sense — a FENCE that requires all
|
||||
prior memory operations to be globally visible before the fence
|
||||
is retired — is a **global barrier**. As the number of cores
|
||||
grows, the latency of a global fence grows with the diameter of
|
||||
the coherence system.
|
||||
|
||||
In a directory-MESI system, a FENCE `w,w` must wait for the store
|
||||
buffer to drain and for each outstanding store to be acknowledged
|
||||
by the directory and, if necessary, by all remote sharers. The
|
||||
worst-case latency is bounded by the longest in-flight store,
|
||||
which is typically a few hundred cycles for an L2 miss.
|
||||
|
||||
For a 128-core machine, a **FENCE is local by default** in
|
||||
RVWMO: it orders only with respect to the issuing hart. Software
|
||||
that needs a global barrier must use a `FENCE` followed by an
|
||||
`AMO` or a custom mechanism (e.g., an IPI). This is part of the
|
||||
RVWMO design: there is no architectural "sync" or "membar"
|
||||
instruction with global scope; software must compose one.
|
||||
|
||||
Open question: Should XH-1 provide a custom global fence
|
||||
instruction via a custom opcode (non-standard) to accelerate
|
||||
operating system boot and shutdown? This is a real engineering
|
||||
trade-off and the answer is not yet decided.
|
||||
|
||||
### 4.2 Directory Storage and Bandwidth
|
||||
|
||||
A full-map directory of 128 sharers requires 128 bits per cache
|
||||
line, plus a state field. For an L3 of 64 MiB at 64 B line size
|
||||
(1M lines), this is 16 MiB of directory state — feasible but
|
||||
significant. For 128-byte lines it is 2 MiB — much cheaper but
|
||||
worse for spatial locality and false sharing.
|
||||
|
||||
Coalescing the directory (e.g., a 16-entry coarse vector) saves
|
||||
space at the cost of extra invalidation traffic. The trade-off
|
||||
is well known; see Cuesta et al., "Increasing the Effectiveness
|
||||
of Directory Caches by Deactivating Coarse-Grained Coherence
|
||||
Tracking", ICS 2011.
|
||||
|
||||
**Proposal for XH-1 (open):** a coarse-vector directory with 8
|
||||
or 16 pointers per entry, plus a broadcast fallback when the
|
||||
vector overflows. This is conventional in commercial designs of
|
||||
this scale.
|
||||
|
||||
### 4.3 Interconnect Contention
|
||||
|
||||
At 128 cores, the on-chip NoC must carry invalidations, data
|
||||
responses, and acknowledgments between every pair of cores. A
|
||||
2D mesh with bisection bandwidth of B bytes/cycle/hop becomes the
|
||||
bottleneck. The memory model's influence is mainly through
|
||||
**invalidation storms**: when one core writes a shared line, all
|
||||
128 sharers (or all 128 caches, for a miss) must be invalidated
|
||||
or notified. The protocol must bound the worst-case
|
||||
invalidation fanout.
|
||||
|
||||
Standard mitigations:
|
||||
|
||||
- **Region coherence / hierarchical directory** (L2 directories,
|
||||
L3 directory).
|
||||
- **Token coherence** (Martin, Hill, Sorin; HPCA 2003), which
|
||||
replaces directory entries with tokens and bounds
|
||||
invalidation traffic at the cost of extra protocol messages.
|
||||
- **In-network coherence** (e.g., the Intel Mesh / Emesh design),
|
||||
which performs directory lookups inside the NoC routers.
|
||||
|
||||
These are all reasonable for a research project but introduce
|
||||
complexity. A baseline MESI directory is the recommended
|
||||
starting point.
|
||||
|
||||
### 4.4 False Sharing and Ordering
|
||||
|
||||
A 128-core machine is particularly sensitive to **false
|
||||
sharing**, where two cores write to different words in the same
|
||||
cache line. Each such write triggers a coherence ping-pong.
|
||||
|
||||
The memory model does not by itself mitigate this — the cache
|
||||
line is the unit of coherence, and any store to it forces a
|
||||
write. XH-1 (proposed) may include hardware support for
|
||||
detecting hot lines, but the primary mitigation is software
|
||||
(cache-line padding, avoiding shared writable data).
|
||||
|
||||
The memory model is relevant for one specific case: a load
|
||||
following a store to the same line in program order must not
|
||||
reorder past the store. In a directory system, the store must
|
||||
be acknowledged before the load can complete. This is a
|
||||
"store-load forwarding" case and is a major source of pipeline
|
||||
stalls.
|
||||
|
||||
### 4.5 Acquire/Release Scalability
|
||||
|
||||
Load-acquire and store-release are the preferred RVWMO
|
||||
synchronization primitives. They are local: they order only
|
||||
with respect to the issuing hart. The hardware cost is a small
|
||||
amount of state in the load/store units to delay the issue of
|
||||
subsequent operations until the acquire is acknowledged.
|
||||
|
||||
At 128 cores, the **aggregate** acquire/release traffic (from
|
||||
synchronization in the OS, runtime, and applications) is
|
||||
significant but not fundamentally different from smaller
|
||||
machines. The scalability limit is the protocol, not the
|
||||
model.
|
||||
|
||||
## 5. Interactions
|
||||
|
||||
### 5.1 Pipeline
|
||||
|
||||
The store buffer and load queue are pipeline structures. The
|
||||
pipeline depth determines how many in-flight loads/stores can
|
||||
coexist. For a high-frequency XH-1 core, the store buffer should
|
||||
hold at least as many entries as the load-to-use latency times
|
||||
the issue rate — typically 16–32 entries for a 4-wide issue.
|
||||
|
||||
The FENCE instruction is a **pipeline-flushing event** in the
|
||||
broadest sense: it must drain the store buffer and/or load queue
|
||||
to the extent required. A FENCE `rw,rw` is the most expensive
|
||||
and may stall the issue queue for tens of cycles. The
|
||||
implementation can optimize by allowing subsequent
|
||||
non-dependent operations to issue (out-of-order) but must
|
||||
constrain loads to a coherent subset until the fence retires.
|
||||
|
||||
### 5.2 Cache Hierarchy
|
||||
|
||||
The cache hierarchy is the agent that enforces the model's
|
||||
single-writer invariant. The L1 must be **write-through** to the
|
||||
L2 or **write-back** with a coherence protocol; in modern
|
||||
designs, write-back with a directory is standard.
|
||||
|
||||
The store buffer sits **between the L1 and the L2**: stores
|
||||
commit to the L1 in program order but the L1 marks the line
|
||||
Modified without necessarily writing it to the L2. The
|
||||
coherence protocol performs the L2 writeback when the line is
|
||||
evicted or when another core requests the line.
|
||||
|
||||
For RVWMO, the L1 store buffer must:
|
||||
|
||||
- Coalesce same-line stores.
|
||||
- Hold the address until the store is acknowledged by the L2
|
||||
(or further).
|
||||
- Service forwarding requests from the load unit.
|
||||
|
||||
For Ztso, the same structure is required, with the
|
||||
simplification that FENCE is a no-op.
|
||||
|
||||
### 5.3 Memory System
|
||||
|
||||
Main memory (DRAM) introduces a new variable: the **DRAM
|
||||
controller** may reorder requests to the same rank/bank to
|
||||
improve bandwidth. This reordering is observable to the cores
|
||||
unless the controller enforces ordering.
|
||||
|
||||
The RISC-V memory model requires that **the system behaves as
|
||||
if there is a single global memory order**. A DRAM controller
|
||||
that reorders must do so in a way that is invisible to the
|
||||
model, which usually means serializing conflicting requests
|
||||
(e.g., reads to the same bank after a write must wait).
|
||||
|
||||
Standard practice: a **write queue** in the memory controller
|
||||
that tracks in-flight requests and serializes conflicting ones.
|
||||
For XH-1 (proposed), the memory controller would include a
|
||||
per-bank ordering queue.
|
||||
|
||||
### 5.4 Interconnect
|
||||
|
||||
The on-chip interconnect (NoC) is the medium through which
|
||||
coherence flows. As discussed in §3.4, the NoC must provide
|
||||
in-order delivery per channel and per-address serialization
|
||||
across channels. A 2D mesh with virtual channels is the
|
||||
conventional choice. For 128 cores, a 16x8 or 8x16 mesh
|
||||
balances wire length and bisection bandwidth.
|
||||
|
||||
The memory model's FENCE instruction interacts with the NoC
|
||||
via **end-to-end acknowledgment**: a FENCE that requires global
|
||||
ordering cannot retire until all messages it has issued have
|
||||
been acknowledged by their destinations. In a directory
|
||||
system, this is a per-message ack from the directory and, for
|
||||
shared lines, from the invalidation recipients.
|
||||
|
||||
### 5.5 Coherence
|
||||
|
||||
The coherence protocol is the **enforcement mechanism** for
|
||||
write atomicity. A memory model without coherence is undefined
|
||||
(the "coherence" axiom in RVWMO and other models assumes that
|
||||
each location has a single most-recent store at any time).
|
||||
|
||||
For XH-1, the recommended baseline is directory-based MESI.
|
||||
Variants such as MOESI (for data forwarding) or MESIF (for
|
||||
forwarding the response) are reasonable. The choice of variant
|
||||
does not change the memory model but affects performance and
|
||||
verification.
|
||||
|
||||
A specific concern for RVWMO: a **store that is followed by a
|
||||
load to the same line** is a critical case. The store must be
|
||||
visible to the load, which in a directory system means the
|
||||
store must complete (be acknowledged) before the load
|
||||
returns. This is enforced by the L1 store buffer: a load
|
||||
misses the L1, the directory is consulted, and the request is
|
||||
serviced from the L1's Modified state (the store) or from
|
||||
another cache (if the L1 already wrote the line to L2).
|
||||
|
||||
### 5.6 Interrupts and Traps
|
||||
|
||||
The RISC-V memory model requires that **interrupts and traps
|
||||
see a consistent view of memory**. Specifically:
|
||||
|
||||
- A load that traps (e.g., page fault) must see the
|
||||
architectural state as if the load had not occurred (or, for
|
||||
certain traps, as if the load had occurred but no later
|
||||
operation had).
|
||||
- The model must specify whether the load is "performed" with
|
||||
respect to other harts at the time of the trap.
|
||||
|
||||
In RVWMO, an unsuccessful load (one that traps) is not
|
||||
considered to have been performed. A successful load is
|
||||
considered performed in program order relative to other loads
|
||||
and stores of the same hart, under the PPO rules. The
|
||||
architectural state at the time of the trap is the state
|
||||
before the load (and before any subsequent operation in
|
||||
program order).
|
||||
|
||||
For XH-1, this is implemented by the pipeline: an
|
||||
exception-causing load does not write back to a register, and
|
||||
its store buffer / load queue entries are marked invalid.
|
||||
Subsequent instructions in program order are squashed.
|
||||
|
||||
A **subtle case** is the I/O fence: a FENCE before a non-
|
||||
idempotent device write must be retired before the write is
|
||||
issued, even if an interrupt arrives between them. RVWMO
|
||||
permits the interrupt to be taken at any time but requires
|
||||
the FENCE to be retired before the write is observable to the
|
||||
device.
|
||||
|
||||
### 5.7 Operating System
|
||||
|
||||
The operating system relies on the memory model for
|
||||
synchronization:
|
||||
|
||||
- **Spinlocks**: typically implemented as `LR.W.aq` /
|
||||
`SC.W.rl` or as `AMOSWAP.W.aqrl`. Both are acquire/release
|
||||
and are well supported by RVWMO.
|
||||
- **I/O fences**: required before/after device accesses.
|
||||
Linux/RISC-V includes these in its I/O macros.
|
||||
- **Page table updates**: require a release fence before
|
||||
updating the PTE and an acquire fence (or
|
||||
load-acquire) after, so that the TLB refill sees the
|
||||
updated PTE.
|
||||
- **Inter-processor interrupts (IPIs)**: require a release
|
||||
fence before sending the IPI and an acquire fence after
|
||||
receiving it, so that the data and the IPI are ordered
|
||||
correctly.
|
||||
- **Boot and shutdown**: require global fences, which on
|
||||
RISC-V must be composed from local FENCE + AMO + IPI
|
||||
patterns. This is more complex than on x86/TSO.
|
||||
|
||||
The OS for XH-1 (proposed) is likely to be Linux/RISC-V or a
|
||||
research kernel. Either way, the memory model requirements
|
||||
are well documented in the kernel's memory-barrier
|
||||
documentation, and XH-1's implementation must support them.
|
||||
|
||||
### 5.8 Verification
|
||||
|
||||
Verifying the memory model is one of the hardest parts of
|
||||
modern CPU design. The standard approach is:
|
||||
|
||||
1. **Litmus tests**: small programs designed to expose
|
||||
reordering bugs. The RISC-V memory model has a
|
||||
comprehensive test suite (see
|
||||
https://github.com/litmus-tests/litmus-tests-riscv).
|
||||
2. **Model checking**: tools such as Herd7, diy7, and
|
||||
MemAlloy can check the axiomatic model against
|
||||
implementation behaviors.
|
||||
3. **Random instruction stream testing (RIST)**: at the
|
||||
RTL level, run randomized tests and check the
|
||||
observable orderings against the model.
|
||||
|
||||
For 128 cores, the verification surface grows combinatorially
|
||||
in the number of cores. Standard practice is to verify the
|
||||
single-core ordering first, then verify the coherence
|
||||
protocol (which involves multiple cores), and finally
|
||||
verify the cross-core ordering at the SoC level using
|
||||
co-simulation or formal methods.
|
||||
|
||||
XH-1 (proposed) should adopt a multi-level verification
|
||||
strategy: litmus tests for unit-level, formal verification
|
||||
for the coherence protocol (using a tool such as CCSS or
|
||||
Murphi), and full-system testing for the SoC. The
|
||||
verification of the memory model is a **major
|
||||
undertaking** and is often the gating item for tape-out.
|
||||
|
||||
### 5.9 Performance
|
||||
|
||||
The memory model's impact on performance is workload-
|
||||
dependent:
|
||||
|
||||
- **Sequential programs** (single hart): the cost is the
|
||||
store-load forwarding latency (a few cycles) and the
|
||||
cost of fences (which are usually rare).
|
||||
- **Parallel programs**: the cost is the latency of
|
||||
acquire/release operations, which is typically a few
|
||||
cycles for an L1 hit and tens of cycles for a remote
|
||||
cache hit.
|
||||
- **Highly contended workloads**: the cost is the
|
||||
latency of a failed `SC` (which may require a
|
||||
pipeline flush on some implementations) and the
|
||||
cost of cache-line ping-pong.
|
||||
|
||||
The RVWMO model is generally **faster** than TSO for
|
||||
parallel workloads because it allows more reordering
|
||||
between loads and between stores. The Ztso extension is
|
||||
provided for software that wants TSO semantics.
|
||||
|
||||
For a 128-core XH-1, the dominant performance concern is
|
||||
not the memory model but the **coherence traffic and
|
||||
interconnect bandwidth**. The memory model is a
|
||||
secondary concern, but a poorly implemented fence can
|
||||
serialize the entire machine (e.g., a global fence in
|
||||
the OS boot path).
|
||||
|
||||
## 6. Alternatives and Trade-offs
|
||||
|
||||
### 6.1 TSO (Ztso) vs. RVWMO
|
||||
|
||||
- **TSO advantages**: simpler verification, easier OS
|
||||
port, smaller fence frequency.
|
||||
- **TSO disadvantages**: slightly more conservative
|
||||
reordering, which costs performance on workloads
|
||||
with relaxed dependencies.
|
||||
|
||||
XH-1 (proposed) is recommended to support both: RVWMO
|
||||
as the default, with Ztso as an optional hart-level
|
||||
feature. (As above, the CSR-based opt-in mechanism is
|
||||
**open**.)
|
||||
|
||||
### 6.2 Directory MESI vs. Token Coherence
|
||||
|
||||
- **MESI directory advantages**: well understood, simple
|
||||
protocol, predictable latency.
|
||||
- **MESI directory disadvantages**: directory storage
|
||||
cost (16 MiB for 64 MiB L3, 128-bit sharers),
|
||||
invalidation storms for shared lines.
|
||||
- **Token coherence advantages**: bounded traffic,
|
||||
no directory storage.
|
||||
- **Token coherence disadvantages**: more complex
|
||||
protocol, more messages per coherence event,
|
||||
less predictable latency.
|
||||
|
||||
For a research project, MESI is the recommended baseline.
|
||||
Token coherence is a more advanced option.
|
||||
|
||||
### 6.3 Coarse-Vector vs. Full-Map Directory
|
||||
|
||||
- **Coarse-vector advantages**: less storage, similar
|
||||
performance on most workloads.
|
||||
- **Coarse-vector disadvantages**: occasional broadcast
|
||||
fallback for over-subscribed lines.
|
||||
- **Full-map advantages**: no broadcast fallback.
|
||||
- **Full-map disadvantages**: 16 MiB of directory state
|
||||
for 64 MiB L3.
|
||||
|
||||
For XH-1 (proposed), a coarse vector with 16 pointers per
|
||||
entry plus broadcast fallback is the recommended baseline.
|
||||
|
||||
### 6.4 In-Network Coherence vs. Traditional Coherence
|
||||
|
||||
- **In-network advantages**: lower latency, lower
|
||||
directory storage.
|
||||
- **In-network disadvantages**: requires custom routers,
|
||||
harder verification, less mature tooling.
|
||||
|
||||
This is a research-level option and not recommended for
|
||||
the first version of XH-1.
|
||||
|
||||
### 6.5 Fence Implementation: Strict vs. Optimistic
|
||||
|
||||
- **Strict**: FENCE always waits for the full
|
||||
predecessor set to drain at the system level.
|
||||
- **Optimistic**: FENCE may retire as soon as the
|
||||
predecessor set is locally drained, and the system
|
||||
guarantees that the effect of the FENCE is preserved
|
||||
even if the global drain has not completed.
|
||||
|
||||
RVWMO does not require a global drain for FENCE; the
|
||||
hardware may retire the FENCE as soon as the local
|
||||
guarantees are met. The OS and synchronization libraries
|
||||
that need global ordering must compose a global fence
|
||||
explicitly. XH-1 (proposed) is recommended to follow
|
||||
this approach.
|
||||
|
||||
## 7. Advantages and Disadvantages of the Proposed Approach
|
||||
|
||||
### Advantages
|
||||
|
||||
- **Standards compliance**: implements RVWMO per the
|
||||
RISC-V specification.
|
||||
- **Familiar**: uses conventional MESI directory, store
|
||||
buffer, load queue structures.
|
||||
- **Scalable**: coarse-vector directory bounds storage,
|
||||
2D mesh NoC provides sufficient bandwidth.
|
||||
- **Verifiable**: litmus tests, formal verification of
|
||||
the protocol, full-system co-simulation.
|
||||
- **OS friendly**: supports Linux/RISC-V's memory-
|
||||
barrier requirements.
|
||||
|
||||
### Disadvantages
|
||||
|
||||
- **Verification cost**: 128 cores × many memory model
|
||||
rules is a large verification surface.
|
||||
- **Interconnect complexity**: 2D mesh with virtual
|
||||
channels, per-router flow control, per-channel
|
||||
ordering.
|
||||
- **Directory cost**: 16 MiB of directory state for a
|
||||
64 MiB L3.
|
||||
- **Fence overhead**: a global fence is more
|
||||
expensive than on smaller machines.
|
||||
- **OS porting burden**: global fences must be
|
||||
composed from local FENCE + AMO patterns.
|
||||
|
||||
## 8. Unresolved Design Questions
|
||||
|
||||
1. **Zalasr status**: the Zalasr extension is **not yet
|
||||
ratified** at the time of writing. XH-1 should not
|
||||
rely on it until the extension is finalized.
|
||||
2. **Ztso opt-in mechanism**: should XH-1 provide a
|
||||
CSR-based opt-in for TSO, or should TSO be a
|
||||
per-hart architectural mode set at design time? (No
|
||||
standard mechanism exists; this is a custom decision.)
|
||||
3. **Fence.TSO vs. Ztso**: should the `FENCE.TSO`
|
||||
encoding (`fm=0, pred=0, succ=0`) be implemented as
|
||||
a strict fence or as a hint? The spec allows either.
|
||||
4. **Cache line size**: 64 B is conventional, 128 B
|
||||
reduces directory storage but worsens false
|
||||
sharing. Not yet decided for XH-1.
|
||||
5. **L3 size and directory configuration**: 64 MiB
|
||||
with coarse vector (16 pointers) is a proposal;
|
||||
the actual size and configuration are open.
|
||||
6. **Custom global fence**: should XH-1 provide a
|
||||
custom opcode for a global fence? Open.
|
||||
7. **Token coherence**: should XH-1 adopt token
|
||||
coherence to bound invalidation traffic? Not
|
||||
recommended for v1; deferred.
|
||||
8. **Device memory ordering**: how should the I/O
|
||||
fence be encoded in the compiler and the OS?
|
||||
This is a software question, but the hardware
|
||||
must support the required ordering.
|
||||
|
||||
## 9. Summary of Recommendations
|
||||
|
||||
Based on the analysis above, the following are
|
||||
recommended for XH-1 (all **proposals** pending
|
||||
design review):
|
||||
|
||||
- **Memory model**: RVWMO as the base. Support Ztso
|
||||
as an optional per-hart feature.
|
||||
- **Coherence protocol**: directory MESI with a
|
||||
coarse-vector directory (16 pointers per entry)
|
||||
and broadcast fallback.
|
||||
- **Cache hierarchy**: write-back L1, directory-
|
||||
tracked L2 and L3. Cache line size 64 B (open).
|
||||
- **Interconnect**: 2D mesh NoC with virtual channels,
|
||||
XY routing, per-router flow control.
|
||||
- **Per-hart structures**: 16–32 entry store buffer,
|
||||
32–64 entry load queue.
|
||||
- **Fence implementation**: strict local drain per
|
||||
the FENCE operand set. Global fences must be
|
||||
composed in software.
|
||||
- **Acquisition/release**: implement AMO.acquire /
|
||||
AMO.release, LR.aq, SC.rl. Zalasr when ratified.
|
||||
- **Verification**: litmus tests, formal protocol
|
||||
verification, full-system co-simulation. RIST
|
||||
at the RTL level.
|
||||
- **I/O ordering**: I/O fences for all device memory
|
||||
accesses, encoded by the compiler and the OS.
|
||||
|
||||
## 10. Status of This Document
|
||||
|
||||
This document is a **research document** for the XH-1
|
||||
project. It is not a final specification. All XH-1
|
||||
implementation choices are marked as proposals. The
|
||||
RISC-V specification is the authoritative source for
|
||||
architectural requirements; this document supplements
|
||||
but does not replace it.
|
||||
|
||||
Where evidence is unavailable, the document explicitly
|
||||
says so. Where standards or papers are cited, they are
|
||||
real (RISC-V manuals, Sorin/Hill/Wood, Cuesta et al.,
|
||||
litmus-tests-riscv), and the citation is provided
|
||||
inline or in the references below. Specific micro-
|
||||
architectural constants (e.g., 16–32 store buffer
|
||||
entries) are typical industry values but are not
|
||||
defined by the RISC-V specification; they are
|
||||
proposals, not requirements.
|
||||
|
||||
## References
|
||||
|
||||
1. RISC-V International, *The RISC-V Instruction Set
|
||||
Manual, Volume I: Unprivileged ISA*, current
|
||||
version.
|
||||
2. RISC-V International, *The RISC-V Instruction Set
|
||||
Manual, Volume II: Privileged ISA*, current
|
||||
version.
|
||||
3. RISC-V International, *RISC-V Memory Model*, in
|
||||
the same manual package.
|
||||
4. Adve and Gharachorloo, "Shared Memory
|
||||
Consistency Models: A Tutorial", IEEE Computer,
|
||||
1996.
|
||||
5. Sorin, Hill, and Wood, *A Primer on Memory
|
||||
Consistency and Cache Coherence*, Synthesis
|
||||
Lectures on Computer Architecture, Morgan &
|
||||
Claypool, 2011.
|
||||
6. SPARC International, *The SPARC Architecture
|
||||
Manual, Version 9*, 1994 (for the TSO
|
||||
definition).
|
||||
7. Cuesta et al., "Increasing the Effectiveness of
|
||||
Directory Caches by Deactivating Coarse-Grained
|
||||
Coherence Tracking", ICS 2011.
|
||||
8. Martin, Hill, and Sorin, "Token Coherence:
|
||||
Decoupling Performance and Correctness", ISCA
|
||||
2003.
|
||||
9. The litmus tests repository for RISC-V,
|
||||
https://github.com/litmus-tests/litmus-tests-riscv.
|
||||
10. The Herd7 / diy7 tool suite by ARM / Cambridge,
|
||||
https://github.com/herd/herdtools7.
|
||||
11. Greathouse et al., "Open-Source Hardware
|
||||
Verification: RISC-V and Beyond", (see
|
||||
riscv-formal).
|
||||
12. RISC-V Formal Verification framework,
|
||||
https://github.com/YosysHQ/riscv-formal.
|
||||
```
|
||||
|
||||
Reference in New Issue
Block a user