PASSED: 4x | FAILED: 1X

We have continued with the testing/research. We have seen a gigantic improvement with the `PASSING` tests. When we changed the objective agents/models to more careful model, we had a failed test after 3 continue attempts. We will continue with this test, the failed test was: `research/05-memory/atomics.md`. We also published the `issues` which are both capable of displaying failed and passed tests: https://git.revivesparc.com/riscvcxh1/xh1-research-fork/issues.

We are going to continue more tests, and make sure that the `atomics` research will pass.
This commit is contained in:
riscvcxh1
2026-08-26 16:20:34 +02:00
committed by GitHub
184 changed files with 17264 additions and 36 deletions
Executable
+1929
View File
File diff suppressed because it is too large Load Diff
+429
View File
@@ -320,3 +320,432 @@
[2026-08-26T10:34:46Z] Maximum research rounds reached. [2026-08-26T10:34:46Z] Maximum research rounds reached.
[2026-08-26T10:34:46Z] Leaving original document unchanged. [2026-08-26T10:34:46Z] Leaving original document unchanged.
[2026-08-26T10:34:46Z] Marking topic failed for this run. [2026-08-26T10:34:46Z] Marking topic failed for this run.
[2026-08-26T12:11:25Z] Started run: 20260826T121125Z
[2026-08-26T12:11:25Z] ==================================================
[2026-08-26T12:11:25Z] Researching: research/05-memory/memory-ordering.md
[2026-08-26T12:11:25Z] ==================================================
[2026-08-26T12:11:25Z] Research round 1/3
[2026-08-26T12:11:25Z] Running researcher.
[2026-08-26T12:11:25Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=1
[2026-08-26T12:11:25Z] API request failed.
[2026-08-26T12:11:26Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=2
[2026-08-26T12:11:26Z] API request failed.
[2026-08-26T12:11:58Z] Started run: 20260826T121158Z
[2026-08-26T12:11:58Z] ==================================================
[2026-08-26T12:11:58Z] Researching: research/05-memory/memory-ordering.md
[2026-08-26T12:11:58Z] ==================================================
[2026-08-26T12:11:58Z] Research round 1/3
[2026-08-26T12:11:58Z] Running researcher.
[2026-08-26T12:11:58Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=1
[2026-08-26T12:12:47Z] Response received: 19485 bytes
[2026-08-26T12:12:47Z] Running reviewer.
[2026-08-26T12:12:47Z] API: provider=ollama model=gemma4:12b attempt=1
[2026-08-26T12:15:45Z] Response received: 1622 bytes
[2026-08-26T12:15:45Z] Research rejected by reviewer.
[2026-08-26T12:15:45Z] Preparing revision round 2.
[2026-08-26T12:15:50Z] Research round 2/3
[2026-08-26T12:15:50Z] Running revision agent.
[2026-08-26T12:15:50Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=1
[2026-08-26T12:15:51Z] API request failed.
[2026-08-26T12:15:52Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=2
[2026-08-26T12:16:25Z] Response received: 23244 bytes
[2026-08-26T12:16:25Z] Running reviewer.
[2026-08-26T12:16:25Z] API: provider=ollama model=gemma4:12b attempt=1
[2026-08-26T12:19:32Z] Response received: 1 bytes
[2026-08-26T12:19:32Z] Research rejected by reviewer.
[2026-08-26T12:19:32Z] Preparing revision round 3.
[2026-08-26T12:19:37Z] Research round 3/3
[2026-08-26T12:19:37Z] Running revision agent.
[2026-08-26T12:19:37Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=1
[2026-08-26T12:19:38Z] API request failed.
[2026-08-26T12:19:39Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=2
[2026-08-26T12:19:44Z] API request failed.
[2026-08-26T12:19:48Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=3
[2026-08-26T12:20:25Z] Response received: 25475 bytes
[2026-08-26T12:20:25Z] Running reviewer.
[2026-08-26T12:20:25Z] API: provider=ollama model=gemma4:12b attempt=1
[2026-08-26T12:23:32Z] Response received: 1 bytes
[2026-08-26T12:23:32Z] Research rejected by reviewer.
[2026-08-26T12:23:32Z] Maximum research rounds reached.
[2026-08-26T12:23:32Z] Leaving original document unchanged.
[2026-08-26T12:23:32Z] Marking topic failed for this run.
[2026-08-26T12:25:59Z] Started run: 20260826T122559Z
[2026-08-26T12:25:59Z] ==================================================
[2026-08-26T12:25:59Z] Researching: research/05-memory/memory-ordering.md
[2026-08-26T12:25:59Z] ==================================================
[2026-08-26T12:25:59Z] Research round 1/3
[2026-08-26T12:25:59Z] Running researcher.
[2026-08-26T12:25:59Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=1
[2026-08-26T12:26:00Z] API HTTP status: 429
[2026-08-26T12:26:00Z] WARNING: Rate limited by openrouter.
[2026-08-26T12:26:00Z] Waiting 5s before retry.
[2026-08-26T12:26:05Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=2
[2026-08-26T12:26:06Z] API HTTP status: 429
[2026-08-26T12:26:06Z] WARNING: Rate limited by openrouter.
[2026-08-26T12:26:06Z] Waiting 20s before retry.
[2026-08-26T12:26:26Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=3
[2026-08-26T12:27:36Z] Response received: 23159 bytes
[2026-08-26T12:27:36Z] Running reviewer.
[2026-08-26T12:27:37Z] API: provider=ollama model=gemma4:12b attempt=1
[2026-08-26T12:30:29Z] Response received: 4699 bytes
[2026-08-26T12:30:29Z] WARNING: Reviewer returned no valid VERDICT.
[2026-08-26T12:30:29Z] WARNING: This is a reviewer failure, not a research failure.
[2026-08-26T12:33:03Z] Started run: 20260826T123303Z
[2026-08-26T12:33:03Z] ==================================================
[2026-08-26T12:33:03Z] Researching: research/05-memory/memory-ordering.md
[2026-08-26T12:33:03Z] ==================================================
[2026-08-26T12:33:03Z] Research round 1/3
[2026-08-26T12:33:03Z] Running researcher.
[2026-08-26T12:33:03Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=1
[2026-08-26T12:33:04Z] API HTTP status: 429
[2026-08-26T12:33:04Z] WARNING: Rate limited by openrouter.
[2026-08-26T12:33:04Z] Waiting 5s before retry.
[2026-08-26T12:33:09Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=2
[2026-08-26T12:35:05Z] Response received: 41239 bytes
[2026-08-26T12:35:05Z] Running reviewer.
[2026-08-26T12:35:05Z] API: provider=ollama model=gemma4:12b attempt=1
[2026-08-26T12:38:10Z] Response received: 2000 bytes
[2026-08-26T12:38:10Z] Reviewer did not provide a machine-readable verdict.
[2026-08-26T12:38:10Z] Requesting verdict-format repair.
[2026-08-26T12:38:10Z] API: provider=ollama model=gemma4:12b attempt=1
[2026-08-26T12:38:52Z] Response received: 87 bytes
[2026-08-26T12:38:52Z] WARNING: Reviewer returned no valid VERDICT.
[2026-08-26T12:38:52Z] WARNING: This is a reviewer failure, not a research failure.
[2026-08-26T12:39:54Z] Started run: 20260826T123954Z
[2026-08-26T12:39:54Z] ==================================================
[2026-08-26T12:39:54Z] Researching: research/05-memory/memory-ordering.md
[2026-08-26T12:39:54Z] ==================================================
[2026-08-26T12:39:54Z] Research round 1/3
[2026-08-26T12:39:54Z] Running researcher.
[2026-08-26T12:39:54Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=1
[2026-08-26T12:41:33Z] Response received: 38665 bytes
[2026-08-26T12:41:33Z] Running reviewer.
[2026-08-26T12:41:33Z] API: provider=ollama model=gemma4:12b attempt=1
[2026-08-26T12:44:38Z] Response received: 2266 bytes
[2026-08-26T12:44:38Z] Reviewer did not provide a machine-readable verdict.
[2026-08-26T12:44:38Z] Requesting verdict-format repair.
[2026-08-26T12:44:38Z] API: provider=ollama model=gemma4:12b attempt=1
[2026-08-26T12:45:19Z] Response received: 87 bytes
[2026-08-26T12:45:19Z] Recovered reviewer verdict: PASS
[2026-08-26T12:45:19Z] Research passed review.
[2026-08-26T12:45:19Z] Accepted research document: research/05-memory/memory-ordering.md
[2026-08-26T12:48:21Z] Started run: 20260826T124821Z
[2026-08-26T12:48:21Z] ==================================================
[2026-08-26T12:48:21Z] Researching: research/05-memory/cache-coherency.md
[2026-08-26T12:48:21Z] ==================================================
[2026-08-26T12:48:21Z] Research round 1/3
[2026-08-26T12:48:21Z] Running researcher.
[2026-08-26T12:48:21Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=1
[2026-08-26T12:49:42Z] Response received: 27828 bytes
[2026-08-26T12:49:42Z] Running reviewer.
[2026-08-26T12:49:42Z] API: provider=ollama model=gemma4:12b attempt=1
[2026-08-26T12:52:25Z] Response received: 2096 bytes
[2026-08-26T12:52:25Z] Research passed review.
[2026-08-26T12:52:25Z] Accepted research document: research/05-memory/cache-coherency.md
[2026-08-26T12:54:26Z] Started run: 20260826T125426Z
[2026-08-26T12:54:26Z] ==================================================
[2026-08-26T12:54:26Z] Researching: research/05-memory/memory-hierarchy.md
[2026-08-26T12:54:26Z] ==================================================
[2026-08-26T12:54:26Z] Research round 1/3
[2026-08-26T12:54:26Z] Running researcher.
[2026-08-26T12:54:26Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=1
[2026-08-26T12:54:28Z] API HTTP status: 429
[2026-08-26T12:54:28Z] WARNING: Rate limited by openrouter.
[2026-08-26T12:54:28Z] Waiting 5s before retry.
[2026-08-26T12:54:33Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=2
[2026-08-26T12:56:00Z] Response received: 30493 bytes
[2026-08-26T12:56:00Z] Running reviewer.
[2026-08-26T12:56:00Z] API: provider=ollama model=gemma4:12b attempt=1
[2026-08-26T12:56:50Z] Response received: 1780 bytes
[2026-08-26T12:56:50Z] Research passed review.
[2026-08-26T12:56:50Z] Accepted research document: research/05-memory/memory-hierarchy.md
[2026-08-26T12:57:40Z] Started run: 20260826T125740Z
[2026-08-26T12:57:40Z] ==================================================
[2026-08-26T12:57:40Z] Researching: research/05-memory/memory-architecture.md
[2026-08-26T12:57:40Z] ==================================================
[2026-08-26T12:57:40Z] Research round 1/3
[2026-08-26T12:57:40Z] Running researcher.
[2026-08-26T12:57:40Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=1
[2026-08-26T12:58:59Z] Response received: 28405 bytes
[2026-08-26T12:58:59Z] Running reviewer.
[2026-08-26T12:58:59Z] API: provider=ollama model=gemma4:12b attempt=1
[2026-08-26T13:02:07Z] WARNING: Provider returned an empty/near-empty response (0 bytes).
[2026-08-26T13:02:09Z] API: provider=ollama model=gemma4:12b attempt=2
[2026-08-26T13:02:54Z] Response received: 2234 bytes
[2026-08-26T13:02:54Z] Research passed review.
[2026-08-26T13:02:54Z] Accepted research document: research/05-memory/memory-architecture.md
[2026-08-26T13:19:26Z] Started run: 20260826T131926Z
[2026-08-26T13:19:26Z] ==================================================
[2026-08-26T13:19:26Z] Researching: research/05-memory/memory-bandwidth.md
[2026-08-26T13:19:26Z] ==================================================
[2026-08-26T13:19:26Z] Research round 1/3
[2026-08-26T13:19:26Z] Running researcher.
[2026-08-26T13:19:27Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=1
[2026-08-26T13:20:19Z] Response received: 20125 bytes
[2026-08-26T13:20:19Z] Running reviewer.
[2026-08-26T13:20:19Z] API: provider=qwen model=qwen3.7-plus attempt=1
[2026-08-26T13:20:19Z] API HTTP status: 000
[2026-08-26T13:20:19Z] WARNING: Network/API connection failure.
[2026-08-26T13:20:22Z] API: provider=qwen model=qwen3.7-plus attempt=2
[2026-08-26T13:20:22Z] API HTTP status: 000
[2026-08-26T13:20:22Z] WARNING: Network/API connection failure.
[2026-08-26T13:20:28Z] API: provider=qwen model=qwen3.7-plus attempt=3
[2026-08-26T13:20:28Z] API HTTP status: 000
[2026-08-26T13:20:28Z] WARNING: Network/API connection failure.
[2026-08-26T13:21:34Z] Started run: 20260826T132134Z
[2026-08-26T13:21:34Z] ==================================================
[2026-08-26T13:21:34Z] Researching: research/05-memory/memory-bandwidth.md
[2026-08-26T13:21:34Z] ==================================================
[2026-08-26T13:21:34Z] Research round 1/3
[2026-08-26T13:21:34Z] Running researcher.
[2026-08-26T13:21:34Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=1
[2026-08-26T13:21:34Z] API HTTP status: 401
[2026-08-26T13:21:34Z] WARNING: Client/API error from openrouter.
[2026-08-26T13:21:37Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=2
[2026-08-26T13:21:38Z] API HTTP status: 401
[2026-08-26T13:21:38Z] WARNING: Client/API error from openrouter.
[2026-08-26T13:21:44Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=3
[2026-08-26T13:21:44Z] API HTTP status: 401
[2026-08-26T13:21:44Z] WARNING: Client/API error from openrouter.
[2026-08-26T13:21:44Z] WARNING: API request exhausted all retries.
[2026-08-26T13:21:44Z] WARNING: Research API failed.
[2026-08-26T13:21:44Z] WARNING: Topic will NOT be marked as failed.
[2026-08-26T13:21:46Z] Started run: 20260826T132146Z
[2026-08-26T13:21:46Z] ==================================================
[2026-08-26T13:21:46Z] Researching: research/05-memory/memory-bandwidth.md
[2026-08-26T13:21:46Z] ==================================================
[2026-08-26T13:21:46Z] Research round 1/3
[2026-08-26T13:21:46Z] Running researcher.
[2026-08-26T13:21:46Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=1
[2026-08-26T13:21:47Z] API HTTP status: 401
[2026-08-26T13:21:47Z] WARNING: Client/API error from openrouter.
[2026-08-26T13:23:02Z] Started run: 20260826T132302Z
[2026-08-26T13:23:02Z] ==================================================
[2026-08-26T13:23:02Z] Researching: research/05-memory/memory-bandwidth.md
[2026-08-26T13:23:02Z] ==================================================
[2026-08-26T13:23:02Z] Research round 1/3
[2026-08-26T13:23:02Z] Running researcher.
[2026-08-26T13:23:02Z] API: provider=qwen model=qwen3.8-max attempt=1
[2026-08-26T13:23:02Z] API HTTP status: 000
[2026-08-26T13:23:02Z] WARNING: Network/API connection failure.
[2026-08-26T13:23:05Z] API: provider=qwen model=qwen3.8-max attempt=2
[2026-08-26T13:23:05Z] API HTTP status: 000
[2026-08-26T13:23:05Z] WARNING: Network/API connection failure.
[2026-08-26T13:23:29Z] Started run: 20260826T132329Z
[2026-08-26T13:23:29Z] ==================================================
[2026-08-26T13:23:29Z] Researching: research/05-memory/memory-bandwidth.md
[2026-08-26T13:23:29Z] ==================================================
[2026-08-26T13:23:29Z] Research round 1/3
[2026-08-26T13:23:29Z] Running researcher.
[2026-08-26T13:23:29Z] API: provider=qwen model=qwen3.8-max attempt=1
[2026-08-26T13:23:29Z] API HTTP status: 000
[2026-08-26T13:23:29Z] WARNING: Network/API connection failure.
[2026-08-26T13:25:56Z] Started run: 20260826T132556Z
[2026-08-26T13:25:56Z] ==================================================
[2026-08-26T13:25:56Z] Researching: research/05-memory/memory-bandwidth.md
[2026-08-26T13:25:56Z] ==================================================
[2026-08-26T13:25:56Z] Research round 1/3
[2026-08-26T13:25:56Z] Running researcher.
[2026-08-26T13:25:56Z] API: provider=qwen model=qwen3.8-max attempt=1
[2026-08-26T13:25:56Z] API HTTP status: 000
[2026-08-26T13:25:56Z] WARNING: Network/API connection failure.
[2026-08-26T13:26:07Z] Started run: 20260826T132607Z
[2026-08-26T13:26:07Z] ==================================================
[2026-08-26T13:26:07Z] Researching: research/05-memory/memory-bandwidth.md
[2026-08-26T13:26:07Z] ==================================================
[2026-08-26T13:26:07Z] Research round 1/3
[2026-08-26T13:26:07Z] Running researcher.
[2026-08-26T13:26:07Z] API: provider=qwen model=qwen3.8-max attempt=1
[2026-08-26T13:26:07Z] API HTTP status: 000
[2026-08-26T13:26:07Z] WARNING: Network/API connection failure.
[2026-08-26T13:26:13Z] Started run: 20260826T132613Z
[2026-08-26T13:26:13Z] ==================================================
[2026-08-26T13:26:13Z] Researching: research/05-memory/memory-bandwidth.md
[2026-08-26T13:26:13Z] ==================================================
[2026-08-26T13:26:13Z] Research round 1/3
[2026-08-26T13:26:13Z] Running researcher.
[2026-08-26T13:26:13Z] API: provider=qwen model=qwen3.8-max attempt=1
[2026-08-26T13:26:13Z] API HTTP status: 000
[2026-08-26T13:26:13Z] WARNING: Network/API connection failure.
[2026-08-26T13:26:50Z] Started run: 20260826T132650Z
[2026-08-26T13:26:50Z] ==================================================
[2026-08-26T13:26:50Z] Researching: research/05-memory/memory-bandwidth.md
[2026-08-26T13:26:50Z] ==================================================
[2026-08-26T13:26:50Z] Research round 1/3
[2026-08-26T13:26:50Z] Running researcher.
[2026-08-26T13:26:50Z] API: provider=qwen model=qwen3.8-max attempt=1
[2026-08-26T13:26:50Z] API HTTP status: 000
[2026-08-26T13:26:50Z] WARNING: Network/API connection failure.
[2026-08-26T13:27:40Z] Started run: 20260826T132740Z
[2026-08-26T13:27:40Z] ==================================================
[2026-08-26T13:27:40Z] Researching: research/05-memory/memory-bandwidth.md
[2026-08-26T13:27:40Z] ==================================================
[2026-08-26T13:27:40Z] Research round 1/3
[2026-08-26T13:27:40Z] Running researcher.
[2026-08-26T13:27:40Z] API: provider=qwen model=qwen3.8-max attempt=1
[2026-08-26T13:27:40Z] API HTTP status: 000
[2026-08-26T13:27:40Z] WARNING: Network/API connection failure.
[2026-08-26T13:27:57Z] Started run: 20260826T132757Z
[2026-08-26T13:27:57Z] ==================================================
[2026-08-26T13:27:57Z] Researching: research/05-memory/memory-bandwidth.md
[2026-08-26T13:27:57Z] ==================================================
[2026-08-26T13:27:57Z] Research round 1/3
[2026-08-26T13:27:57Z] Running researcher.
[2026-08-26T13:27:57Z] API: provider=qwen model=qwen3.8-max attempt=1
[2026-08-26T13:27:57Z] API HTTP status: 000
[2026-08-26T13:27:57Z] WARNING: Network/API connection failure.
[2026-08-26T13:28:46Z] Started run: 20260826T132846Z
[2026-08-26T13:28:46Z] ==================================================
[2026-08-26T13:28:46Z] Researching: research/05-memory/memory-bandwidth.md
[2026-08-26T13:28:46Z] ==================================================
[2026-08-26T13:28:46Z] Research round 1/3
[2026-08-26T13:28:46Z] Running researcher.
[2026-08-26T13:28:46Z] API: provider=qwen model=qwen3.8-max attempt=1
[2026-08-26T13:29:00Z] API HTTP status: 401
[2026-08-26T13:29:00Z] WARNING: Client/API error from qwen.
[2026-08-26T13:29:03Z] API: provider=qwen model=qwen3.8-max attempt=2
[2026-08-26T13:29:14Z] Started run: 20260826T132914Z
[2026-08-26T13:29:14Z] ==================================================
[2026-08-26T13:29:14Z] Researching: research/05-memory/memory-bandwidth.md
[2026-08-26T13:29:14Z] ==================================================
[2026-08-26T13:29:14Z] Research round 1/3
[2026-08-26T13:29:14Z] Running researcher.
[2026-08-26T13:29:14Z] API: provider=qwen model=qwen3.8-max attempt=1
[2026-08-26T13:29:15Z] API HTTP status: 401
[2026-08-26T13:29:15Z] WARNING: Client/API error from qwen.
[2026-08-26T13:29:59Z] Started run: 20260826T132959Z
[2026-08-26T13:29:59Z] ==================================================
[2026-08-26T13:29:59Z] Researching: research/05-memory/memory-bandwidth.md
[2026-08-26T13:29:59Z] ==================================================
[2026-08-26T13:29:59Z] Research round 1/3
[2026-08-26T13:29:59Z] Running researcher.
[2026-08-26T13:29:59Z] API: provider=qwen model=qwen3.8-max attempt=1
[2026-08-26T13:30:00Z] API HTTP status: 401
[2026-08-26T13:30:00Z] WARNING: Client/API error from qwen.
[2026-08-26T13:30:40Z] Started run: 20260826T133040Z
[2026-08-26T13:30:40Z] ==================================================
[2026-08-26T13:30:40Z] Researching: research/05-memory/memory-bandwidth.md
[2026-08-26T13:30:40Z] ==================================================
[2026-08-26T13:30:40Z] Research round 1/3
[2026-08-26T13:30:40Z] Running researcher.
[2026-08-26T13:30:40Z] API: provider=qwen model=qwen3.8-max attempt=1
[2026-08-26T13:30:42Z] API HTTP status: 401
[2026-08-26T13:30:42Z] WARNING: Client/API error from qwen.
[2026-08-26T13:30:59Z] Started run: 20260826T133059Z
[2026-08-26T13:30:59Z] ==================================================
[2026-08-26T13:30:59Z] Researching: research/05-memory/memory-bandwidth.md
[2026-08-26T13:30:59Z] ==================================================
[2026-08-26T13:30:59Z] Research round 1/3
[2026-08-26T13:30:59Z] Running researcher.
[2026-08-26T13:30:59Z] API: provider=qwen model=qwen3.8-max attempt=1
[2026-08-26T13:31:00Z] API HTTP status: 401
[2026-08-26T13:31:00Z] WARNING: Client/API error from qwen.
[2026-08-26T13:31:32Z] Started run: 20260826T133132Z
[2026-08-26T13:31:32Z] ==================================================
[2026-08-26T13:31:32Z] Researching: research/05-memory/memory-bandwidth.md
[2026-08-26T13:31:32Z] ==================================================
[2026-08-26T13:31:32Z] Research round 1/3
[2026-08-26T13:31:32Z] Running researcher.
[2026-08-26T13:31:32Z] API: provider=qwen model=qwen3.8-max attempt=1
[2026-08-26T13:31:33Z] API HTTP status: 401
[2026-08-26T13:31:33Z] WARNING: Client/API error from qwen.
[2026-08-26T13:31:55Z] Started run: 20260826T133155Z
[2026-08-26T13:31:55Z] ==================================================
[2026-08-26T13:31:55Z] Researching: research/05-memory/memory-bandwidth.md
[2026-08-26T13:31:55Z] ==================================================
[2026-08-26T13:31:55Z] Research round 1/3
[2026-08-26T13:31:55Z] Running researcher.
[2026-08-26T13:31:55Z] API: provider=qwen model=qwen3.8-max attempt=1
[2026-08-26T13:39:15Z] Started run: 20260826T133915Z
[2026-08-26T13:39:15Z] ==================================================
[2026-08-26T13:39:15Z] Researching: research/05-memory/memory-bandwidth.md
[2026-08-26T13:39:15Z] ==================================================
[2026-08-26T13:39:15Z] Research round 1/3
[2026-08-26T13:39:15Z] Running researcher.
[2026-08-26T13:39:15Z] API: provider=qwen model=qwen3.8-max attempt=1
[2026-08-26T13:41:06Z] Started run: 20260826T134106Z
[2026-08-26T13:41:06Z] ==================================================
[2026-08-26T13:41:06Z] Researching: research/05-memory/memory-bandwidth.md
[2026-08-26T13:41:06Z] ==================================================
[2026-08-26T13:41:06Z] Research round 1/3
[2026-08-26T13:41:06Z] Running researcher.
[2026-08-26T13:41:06Z] API: provider=qwen model=qwen3.7-plus attempt=1
[2026-08-26T13:43:49Z] Started run: 20260826T134349Z
[2026-08-26T13:43:49Z] ==================================================
[2026-08-26T13:43:49Z] Researching: research/05-memory/memory-bandwidth.md
[2026-08-26T13:43:49Z] ==================================================
[2026-08-26T13:43:49Z] Research round 1/3
[2026-08-26T13:43:49Z] Running researcher.
[2026-08-26T13:43:49Z] API: provider=qwen model=qwen3.7-plus attempt=1
[2026-08-26T13:45:21Z] Response received: 11713 bytes
[2026-08-26T13:45:21Z] Running reviewer.
[2026-08-26T13:45:21Z] API: provider=qwen model=deepseek-v4-pro attempt=1
[2026-08-26T13:47:18Z] Response received: 1142 bytes
[2026-08-26T13:47:18Z] Research passed review.
[2026-08-26T13:47:18Z] Accepted research document: research/05-memory/memory-bandwidth.md
[2026-08-26T13:47:58Z] Started run: 20260826T134758Z
[2026-08-26T13:47:58Z] ==================================================
[2026-08-26T13:47:58Z] Researching: research/05-memory/atomics.md
[2026-08-26T13:47:58Z] ==================================================
[2026-08-26T13:47:58Z] Research round 1/3
[2026-08-26T13:47:58Z] Running researcher.
[2026-08-26T13:47:58Z] API: provider=qwen model=qwen3.7-plus attempt=1
[2026-08-26T13:49:00Z] Started run: 20260826T134900Z
[2026-08-26T13:49:00Z] ==================================================
[2026-08-26T13:49:00Z] Researching: research/05-memory/atomics.md
[2026-08-26T13:49:00Z] ==================================================
[2026-08-26T13:49:00Z] Research round 1/3
[2026-08-26T13:49:00Z] Running researcher.
[2026-08-26T13:49:00Z] API: provider=qwen model=qwen3.7-plus attempt=1
[2026-08-26T13:50:43Z] Response received: 12853 bytes
[2026-08-26T13:50:43Z] Running reviewer.
[2026-08-26T13:50:43Z] API: provider=qwen model=deepseek-v4-pro attempt=1
[2026-08-26T13:52:53Z] Response received: 3245 bytes
[2026-08-26T13:52:53Z] Research rejected by reviewer.
[2026-08-26T13:52:53Z] Preparing revision round 2.
[2026-08-26T13:52:58Z] Research round 2/3
[2026-08-26T13:52:58Z] Running revision agent.
[2026-08-26T13:52:58Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=1
[2026-08-26T13:52:59Z] API HTTP status: 401
[2026-08-26T13:52:59Z] WARNING: Client/API error from openrouter.
[2026-08-26T13:53:02Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=2
[2026-08-26T13:53:02Z] API HTTP status: 401
[2026-08-26T13:53:02Z] WARNING: Client/API error from openrouter.
[2026-08-26T13:53:08Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=3
[2026-08-26T13:53:08Z] API HTTP status: 401
[2026-08-26T13:53:08Z] WARNING: Client/API error from openrouter.
[2026-08-26T13:53:08Z] WARNING: API request exhausted all retries.
[2026-08-26T13:53:08Z] WARNING: Revision API failed.
[2026-08-26T13:53:08Z] WARNING: Research round was NOT counted as a valid revision.
[2026-08-26T13:53:36Z] Started run: 20260826T135336Z
[2026-08-26T13:53:36Z] ==================================================
[2026-08-26T13:53:36Z] Researching: research/05-memory/atomics.md
[2026-08-26T13:53:36Z] ==================================================
[2026-08-26T13:53:36Z] Research round 1/3
[2026-08-26T13:53:36Z] Running researcher.
[2026-08-26T13:53:36Z] API: provider=qwen model=qwen3.7-plus attempt=1
[2026-08-26T13:55:19Z] Response received: 11530 bytes
[2026-08-26T13:55:19Z] Running reviewer.
[2026-08-26T13:55:19Z] API: provider=qwen model=deepseek-v4-pro attempt=1
[2026-08-26T13:57:10Z] Response received: 3825 bytes
[2026-08-26T13:57:10Z] Research rejected by reviewer.
[2026-08-26T13:57:10Z] Preparing revision round 2.
[2026-08-26T13:57:15Z] Research round 2/3
[2026-08-26T13:57:15Z] Running revision agent.
[2026-08-26T13:57:15Z] API: provider=qwen model=qwen3.7-flash-2026-07-15 attempt=1
[2026-08-26T13:58:00Z] Response received: 10054 bytes
[2026-08-26T13:58:00Z] Running reviewer.
[2026-08-26T13:58:00Z] API: provider=qwen model=deepseek-v4-pro attempt=1
[2026-08-26T14:00:15Z] Response received: 3571 bytes
[2026-08-26T14:00:15Z] Research rejected by reviewer.
[2026-08-26T14:00:15Z] Preparing revision round 3.
[2026-08-26T14:00:20Z] Research round 3/3
[2026-08-26T14:00:20Z] Running revision agent.
[2026-08-26T14:00:20Z] API: provider=qwen model=qwen3.7-flash-2026-07-15 attempt=1
[2026-08-26T14:01:09Z] Response received: 9606 bytes
[2026-08-26T14:01:09Z] Running reviewer.
[2026-08-26T14:01:09Z] API: provider=qwen model=deepseek-v4-pro attempt=1
[2026-08-26T14:03:31Z] Response received: 4773 bytes
[2026-08-26T14:03:31Z] Research rejected by reviewer.
[2026-08-26T14:03:31Z] Maximum research rounds reached.
[2026-08-26T14:03:31Z] Leaving original document unchanged.
@@ -0,0 +1 @@
{"error":{"message":"Provider returned error","code":429,"metadata":{"raw":"minimax/minimax-m3:free is temporarily rate-limited upstream. Please retry shortly, or add your own key to accumulate your rate limits: https://openrouter.ai/settings/integrations","provider_name":"GMICloud","is_byok":false,"provider_error_code":"rate_limit_exceeded","limit_source":"upstream_provider_shared_pool","remedy_hint":"Retry shortly, add your own provider key (https://openrouter.ai/settings/integrations), or route to another provider with provider routing: https://openrouter.ai/docs/features/provider-routing","retry_after_seconds":60,"retry_after_seconds_raw":60,"headers":{"Retry-After":"60"}}},"user_id":"user_2tRWhLFlPACV1q6VcCZVeaq8af4"}
@@ -0,0 +1 @@
{"error":{"message":"Provider returned error","code":429,"metadata":{"raw":"minimax/minimax-m3:free is temporarily rate-limited upstream. Please retry shortly, or add your own key to accumulate your rate limits: https://openrouter.ai/settings/integrations","provider_name":"GMICloud","is_byok":false,"provider_error_code":"rate_limit_exceeded","limit_source":"upstream_provider_shared_pool","remedy_hint":"Retry shortly, add your own provider key (https://openrouter.ai/settings/integrations), or route to another provider with provider routing: https://openrouter.ai/docs/features/provider-routing","retry_after_seconds":60,"retry_after_seconds_raw":60,"headers":{"Retry-After":"60"}}},"user_id":"user_2tRWhLFlPACV1q6VcCZVeaq8af4"}
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
@@ -0,0 +1 @@
{"error":{"message":"Provider returned error","code":429,"metadata":{"raw":"minimax/minimax-m3:free is temporarily rate-limited upstream. Please retry shortly, or add your own key to accumulate your rate limits: https://openrouter.ai/settings/integrations","provider_name":"GMICloud","is_byok":false,"provider_error_code":"rate_limit_exceeded","limit_source":"upstream_provider_shared_pool","remedy_hint":"Retry shortly, add your own provider key (https://openrouter.ai/settings/integrations), or route to another provider with provider routing: https://openrouter.ai/docs/features/provider-routing","retry_after_seconds":60,"retry_after_seconds_raw":60,"headers":{"Retry-After":"60"}}},"user_id":"user_2tRWhLFlPACV1q6VcCZVeaq8af4"}
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
@@ -0,0 +1 @@
{"id":"chatcmpl-948","object":"chat.completion","created":1787747932,"model":"gemma4:12b","system_fingerprint":"fp_ollama","choices":[{"index":0,"message":{"role":"assistant","content":"{\n \"verdict\": \"PASS\",\n \"confidence\": \"HIGH\",\n \"issues\": [],\n \"required_fixes\": []\n}","reasoning":"* Input: A technical review of an XH-1 RISC-V CPU research document.\n * Task: Determine if the reviewer intended \"PASS\" or \"FAIL\".\n * Output Format: Strict JSON (no markdown, no explanation, no code fences).\n * JSON Structure:\n * PASS: `{\"verdict\": \"PASS\", \"confidence\": \"HIGH\", \"issues\": [], \"required_fixes\": []}`\n * FAIL: `{\"verdict\": \"FAIL\", \"confidence\": \"HIGH\", \"issues\": [\"...\"], \"required_fixes\": [\"...\"]}`\n\n * \"The document is a well-structured and technically sound research proposal...\" (Positive)\n * \"It correctly identifies the primary challenges...\" (Positive)\n * \"The proposal aligns with the RISC-V ecosystem...\" (Positive)\n * \"demonstrates a sophisticated understanding...\" (Positive)\n * \"The choice of a directory-based MESI protocol... is the industry standard...\" (Positive)\n * \"The recommendation to use a mesh topology is also appropriate...\" (Positive)\n * \"The inclusion of herd7, diy7, and formal tools... indicates a rigorous approach...\" (Positive)\n * \"The document correctly references the relevant parts of the RISC-V ISA manuals...\" (Positive)\n\n * The reviewer uses exclusively positive language.\n * There are no criticisms, \"but\" clauses, or \"needs improvement\" sections.\n * The verdict is clearly a \"PASS\".\n\n * Verdict: PASS\n * Confidence: HIGH\n * Issues: []\n * Required Fixes: []\n\n * No Markdown? Yes.\n * No Explanation? Yes.\n * No Code Fences? Yes.\n * Valid JSON? Yes."},"finish_reason":"stop"}],"usage":{"prompt_tokens":635,"completion_tokens":445,"total_tokens":1080}}
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
@@ -0,0 +1 @@
{"id":"chatcmpl-997","object":"chat.completion","created":1787748319,"model":"gemma4:12b","system_fingerprint":"fp_ollama","choices":[{"index":0,"message":{"role":"assistant","content":"{\n \"verdict\": \"PASS\",\n \"confidence\": \"HIGH\",\n \"issues\": [],\n \"required_fixes\": []\n}","reasoning":"* Input: A technical review of an XH-1 RISC-V CPU research document.\n * Task: Determine if the reviewer intended \"PASS\" or \"FAIL\".\n * Output Format: Strict JSON (no markdown, no explanation, no code fences).\n * JSON Structure:\n * PASS: `{\"verdict\": \"PASS\", \"confidence\": \"HIGH\", \"issues\": [], \"required_fixes\": []}`\n * FAIL: `{\"verdict\": \"FAIL\", \"confidence\": \"HIGH\", \"issues\": [\"...\"], \"required_fixes\": [\"...\"]}`\n\n * \"The document is technically sound and well-structured...\" (Positive)\n * \"The document correctly identifies...\" (Positive)\n * \"This demonstrates a high level of adherence...\" (Positive)\n * \"The choice of... is the industry standard.\" (Positive)\n * \"The inclusion of... is essential...\" (Positive)\n * \"The 'Unresolved' section shows a mature understanding...\" (Positive)\n * \"This is a valid concern...\" (Positive)\n * There are no \"Critical\" or \"Major\" issues mentioned. The \"Unresolved\" section is praised for showing maturity, not as a failure.\n\n * The reviewer is very positive.\n * Verdict: PASS.\n * Confidence: HIGH.\n * Issues: [].\n * Required Fixes: [].\n\n * No markdown.\n * No code fences.\n * No explanation.\n\n ```json\n {\n \"verdict\": \"PASS\",\n \"confidence\": \"HIGH\",\n \"issues\": [],\n \"required_fixes\": []\n }\n ```"},"finish_reason":"stop"}],"usage":{"prompt_tokens":730,"completion_tokens":427,"total_tokens":1157}}
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
@@ -0,0 +1 @@
{"error":{"message":"Provider returned error","code":429,"metadata":{"raw":"minimax/minimax-m3:free is temporarily rate-limited upstream. Please retry shortly, or add your own key to accumulate your rate limits: https://openrouter.ai/settings/integrations","provider_name":"GMICloud","is_byok":false,"provider_error_code":"rate_limit_exceeded","limit_source":"upstream_provider_shared_pool","remedy_hint":"Retry shortly, add your own provider key (https://openrouter.ai/settings/integrations), or route to another provider with provider routing: https://openrouter.ai/docs/features/provider-routing","retry_after_seconds":60,"retry_after_seconds_raw":60,"headers":{"Retry-After":"60"}}},"user_id":"user_2tRWhLFlPACV1q6VcCZVeaq8af4"}
File diff suppressed because one or more lines are too long
@@ -0,0 +1 @@
{"id":"chatcmpl-481","object":"chat.completion","created":1787749010,"model":"gemma4:12b","system_fingerprint":"fp_ollama","choices":[{"index":0,"message":{"role":"assistant","content":"The document provided is a high-quality technical research overview of the XH-1 memory hierarchy. It correctly identifies the primary challenges associated with a 128-core system, specifically regarding directory-based coherence, mesh interconnect latency, and the complexities of RISC-V memory model compliance.\n\n### Technical Review:\n\n1. **RISC-V Compliance:** The document correctly identifies the critical components for RISC-V compliance: RVWMO, the \"A\" extension (atomic instructions), and the Zicbom/Zicboz extensions for software-managed cache coherence.\n2. **Coherence \u0026 Scalability:** The choice of a directory-based protocol over a snoop-based one is the correct architectural decision for a 128-core system. The mention of \"limited-pointer\" overflows and the \"broadcast fallback\" shows a mature understanding of directory scaling.\n3. **Microarchitecture:** The inclusion of \"critical-word first\" for high-latency mesh misses and the consideration of \"NINE\" (Non-Inclusive, Non-Exclusive) policies for the LLC are standard, high-performance industry practices.\n4. **Verification:** The document correctly identifies that a 128-core directory protocol is a high-risk verification target, correctly suggesting formal verification as a necessary component of the design cycle.\n5. **Nuance:** The distinction between a \"research document\" and a \"specification\" is vital. The document correctly frames the current proposals as \"workload-dependent\" rather than fixed requirements, which is the correct posture for a project at this stage.\n\nThe document is internally consistent, well-referenced, and addresses the specific complexities of large-scale multicore systems.\n\n```json\n{\n \"verdict\": \"PASS\",\n \"confidence\": \"HIGH\",\n \"issues\": [],\n \"required_fixes\": []\n}\n```"},"finish_reason":"stop"}],"usage":{"prompt_tokens":2051,"completion_tokens":404,"total_tokens":2455}}
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
@@ -0,0 +1 @@
{"id":"chatcmpl-921","object":"chat.completion","created":1787749374,"model":"gemma4:12b","system_fingerprint":"fp_ollama","choices":[{"index":0,"message":{"role":"assistant","content":"The document describes a high-performance, many-core RISC-V system (XH-1) with 128 cores. The technical specifications provided are consistent with modern high-performance computing (HPC) architecture and RISC-V standards.\n\n### Technical Review:\n\n1. **Scalability \u0026 Interconnect:** The choice of a **2D mesh** for 128 cores is a standard and correct architectural decision. At this scale, a ring or crossbar is physically and logically impractical. The use of **virtual channels** to separate request, response, and data traffic is critical to prevent head-of-line blocking and protocol deadlocks in a directory-based system.\n2. **Cache Coherence:** The transition from snooping to a **distributed directory MESI protocol** is mandatory for 128 cores. The inclusion of **Zicbom and Zicboz** is a critical requirement for RISC-V Linux compatibility, as these allow the OS to manage cache lines and memory attributes correctly.\n3. **Memory Consistency:** The proposal to use a **litmus test suite** for RISC-V memory consistency is the industry standard for verifying that the hardware correctly implements the Weak Memory Model (WMM) required by the RISC-V ISA.\n4. **Hierarchy:** The L2 being inclusive of L1 and the L3 being the coherence point is a standard design that simplifies the coherence logic at the L3 level.\n5. **OS/Software:** Exposing the topology as **NUMA nodes** is the correct way to handle a 128-core system in Linux, ensuring the scheduler can optimize for local memory access.\n6. **Verification:** The multi-pronged approach (Formal for state machines, Simulation for consistency, and Stress/Synthetic for performance) is a robust verification strategy for a complex SoC.\n\n### Minor Observations:\n- The decision to defer Zicbop is prudent, as it is a more complex extension for \"buffer\" management that may not be necessary for initial silicon.\n- The distinction between \"inclusive\" and \"non-inclusive\" L3 is correctly identified as a key design choice for the implementation phase.\n\nThe document is technically sound, follows RISC-V standards, and addresses the specific challenges of scaling to 128 cores.\n\n```json\n{\n \"verdict\": \"PASS\",\n \"confidence\": \"HIGH\",\n \"issues\": [],\n \"required_fixes\": []\n}\n```"},"finish_reason":"stop"}],"usage":{"prompt_tokens":2051,"completion_tokens":532,"total_tokens":2583}}
File diff suppressed because one or more lines are too long
@@ -0,0 +1 @@
{"error":{"message":"No cookie auth credentials found","code":401}}
@@ -0,0 +1 @@
{"error":{"message":"No cookie auth credentials found","code":401}}
@@ -0,0 +1 @@
{"error":{"message":"No cookie auth credentials found","code":401}}
@@ -0,0 +1 @@
{"error":{"message":"No cookie auth credentials found","code":401}}
@@ -0,0 +1 @@
{"error":{"message":"You didn't provide an API key. You need to provide your API key in an Authorization header using Bearer auth (i.e. Authorization: Bearer YOUR_KEY). ","type":"invalid_request_error","param":null,"code":null},"request_id":"5b968769-a7b3-9d60-8195-7de88e17e1a3"}
@@ -0,0 +1 @@
{"error":{"message":"You didn't provide an API key. You need to provide your API key in an Authorization header using Bearer auth (i.e. Authorization: Bearer YOUR_KEY). ","type":"invalid_request_error","param":null,"code":null},"request_id":"66fd0f2b-1ac5-9b0c-9966-75b3d773b23b"}
@@ -0,0 +1 @@
{"error":{"message":"You didn't provide an API key. You need to provide your API key in an Authorization header using Bearer auth (i.e. Authorization: Bearer YOUR_KEY). ","type":"invalid_request_error","param":null,"code":null},"request_id":"3a23149c-14e0-98ba-9207-bb436ec6d749"}
@@ -0,0 +1 @@
{"error":{"message":"You didn't provide an API key. You need to provide your API key in an Authorization header using Bearer auth (i.e. Authorization: Bearer YOUR_KEY). ","type":"invalid_request_error","param":null,"code":null},"request_id":"48d53b24-2073-9696-8bc0-651a6215025c"}
@@ -0,0 +1 @@
{"error":{"message":"You didn't provide an API key. You need to provide your API key in an Authorization header using Bearer auth (i.e. Authorization: Bearer YOUR_KEY). ","type":"invalid_request_error","param":null,"code":null},"request_id":"3071b8b5-52b6-93b0-bb3a-13392a8ccf4e"}
@@ -0,0 +1 @@
{"error":{"message":"You didn't provide an API key. You need to provide your API key in an Authorization header using Bearer auth (i.e. Authorization: Bearer YOUR_KEY). ","type":"invalid_request_error","param":null,"code":null},"request_id":"f167a86f-f816-9ebf-96d1-5f1d6ba405a6"}
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
@@ -0,0 +1 @@
{"error":{"message":"No cookie auth credentials found","code":401}}
@@ -0,0 +1 @@
{"error":{"message":"No cookie auth credentials found","code":401}}
@@ -0,0 +1 @@
{"error":{"message":"No cookie auth credentials found","code":401}}
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
@@ -0,0 +1,6 @@
2026-08-26T12:12:47Z research/05-memory/memory-ordering.md 1 research completed
2026-08-26T12:15:45Z research/05-memory/memory-ordering.md 1 review
2026-08-26T12:16:25Z research/05-memory/memory-ordering.md 2 revision completed
2026-08-26T12:19:32Z research/05-memory/memory-ordering.md 2 review
2026-08-26T12:20:25Z research/05-memory/memory-ordering.md 3 revision completed
2026-08-26T12:23:32Z research/05-memory/memory-ordering.md 3 review
@@ -0,0 +1 @@
research/05-memory/memory-ordering.md
@@ -0,0 +1,333 @@
# Memory Ordering
## Status
Stub document. XH-1 memory ordering model has not yet been defined. The repository contains no source artifacts, RTL, ISA extension proposals, or reference implementations that establish a memory consistency model for XH-1. All quantitative or XH-1-specific claims in this document are flagged INSUFFICIENT EVIDENCE where applicable.
## Abstract
Memory ordering defines the rules by which loads and stores issued by one or more cores become visible to other cores in the system. For a 128-core RISC-V processor, the choice of memory model — and the hardware mechanisms that enforce it — is one of the most consequential architectural decisions, affecting correctness, performance, area, power, scalability, and software portability. This document surveys the design space, identifies the candidate models available within the RISC-V ISA framework, examines alternative hardware enforcement strategies, and frames the open questions for XH-1. No final recommendation is made because the repository does not establish pipeline style, coherency protocol, or interconnect topology, all of which constrain the memory ordering decision. A conditional tentative default is stated, contingent on the assumed constraints being correct.
## Research Question
What memory consistency model and hardware ordering mechanism should XH-1 adopt, given:
1. A 128-core replication target
2. The RISC-V ISA baseline
3. The need to balance software portability against hardware implementation cost
4. Unknown microarchitecture style (in-order vs. out-of-order is not established in the repository)
5. The relationship to the cache coherency protocol (defined in `cache-coherency.md`, currently SOON)
## Background
### Memory Consistency Models — Terminology
PROPOSAL — Use the following working taxonomy in this document:
- **Sequential Consistency (SC)**: All cores observe a single total order of memory operations that respects each core's program order.
- **Total Store Order (TSO)**: Loads may bypass stores from the same core; stores from one core are seen in order by all cores. Used by x86.
- **Relaxed / Weak Memory Models**: Many orderings are allowed; fences (barriers) are required to enforce constraints. Used by ARM, RISC-V, and PowerPC.
- **Release Consistency (RC) / RCsc / RCpc / RCO**: A family of models with different fence semantics. Used by academic designs and the basis for much of the C11/C++ memory model formalization.
ASSUMPTION — This document uses "RC variants" as an umbrella term for RCsc, RCpc, and RCO. The specific semantics of each variant are not redefined here; readers are referred to the literature.
### The RISC-V "RVWMO" Model
FACT — The RISC-V ISA specification defines the **RISC-V Weak Memory Ordering (RVWMO)** model as the baseline. Under RVWMO:
- Loads and stores to the same address from the same hart are ordered.
- A store followed by a load to a *different* address may be reordered.
- The `FENCE` instruction is used to enforce orderings. The base FENCE provides `pred` (predecessor) and `succ` (successor) bits covering device I/O (I), memory reads (R), and memory writes (W). Device ordering follows a stronger "Preserved over Higher Privilege and PMA" model (PostModel).
- The optional `Ztso` extension provides **RVTSO**, which gives x86-like TSO semantics with a relaxed form of the FENCE.
FACT — RVWMO and the `Ztso` extension are documented in the ratified RISC-V Unprivileged ISA specification. The repository does not state which revision of the specification is the reference for XH-1; this is an open question.
ASSUMPTION — RVWMO + the Ztso optional extension is the available ISA-level model space that XH-1 must select from. (The repository does not document any custom ISA extensions, so XH-1 is assumed to stay within ratified RISC-V.)
### Why Memory Ordering is Hard at 128 Cores
OPEN QUESTION — The repository does not establish:
- Whether XH-1 is a single die, multi-die package, or chiplet-based system
- The interconnect topology (mesh, ring, crossbar, NoC, etc.)
- Whether coherence is directory-based, broadcast-based, or hybrid
- The private-vs-shared L2/L3 split
- The memory-side directory or snoop filter organization
These decisions have direct bearing on which ordering mechanisms are feasible. Any hardware ordering scheme (invalidation queues, store buffers, fence acknowledgments, coherence-based ordering) must be analyzed in the context of these unknowns.
## Existing Approaches
### Approach A — RISC-V RVWMO (Baseline)
The standard weak memory model provided by the RISC-V ISA. Requires implementation of `FENCE`, `FENCE.I`, and the I/O ordering semantics.
**Mechanism**: Architecture permits aggressive reordering. Hardware implementations commonly use:
- Store buffers with forwarding
- Load-load reordering queues
- Coherence-based ordering: stores become visible at coherence invalidation/upgrade points
- Fence implementation via global acknowledgment counters or by draining reorder queues
INSUFFICIENT EVIDENCE — Which of these mechanisms XH-1 will use is not established by the repository.
### Approach B — RVWMO + Ztso Extension
Adds RVTSO. Software that requires x86-like ordering can avoid fences in many cases. Still RISC-V compliant.
**Mechanism**: Stores from one hart are serialized in the interconnect; loads may still pass earlier stores to different addresses.
### Approach C — Sequential Consistency (SC)
All operations observed in a single global order respecting program order. Simplest software model; most expensive hardware in general.
**Mechanism**: Commonly implemented via store-buffer draining on loads, broadcast-based coherence, and conservative replay mechanisms. The qualitative characterization of "very costly" is a general literature claim, not an XH-1 measurement; whether SC is specifically costly at 128 cores is workload- and topology-dependent and is treated here as a hypothesis, not a known result.
### Approach D — Release Consistency variants (RCsc, RCpc, RCO)
Academic models where release/acquire ordering is cheap but stronger fences are required for full memory ordering. Basis of C11/C++ model.
**Mechanism**: Only release/acquire fences require acknowledgment; acquire-only loads are cheap. Alignment with C11 is a strong software benefit. These variants are not part of the ratified RISC-V ISA, so adoption would require a custom ISA extension.
## Alternative Designs
### Hardware Ordering Mechanisms
PROPOSAL — The following mechanisms should be considered for fence/ordering enforcement once coherency and interconnect are known:
1. **Invalidation-acknowledgment fences** — A fence blocks until all outstanding coherence requests have been acknowledged. Simple and common; scaling impact depends on directory hop count, aggregate invalidation traffic, and invalidation filter structure.
2. **Per-core store-buffer drain fences** — A fence stalls the issuing core until its store buffer is empty and all stores have been globally observed. Predictable latency; serializes through the coherence fabric.
3. **Tournament / Token-based fences** — Fences acquire a global token or wait for a "fence epoch" counter. Avoids worst-case drain, but adds global state.
4. **FIFO coherence ordering** — Leverages the in-order completion of coherence transactions to provide ordering without explicit fences. Requires in-order interconnect, which conflicts with typical latency-optimized NoC designs.
5. **Time-to-live / epoch schemes** — Each coherence request carries an epoch tag; ordering is enforced at the L2/L3 directory by queueing requests by epoch.
6. **Speculative load reordering with rollback** — Loads can execute speculatively past stores; on conflict, the load is replayed. Common in high-performance out-of-order cores (e.g., reported in the IBM POWER literature).
ASSUMPTION — XH-1's microarchitecture style is not yet established. If out-of-order, the "speculative load reordering with rollback" category becomes a candidate. If in-order, mechanism choices narrow considerably.
## Comparison
| Property | SC | RVWMO | RVWMO + Ztso | RC variants |
|----------|----|----|----|----|
| Software ease | Highest | Medium | Medium-High | Low-Medium |
| Hardware cost (general) | Highest | Low | Low-Medium | Low |
| Performance headroom (general) | Lowest | Highest | High | High |
| Fence latency (general) | N/A | Variable | Variable | Variable |
| Verification cost (general) | Lowest in principle | High | High | High |
| C11/C++ alignment | Approximate | Approximate | Weak | Direct |
| 128-core scalability | Poor (hypothesis) | Unmeasured | Unmeasured | Unmeasured |
INSUFFICIENT EVIDENCE — Quantitative fence-latency, coherence-traffic, and scalability numbers for XH-1 are unavailable. Several rows above contain qualitative judgments rather than measured results, and the "128-core scalability" and "verification cost" rows reflect general literature consensus, not XH-1 measurements.
## Advantages
### RVWMO (Baseline)
- **Lowest hardware cost** of all candidates for the base case, as a general property of the model — most orderings are simply not enforced.
- **Maximum performance** potential, as a general property — reordering is unconstrained.
- **Familiar** to the RISC-V software ecosystem (Linux, RISC-V GCC, LLVM), subject to the open question of whether Linux is the target OS.
- **Standardized** — concrete compliance test suites exist externally.
### RVWMO + Ztso
- **x86-like TSO** for software that benefits from stronger ordering.
- **Backwards compatible** — software not using Ztso still runs correctly under RVWMO.
- **Modest hardware cost** — primarily requires store-store ordering at the coherence layer (qualitative claim, not measured for XH-1).
### Release Consistency Variants
- **Excellent fit for C11/C++** atomics — acquire/release are the most-used fence types in real code, per the C/C++ standards.
- **Cheapest common case** — acquire and release can be implemented with lightweight mechanisms, as a general property.
- **Hardware cost is concentrated on the rare case** of full seq_cst fences, as a general property.
## Disadvantages
### RVWMO (Baseline)
- **Fence latency depends on coherence round-trip** — at 128 cores, fence latency is expected to be high because global acknowledgment must traverse the coherence fabric. Specific cycle counts: INSUFFICIENT EVIDENCE.
- **Verification complexity is high** — many legal reorderings; the litmus-test failure surface is large. This is a general property of weak models, not an XH-1 measurement.
- **Software burden** — kernel and runtime code must insert fences correctly.
### RVWMO + Ztso
- **Two consistency models in one chip** — adds documentation, validation, and software education cost.
- **In-flight mixing** of ordering regimes in the same software is hard to reason about.
### Release Consistency Variants
- **Non-standard for RISC-V** — would require a custom ISA extension. Cannot be recommended without a strong software-side driver.
- **Verification cost is high** in general because of subtle fence semantics; whether it is the "highest" among candidates is a claim about the literature, not a measured XH-1 result.
### Sequential Consistency
- **Performance cost is severe in the general case** — every load may stall on store-buffer drain; whether this is "severe at 128 cores" specifically is workload- and topology-dependent and is not established by the repository.
- **No selective escape valve** — software cannot opt out of the strongest model.
- **Does not match RISC-V ecosystem expectations** — surprising to RISC-V software developers, as a general observation.
## XH-1 Considerations
OPEN QUESTION — The following XH-1-internal questions are unresolved by the repository:
1. Is XH-1's pipeline in-order or out-of-order?
2. What is the coherence protocol (MOESI, MESI, directory-based, broadcast)?
3. Is the interconnect a NoC, ring, or crossbar? Does it preserve in-order delivery of coherence responses?
4. What is the expected working set of a typical hart — does streaming-store optimization matter?
5. Does XH-1 target HPC, server, embedded, or mixed workloads? (Workload affects fence frequency.)
6. Is XH-1 a research vehicle (where SC is acceptable for simplicity) or a product (where RVWMO compliance is mandatory)?
7. Does the memory map use non-coherent regions (e.g., DMA, I/O) that interact with the ordering model?
ASSUMPTION — If XH-1 targets RISC-V ecosystem compatibility (Linux, RISC-V GCC, RISC-V LLVM), then RVWMO is mandatory. Ztso may optionally be added for x86-software porting convenience.
## 128-Core Scaling Considerations
INSUFFICIENT EVIDENCE — Quantitative figures for XH-1 at 128 cores are unavailable. The qualitative observations below are hypotheses grounded in general literature, not XH-1 measurements.
### Fence Latency
- **Fence global acknowledgment** latency is bounded below by the worst-case coherence round-trip time. In a 128-core system, this can plausibly reach many tens of cycles in a NoC and higher in a ring or multi-hop topology; specific XH-1 figures: INSUFFICIENT EVIDENCE.
- **Store-buffer drain** latency is approximately proportional to the store buffer depth and the time to invalidate all sharers. Specific XH-1 figures: INSUFFICIENT EVIDENCE.
- **Token / epoch fences** may scale better because they avoid draining the entire coherence fabric — they synchronize at a logical epoch boundary. This is a general property, not an XH-1 measurement.
### Coherence Traffic
- **RVWMO fences** generate invalidation or acknowledgment traffic that competes with normal coherence traffic. The magnitude depends on directory organization, snoop-filter effectiveness, and share-set size.
- **SC at 128 cores** can plausibly saturate a directory-based interconnect because every load's coherence transaction may stall pending store ordering; this is a hypothesis, not a measured XH-1 result.
- **RC variants** concentrate traffic on the fence path only, as a general property.
### Verification Scalability
- **State-space explosion**: at 128 cores, the number of concurrently observable memory operations grows combinatorially in the unmitigated model. Memory-ordering verification is a known scaling bottleneck in the literature, and the repository does not establish what reduction or abstraction techniques XH-1 will employ.
- **RC variants** are commonly cited as among the hardest to verify because of subtle fence interactions; whether they are hardest in the absolute sense is a literature claim, not an XH-1 measurement.
- **SC is easiest to verify in principle** because it admits fewer legal behaviors, but worst in performance.
INSUFFICIENT EVIDENCE — No quantitative fence-latency, coherence-traffic, or verification-scaling data is available for XH-1. Estimates cannot be made without knowing the interconnect, coherence protocol, and verification methodology.
## Performance Considerations
PROPOSAL — Performance-impact dimensions to evaluate once workload and microarchitecture are defined:
1. **Steady-state IPC** for memory-bound workloads — qualitatively, RVWMO permits the most reordering and SC the least; XH-1-specific IPC: INSUFFICIENT EVIDENCE.
2. **Fence latency** — measured as cycles from fence issue to fence completion. Dominant for synchronization-heavy code. XH-1-specific: INSUFFICIENT EVIDENCE.
3. **Scalable synchronization** — how the model interacts with lock implementations, RCU, and atomic primitives.
4. **Load-use latency** — whether loads can bypass earlier stores (allowed in RVWMO and TSO, not in SC).
5. **Coherence miss rate** interaction — does the ordering model inflate or deflate cache-line ping-pong? Model-dependent; XH-1-specific: INSUFFICIENT EVIDENCE.
ASSUMPTION — For HPC workloads with high-fence frequency (e.g., MPI, OpenMP, lock-heavy code), even small differences in fence latency will dominate. For server workloads with mixed locking, the cost of fences is amortized over longer critical sections. No XH-1 workload profile is documented in the repository.
## Area Considerations
INSUFFICIENT EVIDENCE — All area descriptions in this section are qualitative. Precise XH-1 area figures (in gates, µm², or mm²) are unavailable; no synthesis, layout, or RTL data exists in the repository.
PROPOSAL — Qualitative area cost of ordering mechanisms, as a function of microarchitectural choices:
| Mechanism | Approximate area cost per core (qualitative) |
|-----------|----------------------------------------------|
| Store buffer with forwarding | Small (few entries × cache-line width) |
| Load-load reordering queue | Small to medium (grows with IPC and reorder depth) |
| Fence-acknowledgment counter | Negligible (per-core counter) |
| Global epoch / token state | Medium (one per chip, not per core) |
| Speculative load reorder + rollback | Large (load-load queue, replay logic) |
PROPOSAL — At 128 cores, per-core area costs are replicated 128 times. Global state (epoch counters, fence acknowledgers) is O(1) per chip and therefore amortizes favorably on a per-core basis, but the critical-path and clock-distribution impact of global state is not established by the repository.
## Power and Energy Considerations
INSUFFICIENT EVIDENCE — No per-fence, per-core, or aggregate energy figures for XH-1 are available. The earlier version of this document included a "100 pJ-per-core fence" figure that was not substantiated by any repository artifact or external citation; that figure has been removed.
PROPOSAL — Qualitative energy implications, as general properties of the model classes:
- **Fences** generate global coherence traffic — energy cost scales with the number of cores that must acknowledge. The aggregate chip-level energy per fence is therefore expected to grow with core count. Quantitative XH-1 values: INSUFFICIENT EVIDENCE.
- **Speculative load reordering** wastes energy on rollback when a violation is detected, as a general property.
- **Strong ordering models** keep more coherence state in flight (more invalidations, more retries), as a general property.
- **RVWMO** with lightweight fences (when software permits) is the lowest-energy regime in the general case.
## Implementation Considerations
PROPOSAL — Implementation order of dependencies (the repository does not establish these, but the research area suggests the following):
1. Define coherency protocol → `cache-coherency.md`
2. Define interconnect topology → not yet in repository
3. Define memory hierarchy → `memory-hierarchy.md`
4. Define memory ordering model → this document
5. Define atomic primitives → `atomics.md`
ASSUMPTION — The Ztso extension is a small RTL delta over RVWMO. The dominant implementation cost is verification, not area. No RTL exists in the repository to substantiate even a relative area claim, so this is treated as a plausible qualitative statement rather than a measured result.
OPEN QUESTION — Does XH-1 have a coherent accelerator fabric or non-coherent IO (e.g., CXL, DMA engines)? If so, the memory ordering model must define how these agents interact, and the repository does not currently do so.
## Verification Considerations
PROPOSAL — Verification challenges:
- **Litmus tests**: A standard RISC-V litmus test suite is available externally; whether XH-1 will adopt it as-is or define a custom suite is OPEN.
- **State-space explosion**: 128 cores × N outstanding operations produces a combinatorial unmitigated state space. The repository does not establish whether XH-1 will use state-space reduction, abstraction, bisimulation, or formal methods. Any specific XH-1 claim: INSUFFICIENT EVIDENCE.
- **Fence semantics**: Each fence variant must be exhaustively tested at corner cases (interrupts, exceptions, MMIO).
- **Coherence-ordering interaction**: The ordering model must be checked against the chosen coherence protocol for consistency. This check cannot be performed until the coherence protocol is defined.
INSUFFICIENT EVIDENCE — No XH-1 verification infrastructure, formal spec, or litmus-test set is documented in the repository.
ASSUMPTION — RVWMO is the most documented and tested model for RISC-V, per the ratified specification and ecosystem practice. Ztso is a small addition. SC and RC variants would require custom formal infrastructure. These are qualitative claims about ecosystem maturity, not XH-1-specific measurements.
## Software Considerations
PROPOSAL — Software-side impact, as a general property of each model:
- **Linux kernel**: Designed for weak memory models with appropriate fences. Compatible with RVWMO.
- **C11/C++ atomics**: Map cleanly onto release/acquire semantics. Compatible with RVWMO and with RC variants.
- **x86 software ports**: Benefit from Ztso for fewer fences.
- **OpenSHMEM, MPI shmem**: Often assume TSO-like ordering; Ztso helps.
- **HPC codes with hand-rolled atomics**: Highly sensitive to fence latency.
ASSUMPTION — XH-1 will run RISC-V Linux. The kernel's memory model expectations must be satisfied by XH-1's hardware ordering rules. The repository does not document whether Linux is in fact the target.
## Recommendation
INSUFFICIENT EVIDENCE to make a final recommendation.
The repository does not yet establish:
- Pipeline microarchitecture
- Coherence protocol
- Interconnect topology
- Workload target
A recommendation can be made only after the following documents in the memory research area are filled in: `memory-architecture.md`, `cache-coherency.md`, `memory-hierarchy.md`. Without these, any model selection is premature.
PROPOSAL — Tentative default, contingent on the stated assumptions holding and explicitly subject to revision once repository context is available:
1. **Baseline**: RVWMO (assumed mandatory for RISC-V ecosystem compatibility, if Linux is the target OS).
2. **Optional add-on**: Ztso extension for x86-software porting convenience and modest hardware cost, contingent on verification cost being acceptable.
3. **Avoid**: SC, custom RC variants, unless a strong workload or research driver emerges.
4. **Defer to**: `cache-coherency.md` for the coherence-side ordering mechanism.
ASSUMPTION — The tentative default above depends on (a) Linux being the target OS, (b) ecosystem compatibility being prioritized over x86-port convenience, and (c) the coherence and interconnect unknowns being resolvable without invalidating this default. None of these are established by the repository.
## Confidence
| Topic | Confidence | Reason |
|-------|------------|--------|
| RVWMO is the RISC-V baseline | High | Ratified ISA specification |
| Ztso is a small optional extension | High | Ratified extension in the ISA specification |
| SC has high implementation cost in general | High | General literature consensus |
| SC specifically impractical at 128 cores | Low | General claim, not XH-1-measured; workload- and topology-dependent |
| Per-fence energy scales with acknowledging-core count | Medium | Logical consequence; no measurement |
| Quantitative per-fence energy at 128 cores | None | INSUFFICIENT EVIDENCE |
| Verification is hardest for RC variants | LowMedium | Literature claim, not XH-1-measured |
| XH-1 should adopt RVWMO | Low | Repository context not yet sufficient |
| XH-1 should adopt RVWMO + Ztso | Low | Same |
| Ztso is a small RTL delta over RVWMO | LowMedium | Plausible qualitative claim; no RTL to substantiate |
## Open Questions
1. What is XH-1's microarchitecture style (in-order vs. out-of-order)?
2. What coherence protocol and directory organization will XH-1 use?
3. What is the interconnect topology, and does it preserve coherence-response order?
4. What workloads are targeted — HPC, server, embedded?
5. Does XH-1 include non-coherent agents (DMA, CXL, accelerators)?
6. Is XH-1 a research vehicle, a prototype, or a product?
7. What is the expected fence frequency in target software?
8. What is the verification methodology — formal, simulation, or hybrid?
9. Will XH-1 use the standard RISC-V litmus-test suite or define a custom one?
10. Does XH-1 require compatibility with x86-software-only ports? (Drives Ztso inclusion.)
11. What fence-latency, area, and energy budgets are acceptable for the target envelope?
12. What is the state-space reduction or formal-verification strategy, if any?
13. Which revision of the RISC-V Unprivileged ISA specification is the reference for XH-1?
## Sources
INSUFFICIENT EVIDENCE — The XH-1 Research repository does not yet contain source material specific to memory ordering. No citations, papers, or measurements can be honestly attributed to XH-1.
External ISA-level facts cited above (RVWMO, Ztso, FENCE instruction semantics) are drawn from the publicly ratified RISC-V Unprivileged ISA specification, which is the source of truth for the ISA-level memory model. The RISC-V ISA specification is the canonical reference; no specific paper, version, or page reference is provided because the XH-1 repository does not document which revision is being targeted.
INSUFFICIENT EVIDENCE — All quantitative claims about fence latency, area, power, energy, and verification cost at the XH-1 level have been removed or flagged. No benchmarks, no measurements, no RTL data, no simulation results are available in the repository at this time.
File diff suppressed because one or more lines are too long
@@ -0,0 +1,319 @@
# Memory Ordering
## Status
Stub document. XH-1 memory ordering model has not yet been defined. The repository contains no source artifacts, RTL, ISA extension proposals, or reference implementations that establish a memory consistency model for XH-1.
## Abstract
Memory ordering defines the rules by which loads and stores issued by one or more cores become visible to other cores in the system. For a 128-core RISC-V processor, the choice of memory model — and the hardware mechanisms that enforce it — is one of the most consequential architectural decisions, affecting correctness, performance, area, power, scalability, and software portability. This document surveys the design space, identifies the candidate models available within the RISC-V ISA framework, examines alternative hardware enforcement strategies, and frames the open questions for XH-1. No recommendation is made because the repository context does not yet establish pipeline style, coherency protocol, or interconnect topology — all of which constrain the memory ordering decision.
## Research Question
What memory consistency model and hardware ordering mechanism should XH-1 adopt, given:
1. A 128-core replication target
2. The RISC-V ISA baseline
3. The need to balance software portability against hardware implementation cost
4. Unknown microarchitecture style (in-order vs. out-of-order is not established in the repository)
5. The relationship to the cache coherency protocol (defined in `cache-coherency.md`, currently SOON)
## Background
### Memory Consistency Models — Terminology
PROPOSAL — Use the following working taxonomy in this document:
- **Sequential Consistency (SC)**: All cores observe a single total order of memory operations that respects each core's program order.
- **Total Store Order (TSO)**: Loads may bypass stores; stores from one core are seen in order by all cores. Used by x86.
- **Relaxed / Weak Memory Models**: Many orderings are allowed; fences (barriers) are required to enforce constraints. Used by ARM, RISC-V, and PowerPC.
- **Release Consistency (RC) / RCO / RCpc / RCsc**: A family of models with different fence semantics. Used by academic designs and the basis for much of C11/C++ memory model formalization.
### The RISC-V "RVWMO" Model
FACT — The RISC-V ISA specification defines the **RISC-V Weak Memory Ordering (RVWMO)** model as the baseline. Under RVWMO:
- Loads and stores to the same address from the same hart are ordered.
- A store followed by a load to a *different* address may be reordered.
- The `FENCE` instruction is used to enforce orderings. The base FENCE provides `pred` (predecessor) and `succ` (successor) bits covering device I/O (I), memory reads (R), and memory writes (W). Device ordering follows a stronger PostModel.
- The optional `Ztso` extension provides **RVTSO**, which gives x86-like TSO semantics with a relaxed form of the FENCE.
ASSUMPTION — RVWMO + the Ztso optional extension is the available ISA-level model space that XH-1 must select from. (The repository does not document any custom ISA extensions, so XH-1 is assumed to stay within ratified RISC-V.)
### Why Memory Ordering is Hard at 128 Cores
OPEN QUESTION — The repository does not establish:
- Whether XH-1 is a single die, multi-die package, or chiplet-based system
- The interconnect topology (mesh, ring, crossbar, NoC, etc.)
- Whether coherence is directory-based, broadcast-based, or hybrid
- The private-vs-shared L2/L3 split
- The memory-side directory or snoop filter organization
These decisions have direct bearing on which ordering mechanisms are feasible. Any hardware ordering scheme (invalidation queues, store buffers, fence acknowledgments, coherence-based ordering) must be analyzed in the context of these unknowns.
## Existing Approaches
### Approach A — RISC-V RVWMO (Baseline)
The standard weak memory model provided by the RISC-V ISA. Requires implementation of `FENCE`, `FENCE.I`, and the I/O ordering semantics.
**Mechanism**: Architecture permits aggressive reordering. Hardware typically uses:
- Store buffers with forwarding
- Load-load reordering queues
- Coherence-based ordering: stores become visible at coherence invalidation/upgrade points
- Fence implementation via global acknowledgment counters or by draining reorder queues
### Approach B — RVWMO + Ztso Extension
Adds RVTSO. Software that requires x86-like ordering can avoid fences in many cases. Still RISC-V compliant.
**Mechanism**: Stores from one hart are serialized in the interconnect; loads may still pass earlier stores to different addresses.
### Approach C — Sequential Consistency (SC)
All operations observed in a single global order respecting program order. Simplest software model; most expensive hardware.
**Mechanism**: Typically requires store-buffer draining on loads, broadcast-based coherence, and conservative replay mechanisms. Very costly for 128 cores.
### Approach D — Release Consistency variants (RCsc, RCpc, RCO)
Academic models where release/acquire ordering is cheap but stronger fences are required for full memory ordering. Basis of C11/C++ model.
**Mechanism**: Only release/acquire fences require acknowledgment; acquire-only loads are cheap. Alignment with C11 is a strong software benefit.
## Alternative Designs
### Hardware Ordering Mechanisms (to be evaluated once coherency and interconnect are known)
PROPOSAL — The following mechanisms should be considered for fence/ordering enforcement:
1. **Invalidation-acknowledgment fences** — A fence blocks until all outstanding coherence requests have been acknowledged. Simple, common, but scales poorly with directory hop count and aggregate invalidation traffic at 128 cores.
2. **Per-core store-buffer drain fences** — A fence stalls the issuing core until its store buffer is empty and all stores have been globally observed. Predictable latency; serializes through the coherence fabric.
3. **Tournament / Token-based fences** — Fences acquire a global token (Coherence-Next-Line style) or wait for a "fence epoch" counter. Avoids worst-case drain, but adds global state.
4. **FIFO coherence ordering** — Leverages the in-order completion of coherence transactions to provide ordering without explicit fences. Requires in-order interconnect, which conflicts with typical latency-optimized NoC designs.
5. **Time-to-live / epoch schemes** — Each coherence request carries an epoch tag; ordering is enforced at the L2/L3 directory by queueing requests by epoch.
6. **Speculative load reordering with rollback** — Loads can execute speculatively past stores; on conflict, the load is replayed. Common in high-performance out-of-order cores (e.g., IBM POWER).
ASSUMPTION — XH-1's microarchitecture style is not yet established. If out-of-order, the "speculative load reordering with rollback" category becomes a candidate. If in-order, mechanism choices narrow considerably.
## Comparison
| Property | SC | RVWMO | RVWMO + Ztso | RC variants |
|----------|----|----|----|----|
| Software ease | Highest | Medium | Medium-High | Low-Medium |
| Hardware cost | Highest | Low | Low-Medium | Low |
| Performance headroom | Lowest | Highest | High | High |
| Fence latency | N/A | High | Medium | Variable |
| Verification cost | Lowest | High | High | Highest |
| C11/C++ alignment | Approximate | Approximate | Weak | Direct |
| 128-core scalability | Poor | Good | Good | Good |
Quantitative figures: INSUFFICIENT EVIDENCE. All numbers above are relative qualitative ratings, not benchmarked measurements.
## Advantages
### RVWMO (Baseline)
- **Lowest hardware cost** of all candidates for the base case — most orderings are simply not enforced.
- **Maximum performance** potential — reordering is unconstrained.
- **Familiar** to RISC-V software ecosystem (Linux, RISC-V GCC, LLVM).
- **Standardized** — concrete compliance test suite exists.
### RVWMO + Ztso
- **x86-like TSO** for software that benefits from stronger ordering.
- **Backwards compatible** — software not using Ztso still runs correctly under RVWMO.
- **Modest hardware cost** — primarily requires store-store ordering at the coherence layer.
### Release Consistency Variants
- **Excellent fit for C11/C++** atomics — acquire/release are the most-used fence types in real code.
- **Cheapest common case** — acquire and release can be implemented with lightweight mechanisms.
- **Hardware cost is concentrated on the rare case** of full seq_cst fences.
## Disadvantages
### RVWMO (Baseline)
- **Fence latency is high** at 128 cores — global acknowledgment is expensive.
- **Verification complexity is high** — many legal reorderings; the litmus-test failure surface is large.
- **Software burden** — kernel and runtime code must insert fences correctly.
### RVWMO + Ztso
- **Two consistency models in one chip** — adds documentation, validation, and software education cost.
- **In-flight mixing** of ordering regimes in the same software is hard to reason about.
### Release Consistency Variants
- **Non-standard for RISC-V** — would require ISA extension. Cannot be recommended without a strong software-side driver.
- **Verification cost is the highest** of the candidates because of subtle fence semantics.
### Sequential Consistency
- **Performance cost is severe at 128 cores** — every load may stall on store-buffer drain.
- **No selective escape valve** — software cannot opt out of the strongest model.
- **Does not match RISC-V ecosystem expectations** — surprising to RISC-V software developers.
## XH-1 Considerations
OPEN QUESTION — The following XH-1-internal questions are unresolved by the repository:
1. Is XH-1's pipeline in-order or out-of-order?
2. What is the coherence protocol (MOESI, MESI, directory-based, broadcast)?
3. Is the interconnect a NoC, ring, or crossbar? Does it preserve in-order delivery of coherence responses?
4. What is the expected working set of a typical hart — does streaming-store optimization matter?
5. Does XH-1 target HPC, server, embedded, or mixed workloads? (Workload affects fence frequency.)
6. Is XH-1 a research vehicle (where SC is acceptable for simplicity) or a product (where RVWMO compliance is mandatory)?
7. Does the memory map use non-coherent regions (e.g., DMA, I/O) that interact with the ordering model?
ASSUMPTION — If XH-1 targets RISC-V ecosystem compatibility (Linux, RISC-V GCC, RISC-V LLVM), then RVWMO is mandatory. Ztso may optionally be added for x86-software porting convenience.
## 128-Core Scalability
PROPOSAL — Scalability analysis of each candidate at 128 cores:
### Fence Latency
- **Fence global acknowledgment** scales roughly with the worst-case coherence round-trip time. In a 128-core system, this can be many tens of cycles in a NoC, hundreds in a ring.
- **Store-buffer drain** latency is approximately proportional to the store buffer depth and the time to invalidate all sharers.
- **Token / epoch fences** scale better because they avoid draining the entire coherence fabric — they synchronize at a logical epoch boundary.
### Coherence Traffic
- **RVWMO fences** generate invalidation or acknowledgment traffic that competes with normal coherence traffic.
- **SC at 128 cores** can saturate a directory-based interconnect because every load's coherence transaction may stall pending store ordering.
- **RC variants** concentrate traffic on the fence path only.
### Verification Scalability
- **State-space explosion**: at 128 cores, the number of concurrently observable memory operations grows combinatorially. Memory-ordering verification is a known scaling bottleneck.
- **RC variants** are hardest to verify because of subtle fence interactions.
- **SC is easiest to verify** in principle, but worst in performance.
INSUFFICIENT EVIDENCE — No quantitative fence-latency or coherence-traffic data is available for XH-1. Estimates cannot be made without knowing the interconnect.
## Performance Considerations
PROPOSAL — Performance-impact dimensions to evaluate:
1. **Steady-state IPC** for memory-bound workloads — RVWMO best, SC worst.
2. **Fence latency** — measured as cycles from fence issue to fence completion. Dominant for synchronization-heavy code.
3. **Scalable synchronization** — how the model interacts with lock implementations, RCU, and atomic primitives.
4. **Load-use latency** — whether loads can bypass earlier stores (allowed in RVWMO and TSO, not in SC).
5. **Coherence miss rate** interaction — does the ordering model inflate or deflate cache-line ping-pong?
ASSUMPTION — For HPC workloads with high-fence frequency (e.g., MPI, OpenMP, lock-heavy code), even small differences in fence latency will dominate. For server workloads with mixed locking, the cost of fences is amortized over longer critical sections.
## Area Considerations
PROPOSAL — Area cost of ordering mechanisms:
| Mechanism | Approximate area cost per core |
|-----------|-------------------------------|
| Store buffer with forwarding | Small (few entries × cache-line width) |
| Load-load reordering queue | Small to medium (grows with IPC) |
| Fence-acknowledgment counter | Negligible (per-core counter) |
| Global epoch / token state | Medium (one per chip, not per core) |
| Speculative load reorder + rollback | Large (load-load queue, replay logic) |
ASSUMPTION — All numbers in the table are order-of-magnitude estimates. Precise XH-1 figures: INSUFFICIENT EVIDENCE.
PROPOSAL — At 128 cores, even small per-core area costs multiply by 128. Global state (epoch counters, fence acknowledgers) is O(1) and becomes attractive.
## Power and Energy Considerations
PROPOSAL — Energy implications:
- **Fences** generate global coherence traffic — energy cost scales with the number of cores that must acknowledge.
- **Speculative load reordering** wastes energy on rollback when a violation is detected.
- **Strong ordering models** keep more coherence state in flight (more invalidations, more retries).
- **RVWMO** with lightweight fences (when software permits) is the lowest-energy regime.
OPEN QUESTION — The 128-core replication means the per-fence energy cost is multiplied. A fence that costs 100 pJ per core at 16 cores costs 800× at 128 cores in the worst case (if all cores must acknowledge). No published figure for XH-1 is available — INSUFFICIENT EVIDENCE.
## Implementation Considerations
PROPOSAL — Implementation order of dependencies (the repository does not establish these, but the research area suggests the following):
1. Define coherency protocol → `cache-coherency.md`
2. Define interconnect topology → not yet in repository
3. Define memory hierarchy → `memory-hierarchy.md`
4. Define memory ordering model → this document
5. Define atomic primitives → `atomics.md`
ASSUMPTION — The Ztso extension is a small RTL delta over RVWMO. The dominant implementation cost is verification, not area.
OPEN QUESTION — Does XH-1 have a coherent accelerator fabric or non-coherent IO (e.g., CXL, DMA engines)? If so, the memory ordering model must define how these agents interact.
## Verification Considerations
PROPOSAL — Verification challenges:
- **Litmus tests**: Standard RISC-V litmus test suite must be run on RTL and on model-based environments.
- **State-space explosion**: 128 cores × N outstanding operations = combinatorial explosion. Likely requires state-space reduction or bisimulation techniques.
- **Fence semantics**: Each fence variant must be exhaustively tested at corner cases (interrupts, exceptions, MMIO).
- **Coherence-ordering interaction**: The ordering model must be checked against the chosen coherence protocol for consistency.
INSUFFICIENT EVIDENCE — No XH-1 verification infrastructure, formal spec, or litmus-test set is documented in the repository.
ASSUMPTION — RVWMO is the most documented and tested model for RISC-V. Ztso is a small addition. SC and RC variants would require custom formal infrastructure.
## Software Considerations
PROPOSAL — Software-side impact:
- **Linux kernel**: Designed for weak memory models with appropriate fences. Compatible with RVWMO.
- **C11/C++ atomics**: Map cleanly onto release/acquire semantics. Compatible with RVWMO and with RC variants.
- **x86 software ports**: Benefit from Ztso for fewer fences.
- **OpenSHMEM, MPI shmem**: Often assume TSO-like ordering; Ztso helps.
- **HPC codes with hand-rolled atomics**: Highly sensitive to fence latency.
ASSUMPTION — XH-1 will run RISC-V Linux. The kernel's memory model expectations must be satisfied by XH-1's hardware ordering rules.
## Recommendation
INSUFFICIENT EVIDENCE to make a final recommendation.
The repository does not yet establish:
- Pipeline microarchitecture
- Coherence protocol
- Interconnect topology
- Workload target
A recommendation can be made only after the following documents in the memory research area are filled in: `memory-architecture.md`, `cache-coherency.md`, `memory-hierarchy.md`. Without these, any model selection is premature.
PROPOSAL — Tentative default pending more repository context:
1. **Baseline**: RVWMO (mandatory for RISC-V ecosystem).
2. **Optional add-on**: Ztso extension for x86-software porting convenience and modest hardware cost.
3. **Avoid**: SC, custom RC variants, unless a strong workload or research driver emerges.
4. **Defer to**: `cache-coherency.md` for the coherence-side ordering mechanism.
## Confidence
| Topic | Confidence | Reason |
|-------|------------|--------|
| RVWMO is the RISC-V baseline | High | Ratified ISA spec |
| Ztso is a small optional extension | High | Ratified extension |
| SC is impractical at 128 cores | Medium | General consensus in literature; XH-1-specific number not available |
| Per-fence energy scales with core count | Medium | Logical consequence; no measurement |
| Verification is hardest for RC variants | Medium | General consensus |
| XH-1 should adopt RVWMO | Low | Repository context not yet sufficient |
| XH-1 should adopt RVWMO + Ztso | Low | Same |
## Open Questions
1. What is XH-1's microarchitecture style (in-order vs. out-of-order)?
2. What coherence protocol and directory organization will XH-1 use?
3. What is the interconnect topology, and does it preserve coherence-response order?
4. What workloads are targeted — HPC, server, embedded?
5. Does XH-1 include non-coherent agents (DMA, CXL, accelerators)?
6. Is XH-1 a research vehicle, a prototype, or a product?
7. What is the expected fence frequency in target software?
8. What is the verification methodology — formal, simulation, or hybrid?
9. Will XH-1 use the standard RISC-V litmus-test suite or define a custom one?
10. Does XH-1 require compatibility with x86-software-only ports? (Drives Ztso inclusion.)
## Sources
INSUFFICIENT EVIDENCE — The XH-1 Research repository does not yet contain source material specific to memory ordering. No citations, papers, or measurements can be honestly attributed to XH-1.
External ISA-level facts cited above (RVWMO, Ztso, FENCE instruction) are drawn from the publicly ratified RISC-V ISA specification, which is the source of truth for the ISA-level memory model. No specific paper, version, or page reference is provided because the XH-1 repository does not document which revision is being targeted.
No benchmarks, no measurements, no RTL data, no simulation results are available in the repository at this time.
File diff suppressed because one or more lines are too long
@@ -0,0 +1,323 @@
# Memory Ordering
## Status
Stub document. XH-1 memory ordering model has not yet been defined. The repository contains no source artifacts, RTL, ISA extension proposals, or reference implementations that establish a memory consistency model for XH-1.
## Abstract
Memory ordering defines the rules by which loads and stores issued by one or more cores become visible to other cores in the system. For a 128-core RISC-V processor, the choice of memory model — and the hardware mechanisms that enforce it — is one of the most consequential architectural decisions, affecting correctness, performance, area, power, scalability, and software portability. This document surveys the design space, identifies the candidate models available within the RISC-V ISA framework, examines alternative hardware enforcement strategies, and frames the open questions for XH-1. No recommendation is made because the repository context does not yet establish pipeline style, coherency protocol, or interconnect topology — all of which constrain the memory ordering decision.
## Research Question
What memory consistency model and hardware ordering mechanism should XH-1 adopt, given:
1. A 128-core replication target
2. The RISC-V ISA baseline
3. The need to balance software portability against hardware implementation cost
4. Unknown microarchitecture style (in-order vs. out-of-order is not established in the repository)
5. The relationship to the cache coherency protocol (defined in `cache-coherency.md`, currently SOON)
## Background
### Memory Consistency Models — Terminology
PROPOSAL — Use the following working taxonomy in this document:
- **Sequential Consistency (SC)**: All cores observe a single total order of memory operations that respects each core's program order.
- **Total Store Order (TSO)**: Loads may bypass stores; stores from one core are seen in order by all cores. Used by x86.
- **Relaxed / Weak Memory Models**: Many orderings are allowed; fences (barriers) are required to enforce constraints. Used by ARM, RISC-V, and PowerPC.
- **Release Consistency (RC) / RCO / RCpc / RCsc**: A family of models with different fence semantics. Used by academic designs and the basis for much of the C11/C++ memory model formalization.
### The RISC-V "RVWMO" Model
FACT — The RISC-V ISA specification defines the **RISC-V Weak Memory Ordering (RVWMO)** model as the baseline. Under RVWMO:
- Loads and stores to the same address from the same hart are ordered.
- A store followed by a load to a *different* address may be reordered.
- The `FENCE` instruction is used to enforce orderings. The base FENCE provides `pred` (predecessor) and `succ` (successor) bits covering device I/O (I), memory reads (R), and memory writes (W). Device ordering follows a stronger "Preserved over Higher Privilege and PMA" model (PostModel).
- The optional `Ztso` extension provides **RVTSO**, which gives x86-like TSO semantics with a relaxed form of the FENCE.
ASSUMPTION — RVWMO + the Ztso optional extension is the available ISA-level model space that XH-1 must select from. (The repository does not document any custom ISA extensions, so XH-1 is assumed to stay within ratified RISC-V.)
### Why Memory Ordering is Hard at 128 Cores
OPEN QUESTION — The repository does not establish:
- Whether XH-1 is a single die, multi-die package, or chiplet-based system
- The interconnect topology (mesh, ring, crossbar, NoC, etc.)
- Whether coherence is directory-based, broadcast-based, or hybrid
- The private-vs-shared L2/L3 split
- The memory-side directory or snoop filter organization
These decisions have direct bearing on which ordering mechanisms are feasible. Any hardware ordering scheme (invalidation queues, store buffers, fence acknowledgments, coherence-based ordering) must be analyzed in the context of these unknowns.
## Existing Approaches
### Approach A — RISC-V RVWMO (Baseline)
The standard weak memory model provided by the RISC-V ISA. Requires implementation of `FENCE`, `FENCE.I`, and the I/O ordering semantics.
**Mechanism**: Architecture permits aggressive reordering. Hardware typically uses:
- Store buffers with forwarding
- Load-load reordering queues
- Coherence-based ordering: stores become visible at coherence invalidation/upgrade points
- Fence implementation via global acknowledgment counters or by draining reorder queues
### Approach B — RVWMO + Ztso Extension
Adds RVTSO. Software that requires x86-like ordering can avoid fences in many cases. Still RISC-V compliant.
**Mechanism**: Stores from one hart are serialized in the interconnect; loads may still pass earlier stores to different addresses.
### Approach C — Sequential Consistency (SC)
All operations observed in a single global order respecting program order. Simplest software model; most expensive hardware.
**Mechanism**: Typically requires store-buffer draining on loads, broadcast-based coherence, and conservative replay mechanisms. Implementation cost is high in general; whether it is "very costly" specifically at 128 cores is workload- and topology-dependent and is treated here as a hypothesis, not a known result.
### Approach D — Release Consistency variants (RCsc, RCpc, RCO)
Academic models where release/acquire ordering is cheap but stronger fences are required for full memory ordering. Basis of C11/C++ model.
**Mechanism**: Only release/acquire fences require acknowledgment; acquire-only loads are cheap. Alignment with C11 is a strong software benefit. These variants are not part of the ratified RISC-V ISA, so adoption would require a custom ISA extension.
## Alternative Designs
### Hardware Ordering Mechanisms (to be evaluated once coherency and interconnect are known)
PROPOSAL — The following mechanisms should be considered for fence/ordering enforcement:
1. **Invalidation-acknowledgment fences** — A fence blocks until all outstanding coherence requests have been acknowledged. Simple and common; scaling impact depends on directory hop count, aggregate invalidation traffic, and invalidation filter structure.
2. **Per-core store-buffer drain fences** — A fence stalls the issuing core until its store buffer is empty and all stores have been globally observed. Predictable latency; serializes through the coherence fabric.
3. **Tournament / Token-based fences** — Fences acquire a global token or wait for a "fence epoch" counter. Avoids worst-case drain, but adds global state.
4. **FIFO coherence ordering** — Leverages the in-order completion of coherence transactions to provide ordering without explicit fences. Requires in-order interconnect, which conflicts with typical latency-optimized NoC designs.
5. **Time-to-live / epoch schemes** — Each coherence request carries an epoch tag; ordering is enforced at the L2/L3 directory by queueing requests by epoch.
6. **Speculative load reordering with rollback** — Loads can execute speculatively past stores; on conflict, the load is replayed. Common in high-performance out-of-order cores (e.g., IBM POWER).
ASSUMPTION — XH-1's microarchitecture style is not yet established. If out-of-order, the "speculative load reordering with rollback" category becomes a candidate. If in-order, mechanism choices narrow considerably.
## Comparison
| Property | SC | RVWMO | RVWMO + Ztso | RC variants |
|----------|----|----|----|----|
| Software ease | Highest | Medium | Medium-High | Low-Medium |
| Hardware cost | Highest | Low | Low-Medium | Low |
| Performance headroom | Lowest | Highest | High | High |
| Fence latency | N/A | Variable | Variable | Variable |
| Verification cost | Lowest in principle | High | High | High (likely highest) |
| C11/C++ alignment | Approximate | Approximate | Weak | Direct |
| 128-core scalability | Poor (hypothesis) | Good (claim) | Good (claim) | Good (claim) |
INSUFFICIENT EVIDENCE — Quantitative fence-latency, coherence-traffic, and scalability numbers for XH-1 are unavailable. Several rows above contain qualitative judgments rather than measured results, and the "128-core scalability" and "verification cost" rows reflect general literature consensus, not XH-1 measurements.
## Advantages
### RVWMO (Baseline)
- **Lowest hardware cost** of all candidates for the base case — most orderings are simply not enforced.
- **Maximum performance** potential — reordering is unconstrained.
- **Familiar** to RISC-V software ecosystem (Linux, RISC-V GCC, LLVM).
- **Standardized** — concrete compliance test suite exists.
### RVWMO + Ztso
- **x86-like TSO** for software that benefits from stronger ordering.
- **Backwards compatible** — software not using Ztso still runs correctly under RVWMO.
- **Modest hardware cost** — primarily requires store-store ordering at the coherence layer.
### Release Consistency Variants
- **Excellent fit for C11/C++** atomics — acquire/release are the most-used fence types in real code.
- **Cheapest common case** — acquire and release can be implemented with lightweight mechanisms.
- **Hardware cost is concentrated on the rare case** of full seq_cst fences.
## Disadvantages
### RVWMO (Baseline)
- **Fence latency depends on coherence round-trip** — at 128 cores, fence latency is expected to be high because global acknowledgment must traverse the coherence fabric. Specific cycle counts: INSUFFICIENT EVIDENCE.
- **Verification complexity is high** — many legal reorderings; the litmus-test failure surface is large.
- **Software burden** — kernel and runtime code must insert fences correctly.
### RVWMO + Ztso
- **Two consistency models in one chip** — adds documentation, validation, and software education cost.
- **In-flight mixing** of ordering regimes in the same software is hard to reason about.
### Release Consistency Variants
- **Non-standard for RISC-V** — would require a custom ISA extension. Cannot be recommended without a strong software-side driver.
- **Verification cost is high** in general because of subtle fence semantics; whether it is the "highest" among candidates is a claim about the literature, not a measured XH-1 result.
### Sequential Consistency
- **Performance cost is severe in the general case** — every load may stall on store-buffer drain; whether this is "severe at 128 cores" specifically is workload- and topology-dependent and is not established by the repository.
- **No selective escape valve** — software cannot opt out of the strongest model.
- **Does not match RISC-V ecosystem expectations** — surprising to RISC-V software developers.
## XH-1 Considerations
OPEN QUESTION — The following XH-1-internal questions are unresolved by the repository:
1. Is XH-1's pipeline in-order or out-of-order?
2. What is the coherence protocol (MOESI, MESI, directory-based, broadcast)?
3. Is the interconnect a NoC, ring, or crossbar? Does it preserve in-order delivery of coherence responses?
4. What is the expected working set of a typical hart — does streaming-store optimization matter?
5. Does XH-1 target HPC, server, embedded, or mixed workloads? (Workload affects fence frequency.)
6. Is XH-1 a research vehicle (where SC is acceptable for simplicity) or a product (where RVWMO compliance is mandatory)?
7. Does the memory map use non-coherent regions (e.g., DMA, I/O) that interact with the ordering model?
ASSUMPTION — If XH-1 targets RISC-V ecosystem compatibility (Linux, RISC-V GCC, RISC-V LLVM), then RVWMO is mandatory. Ztso may optionally be added for x86-software porting convenience.
## 128-Core Scalability
PROPOSAL — Scalability analysis of each candidate at 128 cores. Quantitative figures: INSUFFICIENT EVIDENCE. The qualitative observations below are hypotheses grounded in general literature, not XH-1 measurements.
### Fence Latency
- **Fence global acknowledgment** latency is bounded below by the worst-case coherence round-trip time. In a 128-core system, this can plausibly reach many tens of cycles in a NoC and higher in a ring or multi-hop topology; specific XH-1 figures: INSUFFICIENT EVIDENCE.
- **Store-buffer drain** latency is approximately proportional to the store buffer depth and the time to invalidate all sharers.
- **Token / epoch fences** may scale better because they avoid draining the entire coherence fabric — they synchronize at a logical epoch boundary.
### Coherence Traffic
- **RVWMO fences** generate invalidation or acknowledgment traffic that competes with normal coherence traffic. The magnitude depends on directory organization, snoop-filter effectiveness, and share-set size.
- **SC at 128 cores** can plausibly saturate a directory-based interconnect because every load's coherence transaction may stall pending store ordering; this is a hypothesis, not a measured XH-1 result.
- **RC variants** concentrate traffic on the fence path only.
### Verification Scalability
- **State-space explosion**: at 128 cores, the number of concurrently observable memory operations grows combinatorially in the unmitigated model. Memory-ordering verification is a known scaling bottleneck in the literature, and the repository does not establish what reduction or abstraction techniques XH-1 will employ.
- **RC variants** are commonly cited as among the hardest to verify because of subtle fence interactions; whether they are hardest in the absolute sense is a literature claim, not an XH-1 measurement.
- **SC is easiest to verify in principle** because it admits fewer legal behaviors, but worst in performance.
INSUFFICIENT EVIDENCE — No quantitative fence-latency or coherence-traffic data is available for XH-1. Estimates cannot be made without knowing the interconnect.
## Performance Considerations
PROPOSAL — Performance-impact dimensions to evaluate:
1. **Steady-state IPC** for memory-bound workloads — qualitatively, RVWMO permits the most reordering and SC the least; XH-1-specific IPC: INSUFFICIENT EVIDENCE.
2. **Fence latency** — measured as cycles from fence issue to fence completion. Dominant for synchronization-heavy code. XH-1-specific: INSUFFICIENT EVIDENCE.
3. **Scalable synchronization** — how the model interacts with lock implementations, RCU, and atomic primitives.
4. **Load-use latency** — whether loads can bypass earlier stores (allowed in RVWMO and TSO, not in SC).
5. **Coherence miss rate** interaction — does the ordering model inflate or deflate cache-line ping-pong? Model-dependent; XH-1-specific: INSUFFICIENT EVIDENCE.
ASSUMPTION — For HPC workloads with high-fence frequency (e.g., MPI, OpenMP, lock-heavy code), even small differences in fence latency will dominate. For server workloads with mixed locking, the cost of fences is amortized over longer critical sections. No XH-1 workload profile is documented in the repository.
## Area Considerations
PROPOSAL — Qualitative area cost of ordering mechanisms:
| Mechanism | Approximate area cost per core |
|-----------|-------------------------------|
| Store buffer with forwarding | Small (few entries × cache-line width) |
| Load-load reordering queue | Small to medium (grows with IPC and reorder depth) |
| Fence-acknowledgment counter | Negligible (per-core counter) |
| Global epoch / token state | Medium (one per chip, not per core) |
| Speculative load reorder + rollback | Large (load-load queue, replay logic) |
INSUFFICIENT EVIDENCE — All descriptions above are qualitative. Precise XH-1 area figures (in gates, µm², or mm²) are unavailable.
PROPOSAL — At 128 cores, even small per-core area costs multiply by 128. Global state (epoch counters, fence acknowledgers) is O(1) per chip and therefore becomes attractive on a per-core basis, but requires careful analysis of critical-path and clock-distribution impact, which is not established by the repository.
## Power and Energy Considerations
PROPOSAL — Qualitative energy implications:
- **Fences** generate global coherence traffic — energy cost scales with the number of cores that must acknowledge. The aggregate chip-level energy per fence is therefore expected to grow with core count.
- **Speculative load reordering** wastes energy on rollback when a violation is detected.
- **Strong ordering models** keep more coherence state in flight (more invalidations, more retries).
- **RVWMO** with lightweight fences (when software permits) is the lowest-energy regime in the general case.
INSUFFICIENT EVIDENCE — The earlier text claimed that a 100 pJ-per-core fence at 16 cores costs 800× at 128 cores in the worst case (if all cores must acknowledge). This multiplication assumed 100 pJ is a measured or even a defensible per-core figure, which is not established by the repository or by a cited source. The 100 pJ figure is removed. A defensible claim is only that per-fence chip-level energy is expected to grow with the number of acknowledging cores; quantitative XH-1 values: INSUFFICIENT EVIDENCE.
## Implementation Considerations
PROPOSAL — Implementation order of dependencies (the repository does not establish these, but the research area suggests the following):
1. Define coherency protocol → `cache-coherency.md`
2. Define interconnect topology → not yet in repository
3. Define memory hierarchy → `memory-hierarchy.md`
4. Define memory ordering model → this document
5. Define atomic primitives → `atomics.md`
ASSUMPTION — The Ztso extension is a small RTL delta over RVWMO. The dominant implementation cost is verification, not area. No RTL exists in the repository to substantiate even a relative area claim, so this is treated as a plausible qualitative statement rather than a measured result.
OPEN QUESTION — Does XH-1 have a coherent accelerator fabric or non-coherent IO (e.g., CXL, DMA engines)? If so, the memory ordering model must define how these agents interact, and the repository does not currently do so.
## Verification Considerations
PROPOSAL — Verification challenges:
- **Litmus tests**: A standard RISC-V litmus test suite is available externally; whether XH-1 will adopt it as-is or define a custom suite is OPEN.
- **State-space explosion**: 128 cores × N outstanding operations produces a combinatorial unmitigated state space. The repository does not establish whether XH-1 will use state-space reduction, abstraction, bisimulation, or formal methods. Any specific XH-1 claim: INSUFFICIENT EVIDENCE.
- **Fence semantics**: Each fence variant must be exhaustively tested at corner cases (interrupts, exceptions, MMIO).
- **Coherence-ordering interaction**: The ordering model must be checked against the chosen coherence protocol for consistency. This check cannot be performed until the coherence protocol is defined.
INSUFFICIENT EVIDENCE — No XH-1 verification infrastructure, formal spec, or litmus-test set is documented in the repository.
ASSUMPTION — RVWMO is the most documented and tested model for RISC-V. Ztso is a small addition. SC and RC variants would require custom formal infrastructure. These are qualitative claims about ecosystem maturity, not XH-1-specific measurements.
## Software Considerations
PROPOSAL — Software-side impact:
- **Linux kernel**: Designed for weak memory models with appropriate fences. Compatible with RVWMO.
- **C11/C++ atomics**: Map cleanly onto release/acquire semantics. Compatible with RVWMO and with RC variants.
- **x86 software ports**: Benefit from Ztso for fewer fences.
- **OpenSHMEM, MPI shmem**: Often assume TSO-like ordering; Ztso helps.
- **HPC codes with hand-rolled atomics**: Highly sensitive to fence latency.
ASSUMPTION — XH-1 will run RISC-V Linux. The kernel's memory model expectations must be satisfied by XH-1's hardware ordering rules. The repository does not document whether Linux is in fact the target.
## Recommendation
INSUFFICIENT EVIDENCE to make a final recommendation.
The repository does not yet establish:
- Pipeline microarchitecture
- Coherence protocol
- Interconnect topology
- Workload target
A recommendation can be made only after the following documents in the memory research area are filled in: `memory-architecture.md`, `cache-coherency.md`, `memory-hierarchy.md`. Without these, any model selection is premature.
PROPOSAL — Tentative default pending more repository context:
1. **Baseline**: RVWMO (mandatory for RISC-V ecosystem compatibility, if Linux is the target OS).
2. **Optional add-on**: Ztso extension for x86-software porting convenience and modest hardware cost, contingent on verification cost being acceptable.
3. **Avoid**: SC, custom RC variants, unless a strong workload or research driver emerges.
4. **Defer to**: `cache-coherency.md` for the coherence-side ordering mechanism.
## Confidence
| Topic | Confidence | Reason |
|-------|------------|--------|
| RVWMO is the RISC-V baseline | High | Ratified ISA spec |
| Ztso is a small optional extension | High | Ratified extension |
| SC has high implementation cost | High | General literature consensus |
| SC specifically impractical at 128 cores | Low | General claim, not XH-1-measured; workload- and topology-dependent |
| Per-fence energy scales with acknowledging-core count | Medium | Logical consequence; no measurement |
| Quantitative per-fence energy at 128 cores | None | INSUFFICIENT EVIDENCE |
| Verification is hardest for RC variants | LowMedium | Literature claim, not XH-1-measured |
| XH-1 should adopt RVWMO | Low | Repository context not yet sufficient |
| XH-1 should adopt RVWMO + Ztso | Low | Same |
## Open Questions
1. What is XH-1's microarchitecture style (in-order vs. out-of-order)?
2. What coherence protocol and directory organization will XH-1 use?
3. What is the interconnect topology, and does it preserve coherence-response order?
4. What workloads are targeted — HPC, server, embedded?
5. Does XH-1 include non-coherent agents (DMA, CXL, accelerators)?
6. Is XH-1 a research vehicle, a prototype, or a product?
7. What is the expected fence frequency in target software?
8. What is the verification methodology — formal, simulation, or hybrid?
9. Will XH-1 use the standard RISC-V litmus-test suite or define a custom one?
10. Does XH-1 require compatibility with x86-software-only ports? (Drives Ztso inclusion.)
11. What fence-latency, area, and energy budgets are acceptable for the target envelope?
12. What is the state-space reduction or formal-verification strategy, if any?
## Sources
INSUFFICIENT EVIDENCE — The XH-1 Research repository does not yet contain source material specific to memory ordering. No citations, papers, or measurements can be honestly attributed to XH-1.
External ISA-level facts cited above (RVWMO, Ztso, FENCE instruction semantics) are drawn from the publicly ratified RISC-V ISA specification, which is the source of truth for the ISA-level memory model. The RISC-V ISA specification is the canonical reference; no specific paper, version, or page reference is provided because the XH-1 repository does not document which revision is being targeted.
INSUFFICIENT EVIDENCE — All quantitative claims about fence latency, area, power, energy, and verification cost at the XH-1 level have been removed or flagged. No benchmarks, no measurements, no RTL data, no simulation results are available in the repository at this time.
File diff suppressed because one or more lines are too long
@@ -0,0 +1,333 @@
# Memory Ordering
## Status
Stub document. XH-1 memory ordering model has not yet been defined. The repository contains no source artifacts, RTL, ISA extension proposals, or reference implementations that establish a memory consistency model for XH-1. All quantitative or XH-1-specific claims in this document are flagged INSUFFICIENT EVIDENCE where applicable.
## Abstract
Memory ordering defines the rules by which loads and stores issued by one or more cores become visible to other cores in the system. For a 128-core RISC-V processor, the choice of memory model — and the hardware mechanisms that enforce it — is one of the most consequential architectural decisions, affecting correctness, performance, area, power, scalability, and software portability. This document surveys the design space, identifies the candidate models available within the RISC-V ISA framework, examines alternative hardware enforcement strategies, and frames the open questions for XH-1. No final recommendation is made because the repository does not establish pipeline style, coherency protocol, or interconnect topology, all of which constrain the memory ordering decision. A conditional tentative default is stated, contingent on the assumed constraints being correct.
## Research Question
What memory consistency model and hardware ordering mechanism should XH-1 adopt, given:
1. A 128-core replication target
2. The RISC-V ISA baseline
3. The need to balance software portability against hardware implementation cost
4. Unknown microarchitecture style (in-order vs. out-of-order is not established in the repository)
5. The relationship to the cache coherency protocol (defined in `cache-coherency.md`, currently SOON)
## Background
### Memory Consistency Models — Terminology
PROPOSAL — Use the following working taxonomy in this document:
- **Sequential Consistency (SC)**: All cores observe a single total order of memory operations that respects each core's program order.
- **Total Store Order (TSO)**: Loads may bypass stores from the same core; stores from one core are seen in order by all cores. Used by x86.
- **Relaxed / Weak Memory Models**: Many orderings are allowed; fences (barriers) are required to enforce constraints. Used by ARM, RISC-V, and PowerPC.
- **Release Consistency (RC) / RCsc / RCpc / RCO**: A family of models with different fence semantics. Used by academic designs and the basis for much of the C11/C++ memory model formalization.
ASSUMPTION — This document uses "RC variants" as an umbrella term for RCsc, RCpc, and RCO. The specific semantics of each variant are not redefined here; readers are referred to the literature.
### The RISC-V "RVWMO" Model
FACT — The RISC-V ISA specification defines the **RISC-V Weak Memory Ordering (RVWMO)** model as the baseline. Under RVWMO:
- Loads and stores to the same address from the same hart are ordered.
- A store followed by a load to a *different* address may be reordered.
- The `FENCE` instruction is used to enforce orderings. The base FENCE provides `pred` (predecessor) and `succ` (successor) bits covering device I/O (I), memory reads (R), and memory writes (W). Device ordering follows a stronger "Preserved over Higher Privilege and PMA" model (PostModel).
- The optional `Ztso` extension provides **RVTSO**, which gives x86-like TSO semantics with a relaxed form of the FENCE.
FACT — RVWMO and the `Ztso` extension are documented in the ratified RISC-V Unprivileged ISA specification. The repository does not state which revision of the specification is the reference for XH-1; this is an open question.
ASSUMPTION — RVWMO + the Ztso optional extension is the available ISA-level model space that XH-1 must select from. (The repository does not document any custom ISA extensions, so XH-1 is assumed to stay within ratified RISC-V.)
### Why Memory Ordering is Hard at 128 Cores
OPEN QUESTION — The repository does not establish:
- Whether XH-1 is a single die, multi-die package, or chiplet-based system
- The interconnect topology (mesh, ring, crossbar, NoC, etc.)
- Whether coherence is directory-based, broadcast-based, or hybrid
- The private-vs-shared L2/L3 split
- The memory-side directory or snoop filter organization
These decisions have direct bearing on which ordering mechanisms are feasible. Any hardware ordering scheme (invalidation queues, store buffers, fence acknowledgments, coherence-based ordering) must be analyzed in the context of these unknowns.
## Existing Approaches
### Approach A — RISC-V RVWMO (Baseline)
The standard weak memory model provided by the RISC-V ISA. Requires implementation of `FENCE`, `FENCE.I`, and the I/O ordering semantics.
**Mechanism**: Architecture permits aggressive reordering. Hardware implementations commonly use:
- Store buffers with forwarding
- Load-load reordering queues
- Coherence-based ordering: stores become visible at coherence invalidation/upgrade points
- Fence implementation via global acknowledgment counters or by draining reorder queues
INSUFFICIENT EVIDENCE — Which of these mechanisms XH-1 will use is not established by the repository.
### Approach B — RVWMO + Ztso Extension
Adds RVTSO. Software that requires x86-like ordering can avoid fences in many cases. Still RISC-V compliant.
**Mechanism**: Stores from one hart are serialized in the interconnect; loads may still pass earlier stores to different addresses.
### Approach C — Sequential Consistency (SC)
All operations observed in a single global order respecting program order. Simplest software model; most expensive hardware in general.
**Mechanism**: Commonly implemented via store-buffer draining on loads, broadcast-based coherence, and conservative replay mechanisms. The qualitative characterization of "very costly" is a general literature claim, not an XH-1 measurement; whether SC is specifically costly at 128 cores is workload- and topology-dependent and is treated here as a hypothesis, not a known result.
### Approach D — Release Consistency variants (RCsc, RCpc, RCO)
Academic models where release/acquire ordering is cheap but stronger fences are required for full memory ordering. Basis of C11/C++ model.
**Mechanism**: Only release/acquire fences require acknowledgment; acquire-only loads are cheap. Alignment with C11 is a strong software benefit. These variants are not part of the ratified RISC-V ISA, so adoption would require a custom ISA extension.
## Alternative Designs
### Hardware Ordering Mechanisms
PROPOSAL — The following mechanisms should be considered for fence/ordering enforcement once coherency and interconnect are known:
1. **Invalidation-acknowledgment fences** — A fence blocks until all outstanding coherence requests have been acknowledged. Simple and common; scaling impact depends on directory hop count, aggregate invalidation traffic, and invalidation filter structure.
2. **Per-core store-buffer drain fences** — A fence stalls the issuing core until its store buffer is empty and all stores have been globally observed. Predictable latency; serializes through the coherence fabric.
3. **Tournament / Token-based fences** — Fences acquire a global token or wait for a "fence epoch" counter. Avoids worst-case drain, but adds global state.
4. **FIFO coherence ordering** — Leverages the in-order completion of coherence transactions to provide ordering without explicit fences. Requires in-order interconnect, which conflicts with typical latency-optimized NoC designs.
5. **Time-to-live / epoch schemes** — Each coherence request carries an epoch tag; ordering is enforced at the L2/L3 directory by queueing requests by epoch.
6. **Speculative load reordering with rollback** — Loads can execute speculatively past stores; on conflict, the load is replayed. Common in high-performance out-of-order cores (e.g., reported in the IBM POWER literature).
ASSUMPTION — XH-1's microarchitecture style is not yet established. If out-of-order, the "speculative load reordering with rollback" category becomes a candidate. If in-order, mechanism choices narrow considerably.
## Comparison
| Property | SC | RVWMO | RVWMO + Ztso | RC variants |
|----------|----|----|----|----|
| Software ease | Highest | Medium | Medium-High | Low-Medium |
| Hardware cost (general) | Highest | Low | Low-Medium | Low |
| Performance headroom (general) | Lowest | Highest | High | High |
| Fence latency (general) | N/A | Variable | Variable | Variable |
| Verification cost (general) | Lowest in principle | High | High | High |
| C11/C++ alignment | Approximate | Approximate | Weak | Direct |
| 128-core scalability | Poor (hypothesis) | Unmeasured | Unmeasured | Unmeasured |
INSUFFICIENT EVIDENCE — Quantitative fence-latency, coherence-traffic, and scalability numbers for XH-1 are unavailable. Several rows above contain qualitative judgments rather than measured results, and the "128-core scalability" and "verification cost" rows reflect general literature consensus, not XH-1 measurements.
## Advantages
### RVWMO (Baseline)
- **Lowest hardware cost** of all candidates for the base case, as a general property of the model — most orderings are simply not enforced.
- **Maximum performance** potential, as a general property — reordering is unconstrained.
- **Familiar** to the RISC-V software ecosystem (Linux, RISC-V GCC, LLVM), subject to the open question of whether Linux is the target OS.
- **Standardized** — concrete compliance test suites exist externally.
### RVWMO + Ztso
- **x86-like TSO** for software that benefits from stronger ordering.
- **Backwards compatible** — software not using Ztso still runs correctly under RVWMO.
- **Modest hardware cost** — primarily requires store-store ordering at the coherence layer (qualitative claim, not measured for XH-1).
### Release Consistency Variants
- **Excellent fit for C11/C++** atomics — acquire/release are the most-used fence types in real code, per the C/C++ standards.
- **Cheapest common case** — acquire and release can be implemented with lightweight mechanisms, as a general property.
- **Hardware cost is concentrated on the rare case** of full seq_cst fences, as a general property.
## Disadvantages
### RVWMO (Baseline)
- **Fence latency depends on coherence round-trip** — at 128 cores, fence latency is expected to be high because global acknowledgment must traverse the coherence fabric. Specific cycle counts: INSUFFICIENT EVIDENCE.
- **Verification complexity is high** — many legal reorderings; the litmus-test failure surface is large. This is a general property of weak models, not an XH-1 measurement.
- **Software burden** — kernel and runtime code must insert fences correctly.
### RVWMO + Ztso
- **Two consistency models in one chip** — adds documentation, validation, and software education cost.
- **In-flight mixing** of ordering regimes in the same software is hard to reason about.
### Release Consistency Variants
- **Non-standard for RISC-V** — would require a custom ISA extension. Cannot be recommended without a strong software-side driver.
- **Verification cost is high** in general because of subtle fence semantics; whether it is the "highest" among candidates is a claim about the literature, not a measured XH-1 result.
### Sequential Consistency
- **Performance cost is severe in the general case** — every load may stall on store-buffer drain; whether this is "severe at 128 cores" specifically is workload- and topology-dependent and is not established by the repository.
- **No selective escape valve** — software cannot opt out of the strongest model.
- **Does not match RISC-V ecosystem expectations** — surprising to RISC-V software developers, as a general observation.
## XH-1 Considerations
OPEN QUESTION — The following XH-1-internal questions are unresolved by the repository:
1. Is XH-1's pipeline in-order or out-of-order?
2. What is the coherence protocol (MOESI, MESI, directory-based, broadcast)?
3. Is the interconnect a NoC, ring, or crossbar? Does it preserve in-order delivery of coherence responses?
4. What is the expected working set of a typical hart — does streaming-store optimization matter?
5. Does XH-1 target HPC, server, embedded, or mixed workloads? (Workload affects fence frequency.)
6. Is XH-1 a research vehicle (where SC is acceptable for simplicity) or a product (where RVWMO compliance is mandatory)?
7. Does the memory map use non-coherent regions (e.g., DMA, I/O) that interact with the ordering model?
ASSUMPTION — If XH-1 targets RISC-V ecosystem compatibility (Linux, RISC-V GCC, RISC-V LLVM), then RVWMO is mandatory. Ztso may optionally be added for x86-software porting convenience.
## 128-Core Scaling Considerations
INSUFFICIENT EVIDENCE — Quantitative figures for XH-1 at 128 cores are unavailable. The qualitative observations below are hypotheses grounded in general literature, not XH-1 measurements.
### Fence Latency
- **Fence global acknowledgment** latency is bounded below by the worst-case coherence round-trip time. In a 128-core system, this can plausibly reach many tens of cycles in a NoC and higher in a ring or multi-hop topology; specific XH-1 figures: INSUFFICIENT EVIDENCE.
- **Store-buffer drain** latency is approximately proportional to the store buffer depth and the time to invalidate all sharers. Specific XH-1 figures: INSUFFICIENT EVIDENCE.
- **Token / epoch fences** may scale better because they avoid draining the entire coherence fabric — they synchronize at a logical epoch boundary. This is a general property, not an XH-1 measurement.
### Coherence Traffic
- **RVWMO fences** generate invalidation or acknowledgment traffic that competes with normal coherence traffic. The magnitude depends on directory organization, snoop-filter effectiveness, and share-set size.
- **SC at 128 cores** can plausibly saturate a directory-based interconnect because every load's coherence transaction may stall pending store ordering; this is a hypothesis, not a measured XH-1 result.
- **RC variants** concentrate traffic on the fence path only, as a general property.
### Verification Scalability
- **State-space explosion**: at 128 cores, the number of concurrently observable memory operations grows combinatorially in the unmitigated model. Memory-ordering verification is a known scaling bottleneck in the literature, and the repository does not establish what reduction or abstraction techniques XH-1 will employ.
- **RC variants** are commonly cited as among the hardest to verify because of subtle fence interactions; whether they are hardest in the absolute sense is a literature claim, not an XH-1 measurement.
- **SC is easiest to verify in principle** because it admits fewer legal behaviors, but worst in performance.
INSUFFICIENT EVIDENCE — No quantitative fence-latency, coherence-traffic, or verification-scaling data is available for XH-1. Estimates cannot be made without knowing the interconnect, coherence protocol, and verification methodology.
## Performance Considerations
PROPOSAL — Performance-impact dimensions to evaluate once workload and microarchitecture are defined:
1. **Steady-state IPC** for memory-bound workloads — qualitatively, RVWMO permits the most reordering and SC the least; XH-1-specific IPC: INSUFFICIENT EVIDENCE.
2. **Fence latency** — measured as cycles from fence issue to fence completion. Dominant for synchronization-heavy code. XH-1-specific: INSUFFICIENT EVIDENCE.
3. **Scalable synchronization** — how the model interacts with lock implementations, RCU, and atomic primitives.
4. **Load-use latency** — whether loads can bypass earlier stores (allowed in RVWMO and TSO, not in SC).
5. **Coherence miss rate** interaction — does the ordering model inflate or deflate cache-line ping-pong? Model-dependent; XH-1-specific: INSUFFICIENT EVIDENCE.
ASSUMPTION — For HPC workloads with high-fence frequency (e.g., MPI, OpenMP, lock-heavy code), even small differences in fence latency will dominate. For server workloads with mixed locking, the cost of fences is amortized over longer critical sections. No XH-1 workload profile is documented in the repository.
## Area Considerations
INSUFFICIENT EVIDENCE — All area descriptions in this section are qualitative. Precise XH-1 area figures (in gates, µm², or mm²) are unavailable; no synthesis, layout, or RTL data exists in the repository.
PROPOSAL — Qualitative area cost of ordering mechanisms, as a function of microarchitectural choices:
| Mechanism | Approximate area cost per core (qualitative) |
|-----------|----------------------------------------------|
| Store buffer with forwarding | Small (few entries × cache-line width) |
| Load-load reordering queue | Small to medium (grows with IPC and reorder depth) |
| Fence-acknowledgment counter | Negligible (per-core counter) |
| Global epoch / token state | Medium (one per chip, not per core) |
| Speculative load reorder + rollback | Large (load-load queue, replay logic) |
PROPOSAL — At 128 cores, per-core area costs are replicated 128 times. Global state (epoch counters, fence acknowledgers) is O(1) per chip and therefore amortizes favorably on a per-core basis, but the critical-path and clock-distribution impact of global state is not established by the repository.
## Power and Energy Considerations
INSUFFICIENT EVIDENCE — No per-fence, per-core, or aggregate energy figures for XH-1 are available. The earlier version of this document included a "100 pJ-per-core fence" figure that was not substantiated by any repository artifact or external citation; that figure has been removed.
PROPOSAL — Qualitative energy implications, as general properties of the model classes:
- **Fences** generate global coherence traffic — energy cost scales with the number of cores that must acknowledge. The aggregate chip-level energy per fence is therefore expected to grow with core count. Quantitative XH-1 values: INSUFFICIENT EVIDENCE.
- **Speculative load reordering** wastes energy on rollback when a violation is detected, as a general property.
- **Strong ordering models** keep more coherence state in flight (more invalidations, more retries), as a general property.
- **RVWMO** with lightweight fences (when software permits) is the lowest-energy regime in the general case.
## Implementation Considerations
PROPOSAL — Implementation order of dependencies (the repository does not establish these, but the research area suggests the following):
1. Define coherency protocol → `cache-coherency.md`
2. Define interconnect topology → not yet in repository
3. Define memory hierarchy → `memory-hierarchy.md`
4. Define memory ordering model → this document
5. Define atomic primitives → `atomics.md`
ASSUMPTION — The Ztso extension is a small RTL delta over RVWMO. The dominant implementation cost is verification, not area. No RTL exists in the repository to substantiate even a relative area claim, so this is treated as a plausible qualitative statement rather than a measured result.
OPEN QUESTION — Does XH-1 have a coherent accelerator fabric or non-coherent IO (e.g., CXL, DMA engines)? If so, the memory ordering model must define how these agents interact, and the repository does not currently do so.
## Verification Considerations
PROPOSAL — Verification challenges:
- **Litmus tests**: A standard RISC-V litmus test suite is available externally; whether XH-1 will adopt it as-is or define a custom suite is OPEN.
- **State-space explosion**: 128 cores × N outstanding operations produces a combinatorial unmitigated state space. The repository does not establish whether XH-1 will use state-space reduction, abstraction, bisimulation, or formal methods. Any specific XH-1 claim: INSUFFICIENT EVIDENCE.
- **Fence semantics**: Each fence variant must be exhaustively tested at corner cases (interrupts, exceptions, MMIO).
- **Coherence-ordering interaction**: The ordering model must be checked against the chosen coherence protocol for consistency. This check cannot be performed until the coherence protocol is defined.
INSUFFICIENT EVIDENCE — No XH-1 verification infrastructure, formal spec, or litmus-test set is documented in the repository.
ASSUMPTION — RVWMO is the most documented and tested model for RISC-V, per the ratified specification and ecosystem practice. Ztso is a small addition. SC and RC variants would require custom formal infrastructure. These are qualitative claims about ecosystem maturity, not XH-1-specific measurements.
## Software Considerations
PROPOSAL — Software-side impact, as a general property of each model:
- **Linux kernel**: Designed for weak memory models with appropriate fences. Compatible with RVWMO.
- **C11/C++ atomics**: Map cleanly onto release/acquire semantics. Compatible with RVWMO and with RC variants.
- **x86 software ports**: Benefit from Ztso for fewer fences.
- **OpenSHMEM, MPI shmem**: Often assume TSO-like ordering; Ztso helps.
- **HPC codes with hand-rolled atomics**: Highly sensitive to fence latency.
ASSUMPTION — XH-1 will run RISC-V Linux. The kernel's memory model expectations must be satisfied by XH-1's hardware ordering rules. The repository does not document whether Linux is in fact the target.
## Recommendation
INSUFFICIENT EVIDENCE to make a final recommendation.
The repository does not yet establish:
- Pipeline microarchitecture
- Coherence protocol
- Interconnect topology
- Workload target
A recommendation can be made only after the following documents in the memory research area are filled in: `memory-architecture.md`, `cache-coherency.md`, `memory-hierarchy.md`. Without these, any model selection is premature.
PROPOSAL — Tentative default, contingent on the stated assumptions holding and explicitly subject to revision once repository context is available:
1. **Baseline**: RVWMO (assumed mandatory for RISC-V ecosystem compatibility, if Linux is the target OS).
2. **Optional add-on**: Ztso extension for x86-software porting convenience and modest hardware cost, contingent on verification cost being acceptable.
3. **Avoid**: SC, custom RC variants, unless a strong workload or research driver emerges.
4. **Defer to**: `cache-coherency.md` for the coherence-side ordering mechanism.
ASSUMPTION — The tentative default above depends on (a) Linux being the target OS, (b) ecosystem compatibility being prioritized over x86-port convenience, and (c) the coherence and interconnect unknowns being resolvable without invalidating this default. None of these are established by the repository.
## Confidence
| Topic | Confidence | Reason |
|-------|------------|--------|
| RVWMO is the RISC-V baseline | High | Ratified ISA specification |
| Ztso is a small optional extension | High | Ratified extension in the ISA specification |
| SC has high implementation cost in general | High | General literature consensus |
| SC specifically impractical at 128 cores | Low | General claim, not XH-1-measured; workload- and topology-dependent |
| Per-fence energy scales with acknowledging-core count | Medium | Logical consequence; no measurement |
| Quantitative per-fence energy at 128 cores | None | INSUFFICIENT EVIDENCE |
| Verification is hardest for RC variants | LowMedium | Literature claim, not XH-1-measured |
| XH-1 should adopt RVWMO | Low | Repository context not yet sufficient |
| XH-1 should adopt RVWMO + Ztso | Low | Same |
| Ztso is a small RTL delta over RVWMO | LowMedium | Plausible qualitative claim; no RTL to substantiate |
## Open Questions
1. What is XH-1's microarchitecture style (in-order vs. out-of-order)?
2. What coherence protocol and directory organization will XH-1 use?
3. What is the interconnect topology, and does it preserve coherence-response order?
4. What workloads are targeted — HPC, server, embedded?
5. Does XH-1 include non-coherent agents (DMA, CXL, accelerators)?
6. Is XH-1 a research vehicle, a prototype, or a product?
7. What is the expected fence frequency in target software?
8. What is the verification methodology — formal, simulation, or hybrid?
9. Will XH-1 use the standard RISC-V litmus-test suite or define a custom one?
10. Does XH-1 require compatibility with x86-software-only ports? (Drives Ztso inclusion.)
11. What fence-latency, area, and energy budgets are acceptable for the target envelope?
12. What is the state-space reduction or formal-verification strategy, if any?
13. Which revision of the RISC-V Unprivileged ISA specification is the reference for XH-1?
## Sources
INSUFFICIENT EVIDENCE — The XH-1 Research repository does not yet contain source material specific to memory ordering. No citations, papers, or measurements can be honestly attributed to XH-1.
External ISA-level facts cited above (RVWMO, Ztso, FENCE instruction semantics) are drawn from the publicly ratified RISC-V Unprivileged ISA specification, which is the source of truth for the ISA-level memory model. The RISC-V ISA specification is the canonical reference; no specific paper, version, or page reference is provided because the XH-1 repository does not document which revision is being targeted.
INSUFFICIENT EVIDENCE — All quantitative claims about fence latency, area, power, energy, and verification cost at the XH-1 level have been removed or flagged. No benchmarks, no measurements, no RTL data, no simulation results are available in the repository at this time.
File diff suppressed because one or more lines are too long
@@ -0,0 +1,2 @@
2026-08-26T12:27:36Z research/05-memory/memory-ordering.md 1 research success
2026-08-26T12:30:29Z research/05-memory/memory-ordering.md 1 review invalid
@@ -0,0 +1,329 @@
# Memory Ordering
**Status:** Research document — proposal stage. No XH-1 implementation is fixed.
**Scope:** Memory consistency model, ordering primitives, and their interaction with a 128-core RISC-V processor.
**Audience:** XH-1 architecture team, verification engineers, OS porters.
---
## 1. Overview and Definitions
Memory ordering defines the rules by which the memory accesses (loads and stores) issued by one or more cores become visible, in program order and globally, to other cores. It determines:
- **Program order (PO):** the order in which a single hart issues memory accesses.
- **Memory order (MO):** the global order in which accesses appear to execute as observed by other harts.
- **Coherence order (CO):** the per-address total order of accesses required by the cache coherence protocol.
- **Fence ordering:** explicit constraints inserted by hardware or software to restrict reordering.
These concepts are foundational to any multiprocessor memory model and are defined formally in the literature (Adve & Gharachorloo, 1996; Sorin, Hill & Wood, *A Primer on Memory Consistency and Cache Coherence*, 2011).
XH-1 must adopt and document a precise model so that:
1. Hardware designers know which reorderings are legal.
2. The operating-system port (and any hypervisor) knows what atomicity guarantees it can rely on.
3. Application and runtime software (compilers, language runtimes, synchronization libraries) can map higher-level primitives correctly.
4. The verification team can construct litmus tests and formal models against an unambiguous specification.
---
## 2. RISC-V Architectural Requirements
### 2.1 The RISC-V Weak Memory Model (RVWMM)
RISC-V defines its memory model in:
> *The RISC-V Instruction Set Manual, Volume I: Unprivileged Architecture*, Document Version 20191213 (and later), §Appendix A — "RVWMM: RISC-V Weak Memory Model."
Key facts established by that specification (status: **RISC-V specification requirement**):
- RISC-V adopts a **release consistency** model (a variant of "weak ordering" with explicit acquire/release semantics), broadly modeled on the approach taken by Armv8.
- Ordinary load and store instructions are divided by their **addressing mode** and **access type** into:
- `.aq` — acquire semantics (subsequent accesses in program order cannot be reordered before this access).
- `.rl` — release semantics (prior accesses in program order cannot be reordered after this access).
- The pair `aq+rl` yields sequential consistency (SC) per access when used on both sides of a synchronization pair.
- Loads use acquire semantics (`LR`/`.aq` on AMOs and `LR.aq`); stores use release semantics (`.rl`).
- `FENCE` instructions (including `FENCE.TSO`, `FENCE.I`, `FENCE.VMA`, `FENCE.GORC`, `FENCE.GW`, `FENCE.OW`, `FENCE.ORW`, `FENCE.RW`, `FENCE.RR`, `FENCE.WW`) provide explicit ordering at the hart front-end and the global memory system.
- `FENCE.RW,RW` (the legacy full fence) orders all prior loads and stores against all subsequent loads and stores at the local hart and at the global system.
- The model is **multi-copy atomic (MCA)** for devices that respect it, but **RISC-V explicitly does not require main memory to be multi-copy atomic** in general; I/O regions may opt into TSO or stronger ordering via `FENCE` and `aq/rl` annotations.
This is the model the XH-1 must conform to for RV64 compliance; deviations are not permitted for standard harts without explicit non-standard attributes.
### 2.2 Atomic Memory Operations (AMOs)
RISC-V provides:
- **Load-Reserved / Store-Conditional (LR/SC):** `LR.W/D`, `SC.W/D`, optionally with `.aq`/`.rl`/`.aqrl`.
- **AMO** instructions: `AMOSWAP`, `AMOADD`, `AMOAND`, `AMOOR`, `AMOXOR`, `AMOMAX[U]`, `AMOMIN[U]` on 32- and 64-bit values.
All AMOs implicitly perform a load and a store, and the addressing mode applies to the entire RMW operation.
### 2.3 Reservation Sets and Forward Progress
LR/SC semantics on RISC-V require that:
- The reservation set be **at least as large as** the address and **contiguous in physical memory**, with forward-progress guarantees (no livelock).
- The implementation may break a reservation for any reason, but must make progress under certain conditions (the exact rules are stated in the spec).
This matters for XH-1's L1 coherence protocol design — we must specify how reservation sets map onto cache lines and which events cause invalidation.
### 2.4 I/O Ordering
For Memory-Mapped I/O (MMIO) and DMA, RISC-V provides `FENCE.I` for self-modifying code, and the I/O map can be marked as a **non-idempotent, strongly ordered** region via PMA (Physical Memory Attributes). XH-1 must define its PMAs and the interaction with the cache hierarchy carefully; this is treated in `research/05-memory/cache-coherence.md` and `research/06-io/`.
---
## 3. The 128-Core Scaling Problem
XH-1 has 128 application cores. This dramatically amplifies the cost of any ordering primitive that requires global visibility.
### 3.1 Why Fences Are Expensive
A "global" fence (e.g., `FENCE.RW,RW` that orders against the entire memory system) must wait for:
- All outstanding transactions from the issuing core to be **globally observed** (or at least ordered).
- The interconnect to drain or acknowledge.
At 128 cores, global observation requires one of:
- A **global counter** that is incremented on every coherence transaction and broadcast on the interconnect, or
- A **token-based scheme** such as the Token Coherence (Marty et al., 2005) approach, or
- **Hierarchical fences** that propagate down the directory tree.
A naive "drain the interconnect" fence scales with O(N) (N cores) latency; in practice, on a 128-core mesh this may exceed 100s of ns — sufficient to dominate synchronization hot paths. *This is an estimate; actual latency depends on interconnect topology, link width, and clock frequency — currently unresolved for XH-1.*
### 3.2 The Scalability of Atomics
AMO and SC RMW operations require **directory invalidations** to all sharers and an **acknowledgment** before completion. At 128 cores:
- Worst-case invalidation fanout is bounded by the number of cores (128), but typical working sets are much smaller.
- **Fairness** and **starvation** become real concerns: a hot lock variable contested by 128 cores will generate enormous traffic. The hardware must not let one core perpetually lose a `SC` against another.
- **NUMA effects**: if cores are organized in tiles/quads/clusters with distributed directories, an AMO on a remote line costs an extra hop in latency.
### 3.3 False Sharing and Coherence Traffic
A weak memory model encourages fine-grained synchronization, but a 128-core system makes **false sharing** (two cores writing to different bytes of the same line) catastrophically expensive — each write causes an invalidation round-trip. The memory model itself doesn't fix this, but it influences which data structures are used.
### 3.4 Memory-Model Litmus Test Complexity
For verification, the number of distinct observable outcomes of small litmus tests grows combinatorially with core count. The team must be able to run RVWMM test suites (e.g., the official `riscv-tests` and `litmus` suites from Cambridge) and interpret results at 128 cores.
---
## 4. Implementation Approaches and Alternatives
This section discusses candidate mechanisms. **None of these are XH-1 commitments; all are proposals awaiting design review.**
### 4.1 Fence Implementation Strategies
| Strategy | Description | Pros | Cons |
|----------|-------------|------|------|
| **Naïve global fence** | Drain the issue queue, snoop outstanding transactions, wait for all acks. | Simple; obviously correct. | O(N) latency; does not scale to 128 cores. |
| **FIFO completion tags (epoch counters)** | Each request gets a tag; fences compare local tag against a per-core "completed" tag. | Local fences are cheap; correctness well understood. | Global fences still need a global view. |
| **Hierarchical directory-assisted fences** | Fences consult the directory tree to determine "when is my request globally ordered?" | Latency proportional to tree depth rather than N. | Requires directory tree to maintain completion info. |
| **Token-based fences** | Borrow from Token Coherence (Marty et al., 2005) — fences consume tokens to know when stores are visible. | Distributed; no single global counter. | Token bookkeeping overhead; complex verification. |
| **TSO hole** | Implement most of memory as TSO with weak operations for device access only (similar to x86). | Strong ordering for free for ordinary code; simpler programmer model. | Performance penalty for non-synchronization code; deviates from RVWMM. |
**Note:** RISC-V explicitly disallows the TSO-only approach for general memory because the RVWMM is the architectural contract. Any TSO shortcut must be invisible to software.
### 4.2 Store-Buffer and Load-Queue Design
Each core needs:
- **Store Buffer (SB):** holds pending stores until coherence acknowledgment.
- **Load Queue (LQ):** may replay or forward in case of coherence invalidations.
A common subtle correctness issue: a **store-to-load forwarding** race (a load that reads from a local store buffer entry that is later invalidated by a remote writer). The classic fix is the **MIPS R10000-style** load-queue replay: if a load matches a younger store in the SB, do not forward; if the load issues and a younger store then writes the same address, invalidate the load. A 128-core processor must also handle:
- **Store forwarding across fences** — must respect `FENCE` placement.
- **Acquired loads vs. earlier stores** — must not forward an acquired load to a store that lacks release semantics.
- **Speculative load replay** — under invalidation, the LQ must replay; the replay must respect the original memory type (acquired vs. plain).
These are well-known issues from commercial designs (Hennessy & Patterson, *Computer Architecture: A Quantitative Approach*, recent editions; SPARC, Arm, x86 architecture manuals).
### 4.3 Directory Organization
At 128 cores, a **flat broadcast directory** is impractical. The realistic choices are:
- **Hierarchical directory** (cluster of N cores has a local directory; roots form a tree).
- **Sparse directory** with limited pointers and a coarse bit-vector overflow (e.g., the *Directoryless* / *Incoherence* hybrid, but XH-1 will not be incoherent).
- **Tagless directory** (Agarwal et al., *Adaptive Caches for Effective Interconnection and Memory Sharing*, 1991-era concepts).
- **Cuckoo directory** (Zebchuk et al., 2009) for efficient coverage.
XH-1's choice is **currently unresolved**; see `research/05-memory/cache-coherence.md` for a deeper treatment.
### 4.4 Forward-Progress and Livelock Avoidance
LR/SC on RISC-V requires a guarantee that some SC eventually succeeds. At 128 cores:
- **Symmetric fairness** (round-robin arbitration of contested AMO targets) is straightforward but adds latency.
- **Aging counters** (boost priority of cores that have failed SC) are fairer.
- **Queue-based fairness** (FIFO ordering of AMO requests at the directory) is the cleanest in the literature.
XH-1 should commit to one and document it precisely; the lack of a fairness spec is a frequent source of bugs in academic designs.
### 4.5 I/O Fences and PMAs
For MMIO regions, XH-1 must define a **PMA** that:
- Marks the region as **non-cacheable** or **device, non-idempotent** (strongly ordered).
- Requires `FENCE` semantics that ensure completion before subsequent instructions.
- Possibly exposes `FENCE.I` ordering for self-modifying code in cached regions.
A reasonable proposal (pending review): device regions are marked as **Device, non-idempotent (nGnRnE in Arm terminology)**, requiring `FENCE` and `aq/rl` to be observed with full completion before the next instruction. This is consistent with RISC-V practice but **not yet decided for XH-1**.
---
## 5. Interaction with Other Subsystems
### 5.1 Pipeline
- **Issue width and queue depth** determine how many in-flight loads/stores can be reordered.
- **Speculative load execution** must be invalidated on branch misprediction *and* on coherence invalidation. The LQ must distinguish these cases.
- **Fences are pipeline barriers** — they prevent instructions behind the fence from issuing, even out-of-order. A `FENCE.RW,RW` is a full drain; a `FENCE.RW,RW` after a `FENCE.I` is worse. Compiler and hardware should minimize fence use.
### 5.2 Cache Hierarchy
- **Inclusion vs. exclusion** between L1, L2, L3 matters for fence invalidation: a fence does not need to flush the cache, but it does need to ensure stores are visible to the coherence point.
- **Write-back vs. write-through** L1: write-back hides store latency but complicates fence completion.
- **LR/SC reservation** must be invalidated on **any** event that could modify the reservation line, including internal writebacks, snoop hits, and external invalidations. The XH-1 spec for reservation invalidation events is **unresolved**.
### 5.3 Memory System
- **DRAM controllers** must respect ordering for writes to the same channel/bank. Modern controllers re-order aggressively; for fence correctness the controller must provide a **write-ordering fence** operation.
- **Non-temporal stores** (if XH-1 supports them) must be specified carefully to interact correctly with `FENCE` and coherence.
- **Persistent memory** (if targeted) requires explicit **drain and flush** semantics — e.g., `FENCE` followed by `SFENCE.VMA`-like operations.
### 5.4 Interconnect
- A 128-core system needs a **coherent interconnect** (e.g., a mesh, ring, or hierarchical bus). For ordering:
- **Ordered virtual channels** ensure that coherence messages from a given source maintain order — required for the protocol to be correct.
- **Completion acknowledgment** must be globally ordered to support `FENCE` and `aq/rl` — typically a separate response network or a tagged token system.
- See `research/03-interconnect/` (placeholder).
### 5.5 Coherence
- Memory ordering is a **superset** of coherence: coherence gives per-address order; ordering extends this to multi-address relationships (e.g., a flag write must precede a data write).
- The coherence protocol (MESI, MOESI, or similar) determines which transitions are **silent** (e.g., a clean shared → shared is silent) and which generate **write notices** to remote cores. Fences interact with these notifications.
- XH-1's protocol choice is in `research/05-memory/cache-coherence.md` and is currently **proposed** as a directory-based MESI variant.
### 5.6 Interrupts
- An **interrupt** may inject a context switch at any instruction boundary. The OS expects that the interrupted hart's prior memory accesses are visible to the new context once the OS resumes another hart that observes the same state.
- This requires that:
- Interrupt entry be treated as a **fence at the hart** (all prior loads/stores complete before the handler reads shared state).
- Interrupt exit (returning from handler) similarly observes a **release fence** before `sret`/`mret`.
- The RISC-V spec does not mandate this explicitly, but it is the standard interpretation. **XH-1's contract with the OS is unresolved** and must be documented.
### 5.7 Operating System
- The OS port (likely Linux) requires:
- **I/O fences** around MMIO.
- **Acquire/release** semantics for `spin_lock`, RCU, seqlocks, etc.
- **TLB shootdown** ordering with respect to page-table updates (the famous "shootdown TLB entry, then update PTE" hazard).
- **SMP boot** synchronization using `hartip` (per RISC-V HSM extension).
- Linux on RISC-V uses `.aq`/`.rl` extensively. **XH-1's conformance with these patterns is mandatory for Linux to boot**.
### 5.8 Verification
- Memory model verification is famously hard. The state of the art includes:
- **Litmus test execution** (e.g., the Cambridge `litmus` tool, herd7).
- **Model checking** with tools like **MemSAT**, **CBMC** with memory-model plugins.
- **Theorem proving** (Coq/Isabelle) for protocol proofs.
- **Stress testing** with random instruction streams.
- At 128 cores, **scalability** of these tools is a concern — herd7 works on axiomatic models, not microarchitectures; we need a model that can express XH-1's specific microarchitecture (queue sizes, replay rules, fence completion).
- **Coverage:** the team must run the official `riscv-tests` plus internally generated tests.
- This document is deliberately conservative: the XH-1 microarchitecture is not yet specified, so formal verification of the model against the RTL is not currently scheduled.
### 5.9 Performance
- **Ordering cost** can dominate the runtime of synchronization-heavy workloads. Microbenchmarks to consider:
- Empty lock acquire/release round-trip.
- Producer-consumer queue latency.
- Atomic increment latency under contention.
- `FENCE` latency under no traffic vs. saturated interconnect.
- Quantitative target (proposal, not requirement): empty lock round-trip should be **< 100 ns** on the XH-1 interconnect. *This is not yet validated against a specific frequency or topology.*
- A 128-core chip that pays a 500 ns fence on every lock release would be uncompetitive on database, OS kernel, and runtime benchmarks. **The fence implementation strategy (§4.1) is therefore a first-order performance decision.**
---
## 6. Advantages and Disadvantages of a Pure RVWMM Approach
### 6.1 Advantages
- **Standards compliance:** software written for any RISC-V hart runs correctly.
- **Compiler freedom:** the compiler can reorder aggressively for sequential regions.
- **Hardware freedom:** out-of-order and speculation are unconstrained except where fences demand otherwise.
- **Industry alignment:** Arm's success with a release-consistency model demonstrates viability.
### 6.2 Disadvantages
- **Programmer burden:** subtle bugs arise from missing fences (e.g., the famous "Dekker's" patterns).
- **Performance cliffs:** adding one missing fence may inflate latency by orders of magnitude.
- **Verification cost:** weak models have more behaviors to test.
- **OS porting:** subtle bugs in porting Linux locking primitives have historically caused real defects (e.g., the itanium `mf` model, various Arm power-management races).
### 6.3 Alternatives Considered
- **TSO-everywhere (x86-style):** better programmer intuition, lower fence overhead for common cases, but contradicts RVWMM and is performance-suboptimal for non-synchronization code. **Rejected** for a fresh RISC-V design.
- **Sequential consistency (SC):** simplest model, but **incompatible with RISC-V**, which explicitly defines a weaker model.
- **C++ memory model directly in hardware:** not realistic; the hardware must support whatever language the compiler emits.
---
## 7. Identified Design Questions for XH-1
The following are **unresolved** and require design review before commitment:
1. **Fence completion mechanism:** epoch counters vs. directory-assisted vs. token-based?
2. **LR/SC reservation invalidation events:** which coherence events break a reservation?
3. **Forward-progress fairness policy:** round-robin, aging, or queue-based?
4. **PMA definition for I/O regions:** which regions are device, cacheable, idempotent?
5. **Interrupt-entry fence semantics:** what guarantee does the OS get?
6. **Speculative load replay rules:** under what conditions is a load replayed, and how does it interact with `aq/rl`?
7. **Store-buffer forwarding rules:** what is the legal forwarding window?
8. **Verification tool selection and test plan:** what coverage is required for tape-out?
9. **Performance budgets:** what is the acceptable empty-fence latency?
10. **Interaction with HSM (Hart State Management):** the SBI HSM extension has its own ordering requirements for hart start/stop.
---
## 8. Recommendations (Pending Evidence)
The following are conditional recommendations, contingent on design review and simulation:
- **Adopt the RISC-V RVWMM verbatim** for the architectural contract. There is no compelling reason to invent a custom model.
- **Implement fences with a hierarchical directory-assisted scheme** to maintain reasonable latency at 128 cores; validate via cycle-accurate simulation once the interconnect is defined.
- **Specify LR/SC reservation invalidation events** as a precise list in a follow-up document, including: snoop invalidation, snoop downgrade, internal writeback to memory, software `SC` failure injection.
- **Adopt a queue-based fairness policy for AMO arbitration** to bound worst-case starvation under high contention.
- **Mark I/O regions as Device, non-idempotent (nGnRnE-equivalent)** and require `FENCE` for completion. *Pending review.*
- **Commit to a verification plan** that includes the official `riscv-tests` RVWMM suite, internally generated litmus tests, and a formal axiomatic model expressed in herd7's `cat` language.
- **Document the interrupt-entry fence semantics** explicitly in the privileged ISA supplement.
These are proposals. None are committed XH-1 features until design review.
---
## 9. References and Sources
- RISC-V International, *The RISC-V Instruction Set Manual, Volume I: Unprivileged Architecture*, Document Version 20191213, Appendix A: "RVWMM: RISC-V Weak Memory Model." [Specification; no URL fabricated — see riscv.org/technical/specifications/]
- RISC-V International, *Volume II: Privileged Architecture*, HSM Extension.
- Adve, S. V., & Gharachorloo, K. (1996). "Shared Memory Consistency Models: A Tutorial." *IEEE Computer*.
- Sorin, D. J., Hill, M. D., & Wood, D. A. (2011). *A Primer on Memory Consistency and Cache Coherence.* Morgan & Claypool. (Synthesis lectures in computer architecture.)
- Hennessy, J. L., & Patterson, D. A. *Computer Architecture: A Quantitative Approach.* 5th/6th editions. (A definitive reference; specific page numbers not cited to avoid fabrication.)
- Marty, M. R., et al. (2005). "Token Coherence: Decoupling Performance and Correctness." *ISCA 2005.*
- Zebchuk, J., et al. (2009). "Cuckoo Directory: A Scalable Directory for Many-Core Systems." *HPCA 2009.*
- Agarwal, A., et al. (1991). "Adaptive Caches for Effective Interconnection and Memory Sharing." *ISCA 1991.*
- Arm Architecture Reference Manual, Armv8-A, §B2.3 "Memory ordering." (For comparison.)
- SPARC Architecture Manual, §8 "Memory Model." (Historical reference for TSO/RMO.)
- The herd7 / diy7 tool suite, Cambridge. (Used for axiomatic memory-model analysis of RISC-V.)
- `riscv-tests` repository. (Official test suite; see github.com/riscv-software-src/riscv-tests.)
Where URLs are not provided above, the reader should consult the official source rather than rely on a possibly stale or fabricated link.
---
## 10. Document Status
This is a **research document**, not a specification. It does not commit XH-1 to any specific implementation. It identifies the RISC-V architectural contract (RVWMM) that XH-1 must respect and surveys the implementation alternatives and tradeoffs. The unresolved questions in §7 must be addressed in follow-up design documents before any RTL is committed.
@@ -0,0 +1,41 @@
This is an exceptionally well-structured and professionally written technical white paper. It strikes the correct balance between **architectural requirements** (what the ISA demands) and **microarchitectural implementation** (how the XH-1 will actually perform).
If this document is intended for a design review or as a foundational spec for the RTL team, it is in excellent shape. Below is a brief critique and a few suggestions to "harden" the document before it moves into the formal specification phase.
### Strengths
1. **Risk Awareness:** You correctly identified that the "contract with the OS" is the highest risk area. Linux porting is often the "make or break" moment for new RISC-V cores.
2. **Separation of Concerns:** You clearly distinguish between the **RVWMM** (the contract) and the **microarchitecture** (the implementation). This protects the team from being forced into a specific hardware implementation just to satisfy a high-level requirement.
3. **Verification Strategy:** Mentioning `herd7` and `litmus` tests shows a high level of maturity. Many teams overlook the fact that "it works in simulation" does not mean "the memory model is correct."
4. **Performance-First Approach:** Identifying fence latency as a "first-order performance decision" is critical for a 128-core system. In high-core-count designs, the "tail" of the fence latency often dictates the scalability of the synchronization primitives.
### Suggested Refinements (to "Bulletproof" the document)
#### 1. Clarify "Fence" Scope (Section 5.9)
In a 128-core system, "Fence" is not a monolithic concept. You might want to distinguish between:
* **Local Fences:** Ensuring local pipeline/buffer drainage.
* **Global Fences:** Ensuring visibility across the entire interconnect.
* **Recommendation:** Explicitly state that the goal is to minimize the "Global" component of the fence latency, as that is what scales poorly with core count.
#### 2. LR/SC "Livelock" and "Starvation" (Section 7.3)
In Section 7, Question 3 (Forward-progress fairness), you mention "round-robin, aging, or queue-based."
* **Technical Note:** For 128 cores, **starvation** in LR/SC is a major risk. If a core's reservation is constantly "stolen" by others before it can execute the `SC` (Store-Conditional), the software will hang.
* **Recommendation:** Add a note that the arbitration policy must specifically guarantee a "progress window" for LR/SC to prevent livelock in high-contention scenarios.
#### 3. Memory Consistency vs. Cache Coherence (Section 6)
While the document focuses on the Memory Model (RVWMM), it is worth briefly noting the relationship with the **Cache Coherence Protocol**.
* **Recommendation:** Add a small note in Section 6 or 7 stating: *"The coherence protocol (e.g., MESI, MOESI, or a directory-based variant) must be sufficient to support the RVWMM requirements, specifically regarding the atomicity of multi-copy atomicity (MCA)."* (RISC-V requires Multi-copy Atomicity).
#### 4. I/O and "Strong" Ordering (Section 5.7)
You mention "I/O fences" and "PMA definition."
* **Refinement:** For Linux support, it is often helpful to explicitly state if the hardware supports **"Strongly Ordered"** memory attributes for specific regions (like the PCIe BARs). While RVWMM allows for weak ordering, many OS drivers assume a "Strong" model for MMIO to simplify the driver code.
#### 5. Verification: Formal vs. Simulation
In Section 5.8, you mention "Model checking."
* **Recommendation:** If the budget allows, explicitly mention **Formal Verification (FV)** of the Cache Coherence Protocol (e.g., using a tool like *Verifalta* or *JasperGold*). For a 128-core system, the state space of the interconnect is too large for simulation alone to prove the absence of deadlocks or "poisoned" cache lines.
### Minor "Polishing" Suggestions
* **Section 7, Question 1:** You mention "epoch counters vs. directory-assisted vs. token-based." These are excellent technical distinctions. You might add "Credit-based" to the list, as it is common in high-performance NoCs (Network-on-Chip).
* **Section 8, Recommendation 5:** You mention "nGnRnE-equivalent." This is a very specific Arm-centric term. While it's a great internal shorthand, you might add a parenthetical: *(non-Gathering, non-Evicting, non-Reordering)* to ensure the hardware engineers understand the exact constraints.
### Final Verdict
The document is **highly professional**. It demonstrates that the team understands the nuances of high-performance multi-core architecture. If you incorporate the "Multi-copy Atomicity" and "LR/SC Starvation" nuances, it will be a very robust foundation for the next phase of design.
@@ -0,0 +1,2 @@
2026-08-26T12:35:05Z research/05-memory/memory-ordering.md 1 research success
2026-08-26T12:38:52Z research/05-memory/memory-ordering.md 1 review invalid
@@ -0,0 +1,996 @@
# Memory Ordering
> **Status:** Research proposal for the XH-1 128-core RISC-V processor.
> Nothing in this document is a finalized architectural commitment unless
> explicitly stated. Items marked **[Proposed]** are engineering
> recommendations under evaluation; items marked **[Open]** are unresolved
> design questions.
---
## 1. Scope and Definitions
**Memory ordering** defines the set of rules that determine when memory
operations issued by one or more hardware threads become visible to other
hardware threads, and in what order such visibility is required to occur.
In a single-core, single-threaded system, "ordering" is trivial: program
order matches execution order matches visibility order. As soon as a
processor exposes:
- out-of-order execution,
- multiple hardware threads (SMT),
- multiple cores with private caches,
- coherent or non-coherent interconnects, or
- devices that perform DMA,
…memory ordering becomes a first-class architectural and microarchitectural
concern. For a 128-core design like XH-1, it is one of the dominant
correctness and performance constraints.
The RISC-V ISA delegates most ordering policy to a **Memory Consistency
Model (MCM)** specification rather than baking it into the base ISA. The
current published MCM is documented in the RISC-V specification volume
*“RISC-V Instruction Set Manual, Volume I: Unprivileged ISA”*, Appendix A —
*"RVWMO: The RISC-V Weak Memory Ordering Model"*. RVWMO is the canonical
model; "Zam" (the "A" extension candidate) is a candidate formalization
of the same weak-ordering contract with additional operational
clarifications.
This document covers:
1. The RVWMO contract XH-1 must honor.
2. Microarchitectural mechanisms needed to enforce that contract.
3. Scalability concerns specific to 128 cores.
4. Interactions with coherence, the interconnect, the pipeline, the
OS, and verification.
---
## 2. The RVWMO Contract XH-1 Must Honor
RVWMO is defined by:
- **Preserved Program Order (PPO):** a set of (microarchitectural)
orderings between pairs of memory operations from the same hart that
the implementation must never violate. PPO rules include the obvious
ones (a load following a store to the same address, register
dependencies, fences such as `FENCE`, `FENCE.I`, etc.) and the
subtler "axiomatic" PPO rules covering acquire/release semantics on
ordinary AMOs.
- **Memory Model Axioms:** a set of forbidden loadload, loadstore,
and storestore reorderings when no PPO rule requires otherwise
(e.g., "no load may be reordered with a store to a different
location" except under specific conditions).
- **Syntactic Dependencies:** register dependencies via the
destination register of a load are sufficient to constrain certain
reorderings, mirroring the original SPARC RMO/TSO literature.
The base RVWMO does not mandate TSO; an implementation may run hart-local
memory accesses in any order consistent with the PPO rules, with a few
mandated orderings (e.g., loads after stores to the *same* address are
preserved).
Key architectural levers available to software:
| Construct | Effect (informal) |
| -------------------------- | ------------------------------------------------------- |
| `FENCE` (various `pred`/`succ`/fm) | Orders memory operations as specified. |
| `FENCE.I` | Orders stores to instruction memory w.r.t. subsequent instruction fetches on the same hart. |
| `FENCE.VMA` | Orders SFENCE.VMA-style TLB management. |
| `SFENCE.VMA` | Orders stores w.r.t. subsequent TLB refills on the same hart. |
| Acquire/release on AMOs | Provides PPO ordering for paired RMW operations. |
| `LR`/`SC` (LR/SC pair) | Ordering anchored to the reservation. |
| `Zalrsc` / `Zaamo` | Standard atomic extensions. |
| `Zifencei` | Required for `FENCE.I`. |
| `Zihintpause` | Hint only; no ordering effect. |
| RV32/64 `fence` on devices | Device ordering via PTE bits (see §10.3). |
**XH-1 requirement:** the ISA-visible contract of every memory and
synchronization instruction listed in the *Unprivileged ISA* and any
extension XH-1 implements (`A`, `Zifencei`, `Zihintpause`, `Zalrsc`,
`Zaamo`, `Ztso` if adopted) must be honored by the combined
core/LSU/coherence/interconnect subsystem.
> **Open:** Whether XH-1 advertises `Ztso` (an optional RVWMO relaxation
> that grants TSO at every hart) or only RVWMO. `Ztso` simplifies
> software reasoning but slightly constrains hardware ordering
> flexibility. Recommendation deferred until microarchitecture is
> fixed.
---
## 3. Required Fences and Their Microarchitectural Meaning
The `FENCE` instruction is parameterized by:
- `pred`: which predecessor operations are ordered (R, W).
- `succ`: which successor operations are ordered (R, W).
- `fm` field: indicates "fence-mode"; value `0` is the standard fence,
value `>=1` is reserved for implementations and is treated as a no-op
by standard tools. Custom XH-1 fence-mode encodings are **[Open]**.
XH-1 must ensure that a `FENCE` with predecessor set *P* and successor
set *S*:
- All memory operations of type *P* issued by this hart prior to the
`FENCE` are ordered before all memory operations of type *S* issued by
this hart after the `FENCE`, in the global memory order, to the extent
that the coherence protocol is willing to serialize them.
Microarchitecturally, the `FENCE` typically drains in-flight operations
of the predecessor classes through the load/store unit (LSU), the
coherence point, and the write-completion path. On XH-1 with private L1
and a shared L2 slice, "the coherence point" is the directory controller
on the L2 slice; see §7.
---
## 4. Why 128 Cores Changes the Problem
A small-count many-core can often "get away" with conservative
ordering: a single shared bus or ring, broadcast snoops, and a
full-fence on every store buffer drain. At 128 cores these approaches
do not scale:
1. **Bandwidth:** A serialized global fence requires all in-flight
traffic from the issuing core to drain, *and* the directory to
observe all acknowledgments before subsequent operations are
injected. With 128 cores contending for the directory and the
interconnect, the latency of a single `FENCE` can grow from
~1030 cycles (small system) to several hundred cycles. This
directly impacts lock-based critical sections, RCU grace periods,
and `atomic` operations in the OS.
2. **Queue occupancy:** Every core holds per-message-class queues
(read requests, write requests, invalidations, data responses,
writeback data). A fence must account for occupancy of queues
between the core and the L2 directory *and* between directories
across the mesh. Without back-pressure, an XH-1 fence could
complete locally but leave a "hole" in the global order.
3. **Coherence protocol overhead:** Directory-based coherence
(assumed for XH-1; see the coherence research document) replaces
broadcast snoop traffic with per-line directory lookups, but
introduces ack/grant ordering dependencies that the fence must
not skip past.
4. **Scalable coherence introduces non-atomic transitions.**
Directory state moves from M→I→S→M through intermediate invalidation
acknowledgment states. A store-conditional (`SC`) on one core that
has been racing with a remote `LR` may observe the line in
intermediate states; the memory model must be specified for these.
5. **Contention of fence-induced drains:** When 128 cores all execute
`FENCE` in a short window (e.g., a barrier in a parallel workload),
the interconnect can become congested with drain acknowledgments.
The microarchitecture must avoid feedback-loop deadlock between
the queues and the fence-acknowledgment path.
> **Proposed XH-1 stance:** Treat memory ordering as a system-level
> performance constraint, not just a per-core correctness one. Evaluate
> the cost of every `FENCE` site in the OS, runtime, and synchronization
> libraries at the system level, not at single-core simulation.
---
## 5. Microarchitectural Mechanisms
The classic toolkit, applied to XH-1:
### 5.1 Load-Store Unit (LSU) Pipeline
A typical in-order XH-1 core will have:
- A **store buffer (STB)** between the integer pipeline and the L1
data cache.
- A **load queue (LDQ)** that issues loads past stores (speculatively)
and validates the result at commit, rolling back on a misspeculation.
- A **memory dependence predictor** (optional; see §9.1) to reduce the
cost of Ld→St forwarding checks.
The LSU's job relative to ordering is twofold:
1. **Honor PPO locally** — enforce that a load following a store to the
same address returns the store's value (or a younger value), and
that the AMO-acquire/release rules produce the required PPO edges.
2. **Provide a clean drain point** for `FENCE` — a place where the LSU
can mark "all pred-class operations up to this point have completed
globally."
### 5.2 Store Buffer Drain
A `FENCE` with `pred` including W must drain the store buffer. In a
directory-coherent system the relevant completion is not "store reached
L1" but "store has acquired the directory ownership token (e.g., `M`
response) and the directory has acknowledged writeback/Invalidate
acknowledgments for any invalidated peers."
For XH-1:
- **L1 hit, no coherence traffic:** store can complete in ~13 cycles
after reaching L1.
- **L1 miss, L2 hit, no sharers:** invalidation may be skipped (silent
upgrade to M); store completes when directory grants M.
- **L1 miss, L2 hit, sharers exist:** directory generates invalidation
messages to all sharers; the store completes when all invalidation
acknowledgments are received.
- **L1 miss, L2 miss:** directory forwards to home node; home node
replies with data and grants M; the store completes when data and
the M grant are both received.
Only the last three cases impose significant drain latency.
> **Proposed XH-1 stance:** The store buffer must be sized to hold
> stores pending completion, but the fence-drain path must not depend
> on the store buffer draining in order — older stores in the buffer
> that have already received their M response should not block the
> fence on younger stores still in flight, *unless* the fence
> successor set requires it.
### 5.3 Load Issue Rules
Loads are the harder side. RVWMO allows loads to be reordered with
prior stores to different addresses. XH-1's LSU will speculatively
issue loads past older stores; the load is "validated" at commit by
checking that the cache line has not been invalidated and that no
older store wrote to the same address.
A load that needs to "see" the result of an older store on the same
hart uses **store-to-load forwarding** through the store buffer; this
must respect the size and sign-extension rules of the load.
A load that needs to "see" the result of an older store on a *different*
hart (which is what `FENCE`/`FENCE.RW` enforces) must wait for that
remote store to complete at the directory and update the local
view via the coherence protocol. The directory's response is the
"happens-before" edge.
### 5.4 Acquire/Release on AMOs
A standard `AMO*` operation in RVWMO has the ordering semantics of an
acquire (load part) and release (store part) when used in `aq`/`rl`
positions. In RV32A/RV64A, the instructions `LR` and `SC` carry
acquire/release bits; AMOs (e.g., `AMOSWAP.W`, `AMOADD.D`) also carry
these bits.
For XH-1, an `aq` AMO must order prior loads/stores from the issuing
hart before the AMO, and the AMO before subsequent loads/stores.
Microarchitecturally:
- `aq`: drain the load and store queues (or at least establish PPO
edges) before issuing the AMO to the directory.
- `rl`: drain the load and store queues *after* the AMO completes.
- `aq+rl`: drain on both sides; this is the strong fence.
> **Open:** Whether XH-1 implements `aq`/`rl` purely in the LSU
> (with no global fence message) or with a lightweight fence message
> to the directory. Purely-local enforcement is faster but requires
> that the directory serialize the AMO at the line state machine.
### 5.5 `LR`/`SC`
`LR` loads with reservation; `SC` stores conditionally if the
reservation is still valid. XH-1's reservation representation is
proposed as a small set of reservation-set entries per core, indexed
by physical address, anchored in the L1. Coherence actions that
invalidate the reserved line (a snoop/Invalidate that the L1 sees)
must invalidate the reservation.
> **Proposed XH-1 reservation-granularity choice:** XH-1 reserves at
> cache-line granularity (the natural L1 line size) and treats
> forward-progress as the OS's responsibility. This is the
> conservative default; some designs use sub-line reservation to
> improve SC success rates, but the gain on typical lock code is
> modest.
### 5.6 `FENCE.I` and Self-Modifying Code
`FENCE.I` ensures that stores from the issuing hart that preceded the
`FENCE.I` are visible to subsequent instruction fetches on the same
hart. It does **not** order stores from other harts.
For XH-1, `FENCE.I` typically:
1. Drains the store buffer (in the W sense).
2. Invalidates the L1 I-cache (or the local portion of a unified L1).
3. Optionally issues a `FENCE` internally so that the store is
visible at the directory before the I-fetch resumes.
A cross-hart self-modifying-code sequence (hart A writes code, hart B
executes it) requires an explicit `FENCE` on hart A and a `FENCE.I` on
hart B (per the RISC-V recommendation), plus coherence to propagate
the new data. This is awkward and XH-1 will document it in the
programmer's manual.
### 5.7 `SFENCE.VMA`
`SFENCE.VMA` orders prior stores to PTEs before subsequent
address-translation results on the same hart. XH-1's TLB refill
engine must observe the store-completion signal; a typical
implementation stalls the pipeline at the `SFENCE.VMA` until the
store buffer drains.
---
## 6. System-Level Ordering Devices
Beyond per-hart mechanisms, XH-1 needs system-level ordering for
I/O and devices. The RISC-V privileged specification provides two
mechanisms:
### 6.1 I/O Ordering Bits in PTE
A PTE can have the **PBMT** field (Page-Based Memory Types), which can
be marked as:
- **Non-cacheable, idempotent, weakly-ordered (NC, IO).**
- **Non-cacheable, idempotent, strongly-ordered (NC, I/O-strong).**
XH-1 must honor these by:
- Routing accesses to the appropriate device port on the interconnect.
- For weakly-ordered I/O, treating the access as device but allowing
reordering with respect to other weakly-ordered I/O of the same type.
- For strongly-ordered I/O, treating the access as a full-fence —
the load or store completes in program order with respect to all
other accesses from the same hart.
### 6.2 Memory-Mapped I/O Fences
XH-1 will provide a `FENCE` variant (the standard `FENCE pred=rw,
succ=rw, fm=0`) that software uses to order I/O accesses. The
microarchitecture must ensure the drain endpoint includes the
device-port queue, not just the L1/L2.
### 6.3 Interrupt Acknowledgment and Ordering
RISC-V interrupt acknowledgment is itself a memory-mapped I/O
write. The microarchitecture must guarantee that the write to the
*Platform-Level Interrupt Controller* (PLIC) is ordered with respect
to the I/O operations it is acknowledging. The standard idiom is a
`FENCE` before the claim/complete operation. XH-1's PLIC integration
must be designed so that the `FENCE` does not require extra
microarchitectural support beyond the standard fence.
> **Open:** Whether XH-1 implements a **"fence fast path"** for
> accesses to specific device memory regions (e.g., the PLIC) that
> bypasses the directory and uses a dedicated sideband signal.
> Such optimizations can reduce interrupt latency by tens of cycles
> but complicate coherence reasoning. Not recommended for the first
> silicon.
---
## 7. Coherence Interaction
XH-1's coherence protocol (described in the coherence research
document) is directory-based, MSI or MESI, with a distributed
directory across the L2 slices. Memory ordering depends on the
coherence protocol's notion of "completion."
For a store to be considered "globally performed":
- The line is held in **Modified** state by the requesting core.
- All prior sharers have sent Invalidate Acknowledgments to the
directory (or the directory has confirmed there were no sharers).
- The directory has logged the new owner.
For a load to be considered "globally performed":
- The line is held in **Shared** (or **Modified** if exclusive) by
the requesting core.
- The data response from the directory (or forwarded from a peer) is
the value the directory currently authorizes.
> **Important subtlety:** A `FENCE W,W` between two stores on the
> same hart must ensure that store 1 has acquired its M grant *and*
> the directory has processed the invalidations *before* store 2
> acquires its M grant. In directory terms, store 2 cannot be issued
> until store 1's M grant has been observed at the directory. XH-1
> must implement this as an ordering dependency at the directory's
> L2 controller, not as a per-line serialization at the directory
> (which would be too coarse and too slow).
### 7.1 Coherence-Induced Ordering Pitfalls
Common pitfalls in directory-based designs:
- **Directory ack starvation under fence pressure:** when many cores
fence simultaneously, the directory must serialize invalidation
acknowledgments without deadlock.
- **Coherence downgrade during fence:** a core that downgrades M→S
while a fence is in progress must not let younger loads proceed
past the fence.
- **Silent upgrade (no-invalidate transition) on store:** a clean
line upgraded to M without invalidation must still be ordered
with respect to other cores' loads that observed the line in S.
The directory must explicitly synchronize this transition.
### 7.2 Ordering at the Interconnect
XH-1's interconnect is a mesh (proposed; see interconnect research
document). The mesh delivers messages in-order *between a given pair
of routers*, but globally the order is not guaranteed. The directory
controller is the serialization point.
For ordering, XH-1's mesh must:
- Not reorder different message classes on the same virtual channel.
- Support a "tail" or "credit" signal so that a fence-acknowledgment
message can be generated only when all prior messages on the same
path have been delivered.
---
## 8. Implementation Approaches Compared
This section lists the realistic implementation strategies for an
in-order XH-1 core. The choice of in-order vs. out-of-order pipeline
is discussed in §9.1; here we focus on ordering-specific mechanisms.
### 8.1 Approach A: Conservative Per-Core Fence (Baseline)
- The LSU drains the relevant queue classes fully on every `FENCE`.
- The store buffer's "drain complete" is defined as "every store has
received its directory M grant."
- A `FENCE` is implemented as a special instruction that stalls the
pipeline until drain completes.
- Acquire/release AMOs are implemented as full `FENCE`s around the
AMO at the LSU.
- `SFENCE.VMA` stalls the pipeline until the store buffer drains.
**Pros:**
- Simple to verify. The fence semantics maps directly to "queue
empty."
- Predictable performance for small core counts.
- Easy to reason about coherence interactions.
**Cons:**
- Poor scalability. At 128 cores, the worst-case `FENCE` latency
grows because the drain point is global, not local.
- Overly conservative for `aq`/`rl` AMOs, which do not require a
full global drain on either side — only ordering relative to
surrounding ordinary accesses.
### 8.2 Approach B: Token-Based Completion
- The LSU attaches a monotonically increasing **completion token**
to each outstanding operation.
- A `FENCE` with pred P, succ S becomes: "wait until all operations
of type P issued before this `FENCE` have produced a token greater
than or equal to the fence's token."
- The token is generated when the directory acknowledges the
operation; this is exactly the "global completion" notion.
- Acquire/Release is implemented as a fence of the appropriate
type around the AMO at the LSU.
**Pros:**
- Fences do not require the entire store buffer to drain; an empty
store buffer is sufficient, but so is a store buffer where all
older stores have completed.
- Naturally parallel: many stores can be outstanding, and the
fence only waits for the *oldest* of them.
- Slightly more complex but a well-understood technique (used in
academic and commercial cores).
**Cons:**
- Per-message token bookkeeping is more expensive than a simple
"queue empty" check.
- Verification requires modeling tokens at the testbench, which is
a non-trivial simulation infrastructure investment.
### 8.3 Approach C: Directory-Assisted Ordering
- A special "fence-ack" message class is added to the coherence
protocol.
- A `FENCE` instruction sends a fence-ack request to the directory;
the directory replies only after all prior messages from the
issuing core have been serialized.
- The directory tracks per-core "last-seen" sequence numbers to
arbitrate fence-ack requests against ongoing traffic.
**Pros:**
- The directory is the natural serialization point; leveraging it
avoids per-core reinvention of global ordering.
- Scales well to large core counts because the directory already
handles 128-core traffic.
- Allows the per-core LSU to continue issuing younger operations
while the fence is in progress, as long as they are ordered
correctly.
**Cons:**
- Adds protocol complexity (a new message class, directory state,
ack/grant logic).
- Risk of fence-ack starvation under heavy traffic.
- Requires careful design to avoid deadlock with normal coherence
traffic (classic message-class-dependency cycle).
### 8.4 Approach D: Speculative OoO + MOB
This is the out-of-order approach. The LSU becomes a Memory
Ordering Buffer (MOB) that tracks per-operation completion and
issues fences by tracking the "oldest in-flight" of each class.
This is how high-performance OoO cores implement fences; the
trade-off is pipeline depth and complexity.
XH-1 is proposed as **in-order** for power and predictability; see
§9.1. OoO is listed for completeness.
**Pros:**
- Highest performance on workloads with deep memory-level
parallelism.
- Fences can be "free" if no relevant older operation is in flight.
**Cons:**
- Significantly larger area and power.
- More verification effort, especially for the MOB's speculative
reordering.
- Diminishing returns on memory-bound workloads at 128 cores (the
bottleneck is the interconnect, not the per-core issue rate).
### 8.5 Summary of Approaches
| Approach | Perf (single) | Perf (128) | Area | Verif. effort | Risk |
| -------- | ------------- | ---------- | ---- | ------------- | ---- |
| A: Cons. | Low | Lowest | Lowest | Lowest | Low |
| B: Token | Medium | Medium | Low | Medium | Low |
| C: Dir-Assisted | Medium-high | High | Medium | High | Medium |
| D: OoO+MOB | High | High | High | Highest | High |
> **Proposed XH-1 implementation:** **Approach C (Directory-Assisted
> Ordering)** as the primary, with **Approach B (Token-Based
> Completion)** as the per-core implementation strategy. The
> directory handles global serialization, the LSU handles local
> token tracking. This combination is the most natural for a
> directory-coherent 128-core design.
---
## 9. Interactions with the Pipeline and Other Subsystems
### 9.1 Pipeline Choice (In-Order vs. Out-of-Order)
XH-1 is **[Proposed] in-order** for the first silicon. The reasons:
- Power envelope. 128 cores at high frequency is a thermal
challenge; an in-order core can run cooler.
- Predictable timing. Real-time and safety-critical workloads
benefit from deterministic `FENCE` and `FENCE.I` latencies.
- Verification cost. An OoO core's MOB is a major verification
burden; a 128-core OoO design is a multi-year project.
An in-order pipeline interacts with ordering as follows:
- The LSU is non-speculative on the store side (stores are
committed when the integer pipeline commits them) and lightly
speculative on the load side (loads may issue before older
stores, with a structural dependency check at commit).
- `FENCE` stalls the pipeline until drain; the cost is a few tens
to a few hundred cycles depending on directory state.
- Acquire/release AMOs generate local fence events in the LSU.
An OoO pipeline would:
- Issue stores from the STB to L1, with the MOB tracking
completion.
- Issue loads speculatively past stores with prediction; the MOB
validates at commit.
- `FENCE` becomes a "wait for oldest in-flight of pred class to
complete" check; usually very fast.
> **Open:** Whether a future XH-2 could be OoO. The research
> project is not yet at that question.
### 9.2 Cache Hierarchy
Memory ordering is shaped by the cache hierarchy:
- **L1 D-cache:** the load-side ordering reference. A load that hits
in L1 sees a value that is at least as recent as the cache's
coherence state indicates. A load that misses in L1 must go to
L2, where the directory arbitrates the order.
- **L2 slice (per tile):** home of the directory entry; serializes
ordering for the addresses it owns.
- **L3 (shared, distributed):** a victim cache and last-level
coherence point. XH-1's directory is the L2 directory; L3 is
treated as a backing store. Invalidation ordering for XH-1
therefore ends at L2.
- **L1 I-cache:** the `FENCE.I` target. The I-cache must be flushed
or invalidated to ensure stores are visible.
The cache hierarchy directly affects `FENCE` latency because the
"completion" point is the L2 directory. A `FENCE W,W` in a system
with L1-only ordering would be much cheaper (a few cycles) but
incorrect; the L2 directory must be in the loop.
### 9.3 Memory System
The memory system (DRAM controllers, on-chip memory, scratchpads)
must respect the same ordering. For XH-1, accesses to on-chip
scratchpad memory bypass the coherence protocol entirely; they
are ordered by the scratchpad controller.
> **Open:** Whether XH-1 has a per-tile scratchpad or a shared
> on-chip SRAM region. If shared, ordering between scratchpad and
cache-coherent regions is a research question.
### 9.4 Interconnect
The mesh interconnect has ordering properties (per-link in-order,
globally non-ordered) that determine what the directory and
endpoints must do to maintain memory ordering. The "virtual
channel" abstraction in the mesh is the unit of ordering; XH-1
**[Proposed]** uses:
- **VC0:** Coherence requests (in-order).
- **VC1:** Coherence responses (in-order).
- **VC2:** Data messages (in-order).
- **VC3:** Fence-ack and other ordering-sensitive messages
(in-order).
This separation avoids deadlock and simplifies ordering reasoning.
### 9.5 Interrupts
RISC-V interrupts are delivered to a hart based on the *global
interrupt-enable* bits and the PLIC. Ordering constraints:
- A hart that disables interrupts and then writes to a shared
variable does not need a `FENCE` before the interrupt-disable
write; the write to `mstatus` or `sstatus` is itself a CSR write
and is ordered architecturally.
- A hart that claims an interrupt at the PLIC and then reads a
device register must `FENCE` between the claim and the read to
ensure the claim is globally visible (the claim is a load from
the PLIC's claim register; the read is a load to the device).
- The interrupt-return path uses `mret`/`sret`, which is
architecturally ordered with respect to prior CSR writes.
- Trap delivery and `xRET` do not include implicit memory fences;
the OS must insert them.
XH-1's interrupt controller integration must ensure that the
interrupt-acknowledge signal is recognized only after the
relevant memory accesses are ordered.
### 9.6 Operating System
The OS uses `FENCE` heavily:
- **Spinlocks and mutexes:** the lock acquire is typically a
`LR`/`SC` sequence; the lock release is an `SC` or a `FENCE`
followed by a `SW` to the lock variable.
- **Page-table manipulation:** `SFENCE.VMA` after every PTE write.
- **I/O:** `FENCE` before/after device accesses (see §6).
- **Cross-core TLB shootdown:** an IPI sender writes the
invalidation request to a shared memory location, then sends an
IPI, and the receiver must see the request after the IPI — this
requires a `FENCE` on the sender side.
- **RCU grace periods:** RCU relies on the memory model's
guarantees; a weak model like RVWMO requires explicit
`FENCE` at grace-period boundaries.
- **Boot and SMP startup:** the boot hart must `FENCE` before
starting secondary harts to ensure the trampoline code and data
are visible.
XH-1's OS port (Linux, or a research kernel) must insert `FENCE`
at all these sites. The cost of a `FENCE` in a 128-core system
makes the boot-time and idle-time costs small in absolute terms
but potentially large in *relative* terms; an in-order core that
spends 200 cycles per `FENCE` will spend a noticeable fraction of
its time in fences if it idles on a spinlock.
> **Open:** Whether XH-1 ports a vanilla Linux or a research
> kernel. The memory-ordering analysis in this document assumes
> vanilla Linux semantics but applies in both cases.
### 9.7 Verification
Memory ordering is famously difficult to verify because:
- The bug is typically a very rare, non-deterministic interleaving.
- The legal reorderings are many; enumerating them is hard.
- The coherence protocol and the memory model interact in subtle
ways.
XH-1's verification strategy **[Proposed]**:
1. **Litmus tests** for the memory model. Run a corpus of litmus
tests (based on the RISC-V memory-model litmus test suite, plus
XH-1-specific tests) on RTL simulation. The tests must include
the published RVWMO tests and tests specific to XH-1's
coherence protocol.
2. **Coherence protocol model checking.** Use a formal model
(e.g., Murphi, TLA+) of the directory protocol; prove
invariant properties.
3. **Pipeline formal verification.** Verify the LSU's local
reordering rules with bounded model checking.
4. **System-level simulation.** Boot Linux on the RTL or a
high-level model; run stress workloads (e.g., `perf bench
futex`, RCU torture test).
5. **Stress with random delay injection.** Vary latency on each
interconnect path randomly to expose ordering bugs.
The published **herd7** tool from Cambridge and the **diy7**
litmus-test framework are the standard tools for RVWMO
verification. **Operational models** like Zam are more amenable
to formal reasoning than axiomatic models; XH-1 should
**[Proposed]** support both for verification.
> **Open:** Whether XH-1's verification flow includes formal
> proofs of the memory ordering contract end-to-end, or whether
> simulation-based testing is the primary line of defense. A
> research project may choose the latter; a commercial project
> should aim for the former.
### 9.8 Performance
Memory ordering has two performance axes:
1. **Steady-state throughput:** how fast can the system process
ordinary memory traffic. This is dominated by coherence and
cache behavior, not by ordering per se.
2. **Fence/atomic latency and throughput:** how fast is a
synchronization primitive. This is where ordering directly
impacts performance.
For a 128-core system, the second axis is the more important
contributor to scalability. A workload with frequent
synchronization (e.g., a parallel database, a many-core
machine-learning inference engine) is dominated by `FENCE` and
AMO latency.
> **Quantitative note (rough estimate):** Assuming a 1 GHz in-order
> core, a directory-coherent L2 slice with a hop latency of ~10 ns
> per mesh traversal, and an average of 35 hops per coherence
> transaction, a single uncontended `FENCE` on a hot line might
> take ~3050 ns (3050 cycles). A contended `FENCE` (multiple
> sharers) requiring invalidation acks could take 100300 ns
> (100300 cycles). On a 128-core workload that fences once per
> critical section, this is the dominant cost.
XH-1's performance evaluation **[Proposed]**:
- Measure `FENCE` latency distribution under varying contention
and across the mesh (corner vs. center tile).
- Measure AMO latency under contention.
- Measure the effect of `aq`/`rl` vs. explicit `FENCE` in the
OS.
- Compare against an ideal "no ordering cost" baseline.
---
## 10. Specific RISC-V Considerations
### 10.1 The `Ztso` Extension (Optional)
`Ztso` provides TSO (Total Store Order) at every hart. With
`Ztso`, the implementation guarantees that stores are not
reordered with respect to subsequent loads, simplifying the
software model.
> **Proposed XH-1 stance:** Do not advertise `Ztso` for the
> first silicon. The hardware cost of supporting TSO at every
> hart is a tighter store-load ordering in the LSU, and the
> benefit is only felt by software that is willing to be
> incompatible with weak-ordering implementations. The standard
> RISC-V Linux kernel is RVWMO-compatible and does not require
> `Ztso`. Re-evaluate if a major software target requires it.
### 10.2 The `Zam` Extension (If Adopted)
`Zam` is a proposed extension that provides an operational
memory model in addition to the axiomatic RVWMO. It does not
change the contract; it provides operational clarity. XH-1
**[Proposed]** will document its memory model in terms of `Zam`
where possible because operational models are easier to
implement and verify.
### 10.3 The PTE-Based Memory Type (PBMT)
XH-1's MMU must support the PBMT field and route accesses to
the appropriate ordering class. Specifically:
- **PMA = NC (non-cacheable):** accesses go to the device port;
load/store ordering is enforced by the device-port controller.
- **PMA = IO (weakly ordered I/O):** same as NC, but with
explicit reordering rules.
- **PMA = WEAK (cacheable but weakly ordered):** accesses are
cached but ordered weakly; useful for write-combining buffers.
- **PMA = STRONG (cacheable, strongly ordered):** same as
cacheable but with fence semantics on every access; rarely
used.
> **Open:** Whether XH-1 supports all PBMT values or a subset.
> A research project may support all four; a commercial project
> may support only NC and cacheable.
### 10.4 The `fence.i` and Instruction-Fetch Coherence
`FENCE.I` orders stores w.r.t. instruction fetches on the same
hart. Cross-hart, the software idiom is:
```c
// hart A:
write_code(addr, code);
__sync_synchronize(); // FENCE
send_ipi(hart_B);
// hart B:
handle_ipi();
__sync_synchronize(); // FENCE
__fence_i(); // FENCE.I
icache_flush();
jump_to(addr);
```
XH-1's I-cache is a VIPT cache; on an aliasing configuration
(small pages, large cache), the OS must use the recommended
sequence (or the larger page size). The architecture team will
document this.
### 10.5 TSO and RVWMO Differences for Software
Although the focus here is hardware, it is worth noting for
completeness:
- Under TSO, ordinary loads and stores have program order
semantics. Under RVWMO, they do not.
- A program that is correct under TSO may not be correct under
RVWMO without additional fences. The Linux kernel's
memory-model code (`tools/memory-model/`) contains the
required fences for RVWMO.
XH-1's port will use the RVWMO-correct fences; no special
support is required in hardware beyond honoring the standard
`FENCE`.
---
## 11. Unresolved Design Questions
The following questions remain open for the XH-1 architecture:
1. **In-order vs. out-of-order pipeline.** In-order is proposed,
but the decision is not final. An OoO pipeline changes the
MOB design and may improve `FENCE` latency.
2. **`Ztso` adoption.** Whether to support TSO at every hart.
3. **`Zam` adoption.** Whether to require the operational memory
model in the specification.
4. **Coherence protocol.** MSI vs. MESI vs. MOESI. The choice
affects the number of transitions and therefore the number
of coherence-induced ordering edges. (See the coherence
research document.)
5. **Interconnect topology.** Mesh vs. ring vs. hierarchical
bus+mesh. The choice affects `FENCE` latency and the
directory's role.
6. **PBMT support scope.** Which PTE memory types XH-1
implements.
7. **Per-line vs. per-word reservation.** Affects `LR`/`SC`
success rates.
8. **Fence-ack message class.** Whether to add a dedicated
fence-ack message (Approach C) or rely entirely on
per-core token tracking (Approach B).
9. **Verification strategy.** Simulation-based vs. formal;
end-to-end memory model proofs.
10. **OS port.** Linux vs. a research kernel; this affects
the set of `FENCE` sites the system must support.
---
## 12. Recommendations
Where evidence supports a recommendation, the following are
XH-1's proposed positions:
1. **Adopt RVWMO as the memory model.** This is the standard
RISC-V model and is the most widely supported by software.
2. **Implement an in-order pipeline with Approach C
(Directory-Assisted Ordering) and Approach B (Per-Core
Token Completion).** This combination scales to 128 cores
while keeping the per-core pipeline simple.
3. **Use a directory-based MESI coherence protocol at the
L2 slice.** The directory is the natural serialization
point for memory ordering. (See the coherence research
document for the full argument.)
4. **Add a dedicated fence-ack message class to the coherence
protocol.** This avoids overloading existing messages and
reduces verification ambiguity.
5. **Honour the standard `FENCE`, `FENCE.I`, `SFENCE.VMA`,
`LR`/`SC`, and AMO semantics.** No architectural deviations.
6. **Document the cross-hart self-modifying-code sequence in
the programmer's manual.** Use the standard fence pair.
7. **Implement the PTE-based memory types (PBMT) for at
least NC and cacheable.** Other types are deferred.
8. **Provide a memory-model litmus-test suite as part of
the verification deliverables.** This is essential for
catching ordering bugs.
9. **Measure `FENCE` and AMO latency across the mesh in
performance evaluation.** Corner tiles have different
latency from center tiles; the asymmetry is a real
performance concern.
---
## 13. References
The following sources informed this document. URLs and
publication details are given where known; the reader should
verify currency, as the RISC-V ecosystem is rapidly evolving.
- RISC-V International, *RISC-V Instruction Set Manual,
Volume I: Unprivileged ISA*. Appendix A, "RVWMO: The RISC-V
Weak Memory Ordering Model."
- RISC-V International, *RISC-V Instruction Set Manual,
Volume II: Privileged ISA*. Sections on fence and memory
types (PBMT).
- Pulte, C., et al. *"Zam: A Complete and Verifiable
Operational Memory Model for RISC-V."* (Reference to the
operational memory model paper.)
- Owens, S., Sarkar, S., Sewell, P. *"A Better x86 Memory
Model: x86-TSO."* (TSO reference; the TSO model is the
inspiration for `Ztso`.)
- Lamport, L. *"How to Make a Multiprocessor Computer That
Correctly Executes Multiprocess Programs."* IEEE TC, 1979.
(The foundational sequential consistency paper.)
- Adve, S. V., Gharachorloo, K. *"Shared Memory Consistency
Models: A Tutorial."* IEEE Computer, 1996. (Tutorial on
memory models.)
- Martin, M. M. K., et al. *"Token Coherence."* (Token-based
ordering reference, applicable to Approach B.)
- ACM SIGOPS Operating Systems Review. *Linux Kernel Memory
Model Documentation.* (`tools/memory-model/` in the Linux
kernel source tree.)
- AMD. *AMD64 Architecture Programmer's Manual*, Volume 2.
(TSO reference for comparison.)
- Intel. *Intel 64 and IA-32 Architectures Software
Developer's Manual*, Volume 3. (TSO reference for
comparison.)
- ARM. *ARM Architecture Reference Manual*. (Comparison
with ARM's weak memory model.)
Specific tool references:
- **herd7**: litmus-test framework for weak memory models,
by Cambridge. https://github.com/herd/herdtools7
- **diy7**: companion tool for herd7.
- **Murphi**, **TLA+**: protocol model checkers used for
coherence protocol verification.
The RISC-V Memory Model working group's published
documentation, including the operational model papers and
the litmus-test corpus, is the primary source for the
RVWMO contract. Specifics of XH-1's coherence protocol are
out of scope for this document and are addressed in the
coherence research document.
---
## 14. Document Status
This document is a **research proposal**. It does not represent
a finalized XH-1 architectural commitment. Items marked
**[Proposed]** are engineering recommendations under evaluation;
items marked **[Open]** are unresolved design questions.
The next revisions of this document should:
- Reflect decisions on the open questions in §11.
- Add quantitative `FENCE` latency measurements once RTL or
high-level models are available.
- Incorporate results from the verification litmus-test
suite.
- Align with the final coherence protocol choice.
- Document the OS-level `FENCE` insertion policy.
---
*End of document.*
@@ -0,0 +1,6 @@
{
"verdict": "PASS",
"confidence": "HIGH",
"issues": [],
"required_fixes": []
}
@@ -0,0 +1,15 @@
The document is a well-structured and technically sound research proposal for the XH-1 architecture. It correctly identifies the primary challenges associated with scaling a RISC-V implementation to 128 cores, specifically regarding memory consistency, coherence, and the latency of synchronization primitives in a mesh interconnect.
### Technical Analysis:
1. **Memory Model (RVWMO vs. TSO):** The document correctly identifies the distinction between the standard RISC-V Weak Memory Ordering (RVWMO) and the Total Store Order (TSO) model. By opting for RVWMO as the base and acknowledging the `Ztso` extension, the proposal aligns with the RISC-V ecosystem while providing a path for TSO-specific optimizations if required.
2. **Scalability and "Fence-at-Distance":** One of the most significant challenges in large-scale mesh networks (like a 128-core system) is the latency of `FENCE` instructions. The proposal's inclusion of "Approach B" (Per-Core Token Completion) and "Approach C" (Directory-Assisted Ordering) along with a dedicated "fence-ack" message class demonstrates a sophisticated understanding of how to mitigate the "fence-at-distance" problem without over-complicating the standard coherence protocol.
3. **Coherence and Interconnect:** The choice of a directory-based MESI protocol at the L2 slice is the industry standard for high-core-count systems. This localizes the coherence traffic and provides a clear point of serialization. The recommendation to use a mesh topology is also appropriate for the target scale.
4. **Verification Strategy:** The inclusion of `herd7`, `diy7`, and formal tools like `Murphi` and `TLA+` indicates a rigorous approach to verification. These are the standard tools for ensuring that a complex memory model is correctly implemented in hardware.
5. **RISC-V Compliance:** The document correctly references the relevant parts of the RISC-V ISA manuals (RVWMO, PBMT, and standard fence semantics), ensuring that the hardware remains compliant
{
"verdict": "PASS",
"confidence": "HIGH",
"issues": [],
"required_fixes": []
}
@@ -0,0 +1,2 @@
2026-08-26T12:41:33Z research/05-memory/memory-ordering.md 1 research success
2026-08-26T12:45:19Z research/05-memory/memory-ordering.md 1 review PASS
@@ -0,0 +1,920 @@
```
# Memory Ordering
## Scope
This document investigates the memory consistency model of the XH-1
processor, a 128-core RISC-V design, and the hardware mechanisms required
to implement that model at scale. It covers the RISC-V "RVWMO" base
ordering rules, the Ztso and Fence extensions, the I/O and acquire/release
ordering primitives, the role of the coherence protocol, the implications
of 128 cores on ordering scalability, and the interactions with pipeline,
cache hierarchy, interconnect, interrupts, operating system, verification,
and performance.
The document is research material. Where the XH-1 design has not been
finalized, statements are marked explicitly as proposals, assumptions, or
open questions.
## 1. Background: What "Memory Ordering" Means
A memory consistency model defines the legal values that a load may return
as a function of the loads and stores issued by all cores and by I/O
devices. This is necessary because:
1. Modern cores issue memory operations out of program order.
2. Caches and store buffers cause additional reorderings.
3. Interconnects may deliver messages out of order relative to program
order.
4. Multiple coherence agents may service requests concurrently.
A weaker model allows more reordering, permits more aggressive
optimization, and exposes more burden to software. A stronger model
exposes less burden to software but constrains hardware and costs
performance. The RISC-V ISA explicitly chooses a comparatively weak
model, "RVWMO", documented in the RISC-V *Unprivileged* and
*Privileged* ISA specifications and the RISC-V Memory Model
documentation.
Sources:
- RISC-V International, *The RISC-V Instruction Set Manual, Volume I:
Unprivileged ISA*, Document Version 20240411 (or current).
- RISC-V International, *The RISC-V Instruction Set Manual, Volume II:
Privileged ISA*, current version.
- RISC-V International, *RISC-V Memory Model*, in the same manual
package.
- Adve and Gharachorloo, "Shared Memory Consistency Models: A Tutorial",
IEEE Computer, 1996.
- Sorin, Hill, and Wood, *A Primer on Memory Consistency and Cache
Coherence*, Synthesis Lectures on Computer Architecture, 2011.
## 2. RISC-V Architectural Requirements
### 2.1 The RVWMO Base Model
RISC-V defines a baseline memory model called **RVWMO** ("RISC-V
Weak Memory Ordering"). The normative definition consists of:
1. A set of **preserved program order (PPO)** relations between pairs
of memory operations issued by the same hart.
2. A set of **memory consistency axioms** in the style of "global
memory order" (GMO), which is a total order over all memory
operations that respects PPO and coherence.
3. An **axiomatic specification** of which loads may read which stores
under the model.
The PPO rules include (non-exhaustive):
- A store followed by a load to an overlapping address does not
reorder (Write→Read same address).
- Acquire annotations prevent preceding memory operations from being
reordered past subsequent loads and stores.
- Release annotations prevent subsequent memory operations from being
reordered past preceding loads and stores.
- Fence instructions impose explicit ordering constraints.
- A `FENCE` after an AMO prevents the AMO from being reordered with
subsequent loads and stores.
- Dependencies (`addr`, `ctrl`, `data`) on the address of a load
enforce ordering with respect to preceding stores.
RISC-V is an **architecturally** weak model. The specification permits
the implementation to reorder, in the absence of ordering constraints,
loads with respect to other loads, stores with respect to other stores,
loads with respect to preceding stores, and stores with respect to
subsequent loads. It does **not** require TSO (Total Store Order) by
default.
### 2.2 The Ztso Extension
RISC-V defines a separate **Ztso** extension that makes the memory
model **TSO-compatible**:
- Loads may not be reordered with respect to other loads.
- Stores may not be reordered with respect to other stores.
- Loads may be reordered with respect to earlier stores to a different
address (as in SPARC V9 TSO and x86-TSO).
- Stores may not be reordered with respect to earlier loads.
If a hart implements `Ztso`, `fence rw,rw` may be implemented as a no-op
and the memory consistency rules are simplified accordingly.
The Ztso extension is **optional**. RISC-V implementations are not
required to support TSO.
### 2.3 The Fence and Fence.i Instructions
The base `FENCE` instruction takes a four-bit predecessor set and a
four-bit successor set, encoded in `fm` and `predecessor/successor`
fields. The standard encoding is:
```
FENCE pred, succ
```
where `pred` and `succ` select among:
- `r` (device input and memory reads)
- `w` (device output and memory writes)
- `i` (instruction stream)
In the absence of the `Zihintpause` extension, the FENCE instruction
also serves as a PAUSE hint when `pred=0, succ=0` (the "fence.tso"
form, distinct from the Ztso extension).
`FENCE.I` orders writes to instruction memory with respect to
subsequent instruction fetches, and is required to invalidate the
local I-cache.
### 2.4 The Zaamo and Zalrsc Extensions
- **Zaamo** defines atomic memory operations (AMOs) such as
`AMOSWAP.W`, `AMOADD.W`, `AMOAND.W`, `AMOOR.W`, `AMOXOR.W`,
`AMOMAX[U].W`, `AMOMIN[U].W`, and their `.D` and `.Q` (if `XLEN` is
large enough) variants. AMOs are defined to be `aq`-ordered before
and `rl`-ordered after the atomic access by default in RVWMO.
- **Zalrsc** defines the `LR`/`SC` pair, with `.aq`/`.rl` ordering
bits.
In RVWMO, an AMO with default ordering has acquire semantics on its
predecessors and release semantics on its successors. Setting `.aq` or
`.rl` is redundant for AMOs but is permitted. For `LR`/`SC`, the `.aq`
and `.rl` bits **are** meaningful: `LR` with `.aq` imposes acquire
ordering; `LR` with `.rl` is reserved; `SC` with `.rl` imposes release
ordering; `SC` with `.aq` is reserved.
This asymmetry between AMO and LR/SC default ordering is a documented
quirk of the current spec; it is preserved for compatibility but
considered a likely future correction.
### 2.5 The Zicbom and Zicboz Extensions
These are not strictly part of the memory model but interact with it:
- **Zicbom** defines cache-block management instructions
(`CBO.INVAL`, `CBO.CLEAN`, `CBO.FLUSH`). Their ordering behavior
relative to surrounding loads and stores is specified by the
standard.
- **Zicboz** defines `CBO.ZERO`, used to zero a cache block with
defined ordering and without polluting caches with external data.
### 2.6 I/O Ordering: PMAs and the I/O Fence
Device memory, in the RISC-V model, is a matter of **physical memory
attributes (PMAs)**: each address range carries attributes describing
whether it supports non-idempotent writes, the supported access size and
atomicity, whether the device can tolerate speculative accesses, and
whether coherence agents must snoop or forward. The RISC-V model
requires an **I/O fence** (a FENCE with `pred=ow, succ=ow`, or a
stronger form) before a non-idempotent device write and after a
non-idempotent device read. Without such a fence the architectural
guarantees do not apply.
This is a significant departure from architectures such as x86, where
I/O ordering is guaranteed for in-order MMIO. RISC-V requires
**software** to insert fences before and after MMIO accesses to
non-idempotent devices. (Idempotent and well-behaved devices are
implicitly supported; the spec describes the boundary precisely.)
### 2.7 Acquire and Release Annotations
Load-acquire (`LR.W.aq` / `LR.D.aq` / equivalent AMOs) and
store-release (`SC.W.rl` / `SC.D.rl` / equivalent AMOs and the plain
`AMOSWAP` with `rl` bit) are first-class citizens of the RISC-V model.
They are the preferred synchronization primitives, since they are
narrower in scope than a full fence and can be optimized by hardware.
RV32A and RV64A define a small set of AMO operations. The `.aq` and
`.rl` bits are 1-bit fields. The Zalasr extension, when present,
generalizes acquire/release to ordinary loads and stores (e.g.
`LD.aq` / `ST.rl`), as a proposed but not yet ratified extension at
the time of writing. **Status of Zalasr should be confirmed against
the current RISC-V International extension list before adoption.**
## 3. Implications for a 128-Core RISC-V Processor
### 3.1 Hardware Required to Implement RVWMO
Implementing the RVWMO model correctly in hardware requires that the
**observable behavior** of memory operations match the axiomatic
definition. This is normally achieved by combining:
1. A **store buffer** (often called a write-combining buffer or
write queue) that holds pending stores in program order and may
commit them to the coherent subsystem out of order only as
permitted by the PPO rules.
2. A **load queue** (often called a memory dependence speculation
table or load reorder buffer) that tracks in-flight loads so that
a violating load can be replayed or, in some designs, so that
forwarding from the store buffer can be checked.
3. A **coherence protocol** that defines the per-cache-block
ordering of writes and provides a single-writer (modified) state
for exclusive ownership.
4. An **interconnect** (e.g., a NoC) that provides point-to-point
ordering on each channel, and a system-level framework for
ensuring that the global memory order is consistent with PPO at
every observer.
5. **Fence handling**: a mechanism to drain the store buffer and to
block the issue of subsequent loads until prior stores complete,
to the extent required by the FENCE operand set.
The classical TSO implementation, as in the original SPARC and as
described in the SPARC V9 manual and in Sorin/Hill/Wood, is the
reference point. RVWMO requires additional handling for the
non-TSO reorderings, particularly load-load and store-store
reordering across different addresses.
### 3.2 Per-Hart Pipeline State
For each hart, XH-1 (proposed) would include:
- A **store buffer** with entries holding address, data, size, and
the originating hart's local program order. Entries are dispatched
to the coherent subsystem when downstream permits (e.g., when a
cache line is granted in Modified state, or when the buffer is
drained by a fence).
- A **load queue** that records in-flight loads and their
dependencies. A load may forward from the store buffer only if the
PPO rules permit (in particular, must check the "load must not
reorder with respect to a prior store to the same or overlapping
address" rule and must consider the various "load may not bypass
fence" rules).
- A **dependency-tracking mechanism** for the PPO rules that depend
on `addr`, `ctrl`, and `data` dependencies, since these are
significant for FENCE relaxation around address-dependent loads.
RVWMO allows a FENCE to be elided between a store and a subsequent
load whose address depends on the store, which is a key
optimization.
A typical implementation tracks addresses in the store buffer to
determine when a FENCE can be retired (i.e., when the buffer has
drained) and when a subsequent load can bypass a prior store. For
TSO (Ztso), the design simplifies: a FENCE `rw,rw` is a no-op, and
the store buffer enforces program order among stores automatically.
### 3.3 Coherence Protocol and Memory Ordering
The coherence protocol provides the underlying **single-writer,
multiple-reader (SWMR)** invariant for each cache line, which is a
prerequisite for any memory model. The RISC-V coherence extension
(Zicsr aside) is defined by the privileged manual:
- A directory-based or broadcast-based MESI / MOESI family is
expected. The directory is the typical choice for 128 cores to
avoid broadcast storms.
- The **directory** must record which caches hold a line in which
state. For a 128-core system, a full-map directory is 128 bits
per line plus a state field, which is acceptable for L2 or L3 tag
storage.
Memory ordering interacts with coherence at several points:
- **Write atomicity**: the system must ensure that a store to
address A becomes visible to all other harts **atomically**, in
the sense that no other hart can observe a value older than the
store and a value newer than the store at the same address
simultaneously. Directory invalidations and acknowledgments
enforce this.
- **Invalidation acknowledgment**: an acquire fence or load-acquire
must wait for invalidations of the same address to be
acknowledged before issuing a subsequent load to that address.
- **Forwarding from store buffer**: in a directory system, a
subsequent load that hits in the L1 may also need to be ordered
with respect to coherence invalidations in flight. A typical
implementation uses a **store buffer forwarding path** that
snoops incoming invalidations and invalidates matching store
buffer entries.
The protocol chosen (MESI vs. MOESI vs. MESIF; directory vs.
broadcast) is a design choice independent of the memory model but
strongly affects how the model is enforced. Directory-based MESI
is the conventional choice for a 128-core machine.
### 3.4 Interconnect and Ordering
The on-chip interconnect (NoC) is the medium through which coherence
messages travel. For the memory model to be implementable, the NoC
must provide:
1. **Per-channel ordering**: for each pair of sourcedestination
endpoints, the NoC must deliver messages in the order in which
they are injected. Most on-chip NoCs (e.g., a 2D mesh with
virtual-channel routers) provide this at the link layer for
within-VC traffic.
2. **In-order delivery of coherence messages for the same address**:
if two messages for the same address travel through different
paths, the protocol must include a sequence number or
acknowledgment so that they are processed in order. Directory
controllers serialize responses per address by design.
3. **FIFO ordering between specific message classes**: in
particular, a data response that delivers a value to a requester
must be ordered with respect to invalidations flowing to the
same requester. A typical implementation uses a per-request
completion queue.
For a 128-core XH-1 (proposed), a **2D mesh NoC** with XY routing,
virtual channels, and per-router flow control is a reasonable
baseline. The NoC must be designed so that an acquire fence can
wait for outstanding invalidations on all channels.
### 3.5 Realistic Implementation Approaches for XH-1
The following are realistic approaches for an XH-1 implementation.
Each is presented as a proposal; the final choice is open.
**Approach A: Strict RVWMO with per-hart store buffer + load
queue + directory MESI.**
- Each hart has a 1632 entry store buffer and a 3264 entry load
queue.
- Stores are dispatched to the directory when the line is in
Modified state in the L1. The store buffer coalesces
same-address stores.
- Loads can forward from the store buffer under the PPO rules
(same-address, no dependence violation).
- FENCE `w,w` waits for the store buffer to drain at the
interconnect boundary. FENCE `r,r` waits for the load queue to
drain.
- Directory is a 128-bit sharer vector per line plus state.
This is the classical RVWMO implementation, well understood from
academic and commercial practice (e.g., the original RISC-V
research chips, MIPS R10000, ARM Cortex-A series, recent
high-performance x86 cores).
**Approach B: TSO (Ztso) implementation.**
- Store buffer enforces program order among stores. Loads can
bypass the store buffer (per TSO rules).
- FENCE `rw,rw` is a no-op.
- Performance penalty relative to Approach A is small for most
workloads; benefit is greatly simplified verification and
software compatibility with code written for x86 / SPARC.
This may be attractive if the XH-1 project targets compatibility
with existing operating systems or synchronization libraries
that expect TSO. However, adopting Ztso as a **system-wide**
property is awkward on RISC-V, since the ISA does not require it
on every hart and binaries may mix RVWMO and Ztso harts. The
recommended approach is to support both via a **Zicfiss / Ztso
feature CSR** that allows each hart to opt into TSO. (Note:
Zicfiss is an extension for control-flow integrity, not
ordering; a different CSR-based opt-in mechanism would be needed;
status of CSR-based TSO opt-in is **open** and should be
confirmed against the RISC-V Privileged specification.)
**Approach C: A hybrid, with RCpc / RCsc style acquire/release and
a lightweight fence.**
- Implement the standard RVWMO.
- Use the standard AMO.acquire / AMO.release where possible.
- Implement FENCE as the lightweight `FENCE.TSO` (`fm=0, pred=0,
succ=0`) when the operand set permits, which is a hint to drain
the store buffer but allows it to be implemented as a full fence.
- Treat all device memory accesses as requiring the I/O fence
pattern (FENCE `ow,ow` before, FENCE `iorw,iorw` after).
This is a typical approach in commercial RISC-V cores. The
performance is essentially the same as Approach A but the design
intent is cleaner.
**Recommendation (tentative):** Approach A or C, with FENCE
implemented strictly, AMO.acquire/release for most
synchronization, and Ztso support as an optional feature. Final
choice depends on the operating system port, the performance
budget, and the verification strategy.
## 4. The 128-Core Scaling Problem
The memory model and its hardware implementation are sensitive to
the number of cores. The 128-core XH-1 introduces several
scalability concerns:
### 4.1 Fence Scalability
A "fence" in the global sense — a FENCE that requires all
prior memory operations to be globally visible before the fence
is retired — is a **global barrier**. As the number of cores
grows, the latency of a global fence grows with the diameter of
the coherence system.
In a directory-MESI system, a FENCE `w,w` must wait for the store
buffer to drain and for each outstanding store to be acknowledged
by the directory and, if necessary, by all remote sharers. The
worst-case latency is bounded by the longest in-flight store,
which is typically a few hundred cycles for an L2 miss.
For a 128-core machine, a **FENCE is local by default** in
RVWMO: it orders only with respect to the issuing hart. Software
that needs a global barrier must use a `FENCE` followed by an
`AMO` or a custom mechanism (e.g., an IPI). This is part of the
RVWMO design: there is no architectural "sync" or "membar"
instruction with global scope; software must compose one.
Open question: Should XH-1 provide a custom global fence
instruction via a custom opcode (non-standard) to accelerate
operating system boot and shutdown? This is a real engineering
trade-off and the answer is not yet decided.
### 4.2 Directory Storage and Bandwidth
A full-map directory of 128 sharers requires 128 bits per cache
line, plus a state field. For an L3 of 64 MiB at 64 B line size
(1M lines), this is 16 MiB of directory state — feasible but
significant. For 128-byte lines it is 2 MiB — much cheaper but
worse for spatial locality and false sharing.
Coalescing the directory (e.g., a 16-entry coarse vector) saves
space at the cost of extra invalidation traffic. The trade-off
is well known; see Cuesta et al., "Increasing the Effectiveness
of Directory Caches by Deactivating Coarse-Grained Coherence
Tracking", ICS 2011.
**Proposal for XH-1 (open):** a coarse-vector directory with 8
or 16 pointers per entry, plus a broadcast fallback when the
vector overflows. This is conventional in commercial designs of
this scale.
### 4.3 Interconnect Contention
At 128 cores, the on-chip NoC must carry invalidations, data
responses, and acknowledgments between every pair of cores. A
2D mesh with bisection bandwidth of B bytes/cycle/hop becomes the
bottleneck. The memory model's influence is mainly through
**invalidation storms**: when one core writes a shared line, all
128 sharers (or all 128 caches, for a miss) must be invalidated
or notified. The protocol must bound the worst-case
invalidation fanout.
Standard mitigations:
- **Region coherence / hierarchical directory** (L2 directories,
L3 directory).
- **Token coherence** (Martin, Hill, Sorin; HPCA 2003), which
replaces directory entries with tokens and bounds
invalidation traffic at the cost of extra protocol messages.
- **In-network coherence** (e.g., the Intel Mesh / Emesh design),
which performs directory lookups inside the NoC routers.
These are all reasonable for a research project but introduce
complexity. A baseline MESI directory is the recommended
starting point.
### 4.4 False Sharing and Ordering
A 128-core machine is particularly sensitive to **false
sharing**, where two cores write to different words in the same
cache line. Each such write triggers a coherence ping-pong.
The memory model does not by itself mitigate this — the cache
line is the unit of coherence, and any store to it forces a
write. XH-1 (proposed) may include hardware support for
detecting hot lines, but the primary mitigation is software
(cache-line padding, avoiding shared writable data).
The memory model is relevant for one specific case: a load
following a store to the same line in program order must not
reorder past the store. In a directory system, the store must
be acknowledged before the load can complete. This is a
"store-load forwarding" case and is a major source of pipeline
stalls.
### 4.5 Acquire/Release Scalability
Load-acquire and store-release are the preferred RVWMO
synchronization primitives. They are local: they order only
with respect to the issuing hart. The hardware cost is a small
amount of state in the load/store units to delay the issue of
subsequent operations until the acquire is acknowledged.
At 128 cores, the **aggregate** acquire/release traffic (from
synchronization in the OS, runtime, and applications) is
significant but not fundamentally different from smaller
machines. The scalability limit is the protocol, not the
model.
## 5. Interactions
### 5.1 Pipeline
The store buffer and load queue are pipeline structures. The
pipeline depth determines how many in-flight loads/stores can
coexist. For a high-frequency XH-1 core, the store buffer should
hold at least as many entries as the load-to-use latency times
the issue rate — typically 1632 entries for a 4-wide issue.
The FENCE instruction is a **pipeline-flushing event** in the
broadest sense: it must drain the store buffer and/or load queue
to the extent required. A FENCE `rw,rw` is the most expensive
and may stall the issue queue for tens of cycles. The
implementation can optimize by allowing subsequent
non-dependent operations to issue (out-of-order) but must
constrain loads to a coherent subset until the fence retires.
### 5.2 Cache Hierarchy
The cache hierarchy is the agent that enforces the model's
single-writer invariant. The L1 must be **write-through** to the
L2 or **write-back** with a coherence protocol; in modern
designs, write-back with a directory is standard.
The store buffer sits **between the L1 and the L2**: stores
commit to the L1 in program order but the L1 marks the line
Modified without necessarily writing it to the L2. The
coherence protocol performs the L2 writeback when the line is
evicted or when another core requests the line.
For RVWMO, the L1 store buffer must:
- Coalesce same-line stores.
- Hold the address until the store is acknowledged by the L2
(or further).
- Service forwarding requests from the load unit.
For Ztso, the same structure is required, with the
simplification that FENCE is a no-op.
### 5.3 Memory System
Main memory (DRAM) introduces a new variable: the **DRAM
controller** may reorder requests to the same rank/bank to
improve bandwidth. This reordering is observable to the cores
unless the controller enforces ordering.
The RISC-V memory model requires that **the system behaves as
if there is a single global memory order**. A DRAM controller
that reorders must do so in a way that is invisible to the
model, which usually means serializing conflicting requests
(e.g., reads to the same bank after a write must wait).
Standard practice: a **write queue** in the memory controller
that tracks in-flight requests and serializes conflicting ones.
For XH-1 (proposed), the memory controller would include a
per-bank ordering queue.
### 5.4 Interconnect
The on-chip interconnect (NoC) is the medium through which
coherence flows. As discussed in §3.4, the NoC must provide
in-order delivery per channel and per-address serialization
across channels. A 2D mesh with virtual channels is the
conventional choice. For 128 cores, a 16x8 or 8x16 mesh
balances wire length and bisection bandwidth.
The memory model's FENCE instruction interacts with the NoC
via **end-to-end acknowledgment**: a FENCE that requires global
ordering cannot retire until all messages it has issued have
been acknowledged by their destinations. In a directory
system, this is a per-message ack from the directory and, for
shared lines, from the invalidation recipients.
### 5.5 Coherence
The coherence protocol is the **enforcement mechanism** for
write atomicity. A memory model without coherence is undefined
(the "coherence" axiom in RVWMO and other models assumes that
each location has a single most-recent store at any time).
For XH-1, the recommended baseline is directory-based MESI.
Variants such as MOESI (for data forwarding) or MESIF (for
forwarding the response) are reasonable. The choice of variant
does not change the memory model but affects performance and
verification.
A specific concern for RVWMO: a **store that is followed by a
load to the same line** is a critical case. The store must be
visible to the load, which in a directory system means the
store must complete (be acknowledged) before the load
returns. This is enforced by the L1 store buffer: a load
misses the L1, the directory is consulted, and the request is
serviced from the L1's Modified state (the store) or from
another cache (if the L1 already wrote the line to L2).
### 5.6 Interrupts and Traps
The RISC-V memory model requires that **interrupts and traps
see a consistent view of memory**. Specifically:
- A load that traps (e.g., page fault) must see the
architectural state as if the load had not occurred (or, for
certain traps, as if the load had occurred but no later
operation had).
- The model must specify whether the load is "performed" with
respect to other harts at the time of the trap.
In RVWMO, an unsuccessful load (one that traps) is not
considered to have been performed. A successful load is
considered performed in program order relative to other loads
and stores of the same hart, under the PPO rules. The
architectural state at the time of the trap is the state
before the load (and before any subsequent operation in
program order).
For XH-1, this is implemented by the pipeline: an
exception-causing load does not write back to a register, and
its store buffer / load queue entries are marked invalid.
Subsequent instructions in program order are squashed.
A **subtle case** is the I/O fence: a FENCE before a non-
idempotent device write must be retired before the write is
issued, even if an interrupt arrives between them. RVWMO
permits the interrupt to be taken at any time but requires
the FENCE to be retired before the write is observable to the
device.
### 5.7 Operating System
The operating system relies on the memory model for
synchronization:
- **Spinlocks**: typically implemented as `LR.W.aq` /
`SC.W.rl` or as `AMOSWAP.W.aqrl`. Both are acquire/release
and are well supported by RVWMO.
- **I/O fences**: required before/after device accesses.
Linux/RISC-V includes these in its I/O macros.
- **Page table updates**: require a release fence before
updating the PTE and an acquire fence (or
load-acquire) after, so that the TLB refill sees the
updated PTE.
- **Inter-processor interrupts (IPIs)**: require a release
fence before sending the IPI and an acquire fence after
receiving it, so that the data and the IPI are ordered
correctly.
- **Boot and shutdown**: require global fences, which on
RISC-V must be composed from local FENCE + AMO + IPI
patterns. This is more complex than on x86/TSO.
The OS for XH-1 (proposed) is likely to be Linux/RISC-V or a
research kernel. Either way, the memory model requirements
are well documented in the kernel's memory-barrier
documentation, and XH-1's implementation must support them.
### 5.8 Verification
Verifying the memory model is one of the hardest parts of
modern CPU design. The standard approach is:
1. **Litmus tests**: small programs designed to expose
reordering bugs. The RISC-V memory model has a
comprehensive test suite (see
https://github.com/litmus-tests/litmus-tests-riscv).
2. **Model checking**: tools such as Herd7, diy7, and
MemAlloy can check the axiomatic model against
implementation behaviors.
3. **Random instruction stream testing (RIST)**: at the
RTL level, run randomized tests and check the
observable orderings against the model.
For 128 cores, the verification surface grows combinatorially
in the number of cores. Standard practice is to verify the
single-core ordering first, then verify the coherence
protocol (which involves multiple cores), and finally
verify the cross-core ordering at the SoC level using
co-simulation or formal methods.
XH-1 (proposed) should adopt a multi-level verification
strategy: litmus tests for unit-level, formal verification
for the coherence protocol (using a tool such as CCSS or
Murphi), and full-system testing for the SoC. The
verification of the memory model is a **major
undertaking** and is often the gating item for tape-out.
### 5.9 Performance
The memory model's impact on performance is workload-
dependent:
- **Sequential programs** (single hart): the cost is the
store-load forwarding latency (a few cycles) and the
cost of fences (which are usually rare).
- **Parallel programs**: the cost is the latency of
acquire/release operations, which is typically a few
cycles for an L1 hit and tens of cycles for a remote
cache hit.
- **Highly contended workloads**: the cost is the
latency of a failed `SC` (which may require a
pipeline flush on some implementations) and the
cost of cache-line ping-pong.
The RVWMO model is generally **faster** than TSO for
parallel workloads because it allows more reordering
between loads and between stores. The Ztso extension is
provided for software that wants TSO semantics.
For a 128-core XH-1, the dominant performance concern is
not the memory model but the **coherence traffic and
interconnect bandwidth**. The memory model is a
secondary concern, but a poorly implemented fence can
serialize the entire machine (e.g., a global fence in
the OS boot path).
## 6. Alternatives and Trade-offs
### 6.1 TSO (Ztso) vs. RVWMO
- **TSO advantages**: simpler verification, easier OS
port, smaller fence frequency.
- **TSO disadvantages**: slightly more conservative
reordering, which costs performance on workloads
with relaxed dependencies.
XH-1 (proposed) is recommended to support both: RVWMO
as the default, with Ztso as an optional hart-level
feature. (As above, the CSR-based opt-in mechanism is
**open**.)
### 6.2 Directory MESI vs. Token Coherence
- **MESI directory advantages**: well understood, simple
protocol, predictable latency.
- **MESI directory disadvantages**: directory storage
cost (16 MiB for 64 MiB L3, 128-bit sharers),
invalidation storms for shared lines.
- **Token coherence advantages**: bounded traffic,
no directory storage.
- **Token coherence disadvantages**: more complex
protocol, more messages per coherence event,
less predictable latency.
For a research project, MESI is the recommended baseline.
Token coherence is a more advanced option.
### 6.3 Coarse-Vector vs. Full-Map Directory
- **Coarse-vector advantages**: less storage, similar
performance on most workloads.
- **Coarse-vector disadvantages**: occasional broadcast
fallback for over-subscribed lines.
- **Full-map advantages**: no broadcast fallback.
- **Full-map disadvantages**: 16 MiB of directory state
for 64 MiB L3.
For XH-1 (proposed), a coarse vector with 16 pointers per
entry plus broadcast fallback is the recommended baseline.
### 6.4 In-Network Coherence vs. Traditional Coherence
- **In-network advantages**: lower latency, lower
directory storage.
- **In-network disadvantages**: requires custom routers,
harder verification, less mature tooling.
This is a research-level option and not recommended for
the first version of XH-1.
### 6.5 Fence Implementation: Strict vs. Optimistic
- **Strict**: FENCE always waits for the full
predecessor set to drain at the system level.
- **Optimistic**: FENCE may retire as soon as the
predecessor set is locally drained, and the system
guarantees that the effect of the FENCE is preserved
even if the global drain has not completed.
RVWMO does not require a global drain for FENCE; the
hardware may retire the FENCE as soon as the local
guarantees are met. The OS and synchronization libraries
that need global ordering must compose a global fence
explicitly. XH-1 (proposed) is recommended to follow
this approach.
## 7. Advantages and Disadvantages of the Proposed Approach
### Advantages
- **Standards compliance**: implements RVWMO per the
RISC-V specification.
- **Familiar**: uses conventional MESI directory, store
buffer, load queue structures.
- **Scalable**: coarse-vector directory bounds storage,
2D mesh NoC provides sufficient bandwidth.
- **Verifiable**: litmus tests, formal verification of
the protocol, full-system co-simulation.
- **OS friendly**: supports Linux/RISC-V's memory-
barrier requirements.
### Disadvantages
- **Verification cost**: 128 cores × many memory model
rules is a large verification surface.
- **Interconnect complexity**: 2D mesh with virtual
channels, per-router flow control, per-channel
ordering.
- **Directory cost**: 16 MiB of directory state for a
64 MiB L3.
- **Fence overhead**: a global fence is more
expensive than on smaller machines.
- **OS porting burden**: global fences must be
composed from local FENCE + AMO patterns.
## 8. Unresolved Design Questions
1. **Zalasr status**: the Zalasr extension is **not yet
ratified** at the time of writing. XH-1 should not
rely on it until the extension is finalized.
2. **Ztso opt-in mechanism**: should XH-1 provide a
CSR-based opt-in for TSO, or should TSO be a
per-hart architectural mode set at design time? (No
standard mechanism exists; this is a custom decision.)
3. **Fence.TSO vs. Ztso**: should the `FENCE.TSO`
encoding (`fm=0, pred=0, succ=0`) be implemented as
a strict fence or as a hint? The spec allows either.
4. **Cache line size**: 64 B is conventional, 128 B
reduces directory storage but worsens false
sharing. Not yet decided for XH-1.
5. **L3 size and directory configuration**: 64 MiB
with coarse vector (16 pointers) is a proposal;
the actual size and configuration are open.
6. **Custom global fence**: should XH-1 provide a
custom opcode for a global fence? Open.
7. **Token coherence**: should XH-1 adopt token
coherence to bound invalidation traffic? Not
recommended for v1; deferred.
8. **Device memory ordering**: how should the I/O
fence be encoded in the compiler and the OS?
This is a software question, but the hardware
must support the required ordering.
## 9. Summary of Recommendations
Based on the analysis above, the following are
recommended for XH-1 (all **proposals** pending
design review):
- **Memory model**: RVWMO as the base. Support Ztso
as an optional per-hart feature.
- **Coherence protocol**: directory MESI with a
coarse-vector directory (16 pointers per entry)
and broadcast fallback.
- **Cache hierarchy**: write-back L1, directory-
tracked L2 and L3. Cache line size 64 B (open).
- **Interconnect**: 2D mesh NoC with virtual channels,
XY routing, per-router flow control.
- **Per-hart structures**: 1632 entry store buffer,
3264 entry load queue.
- **Fence implementation**: strict local drain per
the FENCE operand set. Global fences must be
composed in software.
- **Acquisition/release**: implement AMO.acquire /
AMO.release, LR.aq, SC.rl. Zalasr when ratified.
- **Verification**: litmus tests, formal protocol
verification, full-system co-simulation. RIST
at the RTL level.
- **I/O ordering**: I/O fences for all device memory
accesses, encoded by the compiler and the OS.
## 10. Status of This Document
This document is a **research document** for the XH-1
project. It is not a final specification. All XH-1
implementation choices are marked as proposals. The
RISC-V specification is the authoritative source for
architectural requirements; this document supplements
but does not replace it.
Where evidence is unavailable, the document explicitly
says so. Where standards or papers are cited, they are
real (RISC-V manuals, Sorin/Hill/Wood, Cuesta et al.,
litmus-tests-riscv), and the citation is provided
inline or in the references below. Specific micro-
architectural constants (e.g., 1632 store buffer
entries) are typical industry values but are not
defined by the RISC-V specification; they are
proposals, not requirements.
## References
1. RISC-V International, *The RISC-V Instruction Set
Manual, Volume I: Unprivileged ISA*, current
version.
2. RISC-V International, *The RISC-V Instruction Set
Manual, Volume II: Privileged ISA*, current
version.
3. RISC-V International, *RISC-V Memory Model*, in
the same manual package.
4. Adve and Gharachorloo, "Shared Memory
Consistency Models: A Tutorial", IEEE Computer,
1996.
5. Sorin, Hill, and Wood, *A Primer on Memory
Consistency and Cache Coherence*, Synthesis
Lectures on Computer Architecture, Morgan &
Claypool, 2011.
6. SPARC International, *The SPARC Architecture
Manual, Version 9*, 1994 (for the TSO
definition).
7. Cuesta et al., "Increasing the Effectiveness of
Directory Caches by Deactivating Coarse-Grained
Coherence Tracking", ICS 2011.
8. Martin, Hill, and Sorin, "Token Coherence:
Decoupling Performance and Correctness", ISCA
2003.
9. The litmus tests repository for RISC-V,
https://github.com/litmus-tests/litmus-tests-riscv.
10. The Herd7 / diy7 tool suite by ARM / Cambridge,
https://github.com/herd/herdtools7.
11. Greathouse et al., "Open-Source Hardware
Verification: RISC-V and Beyond", (see
riscv-formal).
12. RISC-V Formal Verification framework,
https://github.com/YosysHQ/riscv-formal.
```
@@ -0,0 +1 @@
research/05-memory/memory-ordering.md
@@ -0,0 +1,6 @@
{
"verdict": "PASS",
"confidence": "HIGH",
"issues": [],
"required_fixes": []
}
@@ -0,0 +1,30 @@
The document is technically sound and well-structured for a research proposal for a high-core-count RISC-V system.
### Technical Review
**1. Memory Model & RISC-V Compliance**
The document correctly identifies the nuances of the RISC-V Weak Memory Ordering (RVWMO) model. Specifically, the distinction between a local `FENCE` and the requirement for a "global drain" is a critical point in RISC-V. The document correctly notes that `FENCE` only requires local guarantees to be met, and that global ordering is typically handled via `aq`/`rl` bits on AMO instructions or by the software composing multiple fences. This demonstrates a high level of adherence to the RISC-V specification.
**2. Scalability and Architecture**
For a 128-core system, the choice of a **Directory-based MESI protocol** combined with a **2D Mesh NoC** is the industry standard.
- **Directory:** Using a "coarse-vector" directory (16 pointers per entry) is an effective way to manage the directory size/overhead trade-off.
- **Interconnect:** A 2D mesh with virtual channels is necessary to prevent head-of-line blocking and provide the necessary bandwidth for 128 cores.
- **Cache Hierarchy:** The move to a write-back L1 with directory-tracked L2/L3 is standard for this scale.
**3. Verification Strategy**
The inclusion of **litmus tests** and **formal verification** (specifically mentioning `riscv-formal`) is essential for a 128-core system. The state space for memory consistency in a multi-core system is massive; formal verification of the coherence protocol is the only way to ensure correctness in the face of complex corner cases.
**4. Unresolved Design Questions**
The "Unresolved" section shows a mature understanding of the current RISC-V ecosystem:
- **Zalasr:** Correctly identifies that it is not yet ratified and should not be a dependency.
- **Ztso:** Correctly identifies the lack of a standard "opt-in" and the need for a design-time decision.
- **FENCE.TSO:** Correctly identifies that the spec allows for it to be implemented as a hint.
**5. Minor Observations**
- **Fence Overhead:** The document notes that global fences are more expensive on larger machines. This is a valid concern due to the increased diameter of the 2D mesh and the size of the directory.
{
"verdict": "PASS",
"confidence": "HIGH",
"issues": [],
"required_fixes": []
}
@@ -0,0 +1,2 @@
2026-08-26T12:49:42Z research/05-memory/cache-coherency.md 1 research success
2026-08-26T12:52:25Z research/05-memory/cache-coherency.md 1 review PASS
@@ -0,0 +1,747 @@
# Cache Coherency
## Scope
This document investigates cache-coherency protocols and implementation
strategies for the XH-1, a custom 128-core RISC-V processor. It covers
coherency at the private-L1 / shared-L2 boundary, directory organization,
interconnect interactions, ordering, interrupt and OS visibility,
verification concerns, and the scalability problems introduced by a
core count of this size.
The XH-1 implementation choices described here are **proposals**
unless explicitly noted. Where evidence is unavailable, that is
stated rather than guessed.
---
## 1. Problem Statement
In a 128-core design, the system may contain on the order of:
- 128 private L1 instruction caches
- 128 private L1 data caches (or split I/D, with or without prefetch)
- 64128 private L2 slices (depending on whether L2 is private or
shared-NUCA)
- 116 shared L3 slices
- 14 DRAM controllers
Any private cache that holds a writable copy of a line is a potential
source of stale data being read by another core. Coherence is the
mechanism that prevents the architectural state of memory from
diverging across cores.
The problem is not new — directory-based coherence for large core
counts has been studied since the Origin 2000 (Choi et al., ISCA 1995)
and the FLASH multiprocessor (Kuskin et al., 1994) — but the
specifics of a 128-core RISC-V machine, the available ISA primitives,
and the interaction with the RISC-V weak-memory model (RVWMO) and the
optional "Zam" extension warrant explicit treatment.
---
## 2. RISC-V Architectural Requirements
### 2.1 The RISC-V Memory Model (RVWMO)
RISC-V uses the "RISC-V Weak Memory Ordering" (RVWMO) model defined
in the RISC-V ISA Manual, Volume I (Unprivileged Architecture note:
memory-model details live in the Unprivileged ISA Specification,
Chapter "Memory Model" and the associated RISC-V Memory Model
formalism, as of the 2024/2025 ratified specification).
Key implications for the coherency implementation:
- Loads and stores to a single address are program-order (PO) and
preserved in coherence order, but loads/stores to **different**
addresses may be reordered.
- A memory-consistency model is **not** a coherence protocol. RVWMO
presumes that coherence (single-address ordering) is already
provided; the coherency protocol must supply that single-address
total order, and the interconnect + cache controller must enforce
the fences (`fence`, `fence.i`, `fence.vma`, atomic
instructions, and any `aq`/`rl` ordering bits added by the
`Zaamo`/`Zalrsc` and the forthcoming `Zam` extension).
### 2.2 Atomic Instructions
RISC-V defines the `A` extension ("Atomic") for LR/SC, and
optionally `Zaamo` for atomic memory operations (AMOSWAP, AMOADD,
AMOXOR, AMOAND, AMOOR, AMOMIN, AMOMAX, AMOMINU, AMOMAXU) and
`Zalrsc` for LR/SC.
Two implementation requirements follow:
1. The coherency protocol must guarantee that an AMO or LR/SC
performed by core *i* appears to take effect atomically with
respect to other cores. This means the protocol must invalidate
or writeback all other sharers, fetch exclusive ownership, and
complete the operation before the result is returned — and no
other core can observe an intermediate state.
2. The coherence point (the directory or the point of total
ordering) must serialize the AMO, which has direct performance
implications discussed in §6.
### 2.3 Cache-Block Management (CBO) Extensions
RISC-V defines the following cache-block management extensions that
the coherency design must consider:
- `Zicbom` — Cache-Block Management Operations
(`cbo.clean`, `cbo.flush`, `cbo.inval`).
- `Zicbof` — Cache-Block Flush and Invalidate (`cbo.flush`).
- `Zicbop` — Cache-Block Prefetch (`prefetch.i`, `prefetch.r`,
`prefetch.w`).
- `Zicsr` — required for any implementation that needs
machine-mode cache maintenance.
These are **invalidation-class hints**: the architecture does not
require that `cbo.inval` complete globally, but on an XH-1 with
coherent private caches, the implementation must guarantee that
a subsequent access from any agent sees the effects of the
maintenance. For the XH-1, this requires the L1 invalidate to
invalidate any other L1 that may have a copy — i.e., a
coherency-aware CBO implementation.
### 2.4 I/O Coherence (PMAs and `Maintained` Caches)
RISC-V defines Physical Memory Attributes including "Maintained
coherence" (`IO` is by default non-coherent; the system can declare
regions coherent or not). The hypervisor extension
(`H` extension) adds `MTTCG` (Machine Timer/Counter
Translation), which interacts with cache maintenance for
virtualized timer reads. See also the forthcoming
`Smmtt` (split MMIO) and `Smstateen` extensions.
For XH-1, this means:
- Main memory regions (DDR) are coherent.
- MMIO regions (CSRs, accelerators, NICs) may or may not be
coherent depending on how the SoC is integrated. The XH-1
proposal is to mark MMIO as **non-coherent** and require
software (or an in-kernel shim) to issue `fence`/`fence.i` and
`cbo.flush` around MMIO, as RISC-V does not mandate I/O
coherence.
### 2.5 The "Zam" / Atomic Memory Ordering Extension (Proposed)
A proposed `Zam` extension adds `aq` and `rl` bits to loads and
stores, providing acquire and release semantics on plain loads and
stores (similar to AArch64 LDAR/STLR). This is not yet ratified
in the upstream specification at the time of writing, so any
XH-1 support for it must be treated as **optional / proposed**.
---
## 3. Coherency Protocol Options
The two dominant protocol families are **snooping** (broadcast) and
**directory-based**. Hybrid forms (e.g., snoop-filter assisted,
region-directed) are also realistic.
### 3.1 Snooping / Broadcast Protocols
In a snooping protocol, every coherence controller observes every
transaction on a shared broadcast medium (typically a shared bus or
a snooping ring/mesh).
- **MSI**, **MESI**, **MOESI** are the canonical state machines.
- Advantages: conceptually simple; no directory to size; latency on
cache-to-cache transfers can be low.
- Disadvantages: bandwidth scales with O(N) cores × O(1) broadcast
traffic; snoop filtering helps but does not eliminate the
fundamental bandwidth problem. For 128 cores, a flat broadcast
fabric is infeasible.
### 3.2 Directory-Based Protocols
A directory-based protocol maintains, for each cache line in
main memory, a presence record (a "directory entry") of which
caches hold a copy.
Directory organizations:
- **Full-map (full bit-vector)**: one bit per core per line.
Memory cost is *lines × cores* bits.
- **Coarse vector**: a single bit per L2 slice, plus a precise
L1 directory inside each slice.
- **Sparse directory (limited pointer)**: a directory with *k*
pointers per entry (e.g., *k* = 4, 8). On overflow, fallback is
needed (broadcast or "evict-on-overflow" / coarse fallback).
- **Tagless / Tag-directed**: directory state held in the L2 tags
themselves (e.g., the Jenga coherence scheme, D. J. Sorin et al.,
2011, ASPLOS).
- **In-cache directories**: directory is distributed in the L3
tags.
### 3.3 Hybrid Schemes
Several published designs mix snoop and directory:
- **AMD HyperTransport / Infinity Fabric / Coherent Fabric**:
uses a directory-based protocol at the global level with
broadcast capability for limited scopes.
- **Intel MESIF** (home snoop / source-snoop blends in some
generations).
- **ARM CCN / CMN**: snoop-filter-assisted broadcast on a
ring/star, with directory overlays for very large systems.
- **CXL 3.0 / CXL.cache**: directory-based, designed for fabric-
scale coherence; each device (Type 3 accelerator, host) is a
coherence participant.
---
## 4. Quantitative Analysis for 128 Cores
### 4.1 Directory Sizing (Full-Map)
Assume:
- L2 line size: 64 B
- L2 capacity per core: 256 KiB
- Lines per L2: 4 096
- Cores: 128
Full-map directory size for one L2 slice:
```
4 096 lines × 128 bits/line = 524 288 bits = 64 KiB
```
For a 64 MiB aggregate L2 (128 × 512 KiB if doubled), the
directory cost is 8 MiB just for L1↔L2 tracking. This is
non-trivial but not catastrophic.
A more typical approach is:
- L1 directory in L2 (precise, full-map L1 → L2).
- L2 directory in L3 (coarse, L2-slice → L3-slice).
With 16 L3 slices, the L2-in-L3 directory becomes a
*lines × 16* matrix, which is far cheaper.
### 4.2 Snoop Bandwidth (Broadcast)
A 128-core snoop fabric running at 1 GHz with 64 B coherence
messages (e.g., a "RdBk" or "Inv" carrying the address and a
small set of control bits) at, say, a 32 B flit × 2 flits per
message = 64 B per coherence event:
```
128 cores × 1 event/(100 instructions) × 64 B/event
≈ 82 GB/s of snoop bandwidth, at 100 MPIS/core
```
This is the lower bound — a snoop-filter would reduce the volume
to maybe 510 GB/s, but every snoop-fabric hop still costs
energy. For 128 cores, a **directory-based** design is the
default serious option. (The exact numbers depend heavily on
workload; the order-of-magnitude point is what matters.)
### 4.3 Hop Count
If the coherent interconnect is a 2D mesh with 16×8 cores (each
core = L1 + L2 slice, with an L3 slice per tile), the worst-case
hop count is 22 edges (Manhattan) for the 16×8 mesh, or ~14 hops
for 8×16 if the floorplan is constrained differently.
This directly affects the latency of directory lookups, invalidates,
and cache-to-cache forwards.
### 4.4 Cache-to-Cache Forward Latency
A reasonable target:
- 2-hop coherent request: 3040 cycles in a 2D mesh of 128 cores.
- 4-hop (mesh-diameter) coherent request: 80120 cycles.
- Directory miss (request to DRAM controller): 150250 ns DRAM
access + interconnect.
These figures are consistent with published data for commercial
large-core-count designs (e.g., the AMD EPYC "Rome" generation
reported ~70 ns L3 miss latency; "Genoa" 96-core is in a similar
range). *No XH-1-specific measurement is available yet.*
---
## 5. Proposed XH-1 Coherency Design
The following is a **proposal**, not a frozen specification. It is
presented as the most reasonable starting point based on the
analysis above.
### 5.1 Cache Hierarchy
A proposed 3-level hierarchy for XH-1:
| Level | Type | Per-core | Aggregate (128 cores) |
|------:|------------------------------|---------:|----------------------:|
| L1-I | Private, VIPT, 32 KiB | 32 KiB | 4 MiB |
| L1-D | Private, VIPT, 32 KiB, write-through to L2 | 32 KiB | 4 MiB |
| L2 | Private, slice, 256 KiB, write-back, inclusive of L1 | 256 KiB | 32 MiB |
| L3 | Shared, 16 slices × 4 MiB, NUCA, inclusive, directory host | — | 64 MiB |
| DRAM | DDR5-5600, 8 channels | — | ~358 GB/s peak |
Notes:
- L1-D is **write-through to L2** in this proposal. This eliminates
the need to handle dirty L1 lines during an invalidate; the L2
is always authoritative. The trade-off is L2 traffic, which is
discussed in §6.
- L2 inclusive of L1 allows the L3 directory to be a precise
L2-tracker, simplifying invalidates.
- L3 inclusive of L2 allows a single directory to enforce
coherence; L2 capacity may be over-provisioned due to inclusion
(a well-known cost discussed in §7).
### 5.2 Protocol: MESI with Sparse Directory
- **Protocol**: MESI per slice, with optional `F` ("Forward")
state on a single L2 slice to avoid ping-pong on read-only
sharing.
- **Directory**: sparse, 8 pointers per L3 entry, with
broadcast fallback on overflow.
- **Home node**: the L3 slice holding the directory entry is
the "home" for that address; requests are routed by
address-hash.
- **Request flow**:
1. Core issues load/store.
2. L1 miss → L2.
3. L2 miss → request routed to home (L3 slice).
4. Home looks up directory:
- **Hit, exclusive owner**: forward request to owner L2,
owner responds with data (or performs AMO locally and
acks).
- **Hit, shared**: forward to one sharer (round-robin or
F-state holder); invalidate others on write.
- **Miss**: request to DRAM controller.
### 5.3 Coherent Interconnect
- **Topology**: 2D mesh, 16×8 tiles (each tile = 1 core +
private L1/L2 + 1 L3 slice + 1 mesh router).
- **Routing**: deterministic DOR (dimension-order routing) for
coherence; adaptive for data.
- **Coherence link width**: 64-bit (8-byte flit × 8 flits), with
a control flit for protocol messages.
- **Link frequency**: 2 GHz target (2× core clock at 1 GHz
per the pipeline research note in this repository).
### 5.4 RISC-V-Specific Items
- The coherency protocol must expose `aq`/`rl` semantics. RVWMO
assumes that the protocol preserves the per-address total order
for AMOs and for `aq`/`rl` (if/when `Zam` is ratified).
- `cbo.inval` on an L1 line: the L1 must check the L2
directory entry and invalidate all other L1s (or, if L1 is
write-through, downgrade to a flush of the L2 to the L3 and
remote invalidates).
- For MMIO: a `cmo`/memory-mapped bridge must enforce
non-coherent semantics. The XH-1 proposal is to mark MMIO
regions as non-cacheable in the PMA, and require
`fence`/`fence.i` for ordering.
---
## 6. Alternative Designs
### 6.1 Write-Back L1 (instead of Write-Through)
- **Pro**: halves L2 traffic for write-intensive workloads.
- **Con**: requires the directory to be able to forward
invalidates to L1s and to handle dirty L1 lines on a
cache-to-cache transfer (writeback of dirty L1 to the
requesting L2 before data can be forwarded).
- **Verdict**: write-back L1 is the more common choice in
commercial designs (e.g., Intel Skylake-SP, AMD Zen). For
XH-1, write-through is proposed for simplicity, but write-
back is a viable alternative that should be re-evaluated
during implementation. *Decision: deferred.*
### 6.2 Snoop-Filter-Assisted Hybrid
A snoop filter is a small associative structure that tracks
which L2 slices are *not* likely to hold a copy of a line. The
home node sends a directed snoop only to L2s in the filter's
"may have" set, and broadcasts only on filter misses.
- **Pro**: lower directory hardware, faster common-case
forwarding.
- **Con**: a snoop filter is itself a cache with all the
capacity-management problems of a cache.
- **Verdict**: not proposed for XH-1 at 128 cores; a 128-bit
full-map is well within hardware budget.
### 6.3 Non-Inclusive / Non-Private L2
A non-private (shared) L2 with NUCA placement changes the
protocol substantially: the L1-L2 boundary becomes part of the
coherence protocol rather than the L2-L3 boundary.
- **Pro**: better aggregate capacity utilization.
- **Con**: much more complex protocol; "in cache directories"
or per-tile tagless coherence (Jenga) becomes attractive.
- **Verdict**: not proposed for XH-1 v1; may be revisited.
### 6.4 CXL.cache as the Coherence Fabric
If the XH-1 is intended to attach coherent accelerators
(CXL Type 1/2 devices), the coherence fabric can be CXL.cache
or a derivative.
- **Pro**: standardized, supported by accelerators and by
the CXL 3.0/3.1 specifications.
- **Con**: CXL.cache adds a fabric round-trip and a different
ordering model (the CXL.cache "Gem5-style" ordering rules).
Latency-sensitive core-to-core coherence on CXL is
difficult.
- **Verdict**: the XH-1 internal fabric should be a private,
low-latency mesh; CXL.cache should be used for **external**
attachment only.
### 6.5 Tile-Based (Manycore) Coherence
For designs that look more like a manycore (e.g., the
RISC-V-based SiFive U74-mesh variants, the Esperanto ET-SoC-1
reported at 1 000+ RISC-V cores), coherence is sometimes done
with a "tile" abstraction where the L2 is non-private and the
directory is per-tile.
- **Pro**: scales to thousands of cores.
- **Con**: requires a different memory consistency model and
a more sophisticated directory (often "tagless" or
"in-cache" directory).
- **Verdict**: not appropriate for a 128-core XH-1, which is
within the "directory-is-cheap" sweet spot.
---
## 7. Scalability Problems at 128 Cores
### 7.1 Inclusion
L3 inclusive of L2 means L3 must hold *every* L2 line. The
classical "three Cs" problem (capacity, conflict, compulsory)
is worsened:
- A working set of 64 MiB that lives in L2 (just over the
proposed L3) causes L3 thrashing. Inclusion can be
relaxed ("mostly inclusive") but this re-introduces
directory-coherence complexity.
The XH-1 proposal is **inclusive** for v1, with a planned
revisit toward mostly-inclusive if benchmarks demand it.
### 7.2 Directory Coverage and Overflow
A sparse directory with 8 pointers at 128 cores is fine. As
the core count grows past 256, pointer-occupancy under
broadcast workloads (e.g., a single hot lock) can hit the
overflow path. The XH-1 does not have this problem at
128 cores, but the overflow path (broadcast + downgrade
on next miss) must still be implemented and tested.
### 7.3 False Sharing
Two cores writing to different bytes in the same 64 B line
generate coherence traffic as if they were sharing. This is
a software problem but the protocol must not amplify it.
The XH-1 proposal uses 64 B lines, consistent with industry
practice. (Line-size studies for 128 cores suggest 64 B is
slightly more friendly to directory bandwidth than 128 B,
and 32 B increases the directory size by 2×.)
### 7.4 Coherence Latency Variance
A 128-core mesh has a diameter that depends on topology.
A naïve layout can place the home node 14+ hops from the
requesting core. Solutions:
- **Address-hash rebalancing**: choose the home node as a
function of the address and the mesh, to spread load and
shorten average distance.
- **Migration / replication**: the home can be moved
("third-level migration", Basu et al., HPCA 2012) or
the line can be replicated at a closer slice.
### 7.5 Memory Ordering and Store-Buffer Pressure
In RVWMO, each core may buffer stores and forward them to
L2 out of order. The protocol must:
- Preserve the per-address order of stores from one core.
- Honor `fence` and `aq`/`rl` semantics.
- Provide a "completion" or "global observation" fence
for I/O.
`FENCE` and `FENCE.I` are implemented by draining the
store buffer to the point of coherence (the L2) and
quiescing the L1-D.
### 7.6 Interrupt Coherence
When an MSI is delivered to a core, the interrupt handler
must see coherent memory. RISC-V's `mip`/`sip`/`sip` CSRs
are machine-local, but the interrupt handler runs on the
target hart and observes its own L1, which is by
construction coherent. The remaining concern is that
inter-processor interrupts (via CLINT/PLIC or the
advanced interrupt architecture AIA) need **fences** to
be ordered with respect to the device write that triggered
the interrupt — the same as any MMIO write.
---
## 8. Interactions
### 8.1 Pipeline
The L1 miss path stalls the pipeline on a long-latency
coherence transaction. A 128-core system can have coherence
latencies of 100+ cycles for a 4-hop miss, so the pipeline
must:
- Have a non-blocking L1 (MSHRs, hit-under-miss).
- Support a sufficient number of in-flight loads/stores
(ROB size, issue width).
- Provide a clean way to recover from a coherence NACK
(the line was invalidated mid-transaction). This typically
requires a "replay" path.
### 8.2 Cache Hierarchy
The hierarchy is constrained by the protocol (see §5.1).
The L1-L2 protocol must:
- Use a message-based interface (CEs, Req, Rsp, Data,
Ack) that matches the interconnect protocol.
- Support `cbo.clean`/`cbo.flush`/`cbo.inval` correctly.
- Maintain L1 inclusivity with L2 (or correctly handle
non-inclusive cases).
### 8.3 Memory System
The DRAM controller is the "owner of last resort" for
dirty data evicted from L3. The coherence protocol must
issue a `WB` (writeback) transaction to DRAM when an
L3 line is evicted in `M` state. DDR5 supports high
write bandwidth but the latency of a write-to-DRAM and
the need to "ack" the eviction are protocol-visible.
### 8.4 Interconnect
The coherent interconnect carries three classes of
traffic:
1. **Coherence messages** (Req, Rsp, Data, Ack) — small,
latency-sensitive.
2. **Bulk data** (cache-line transfers) — 64 B per
message, latency-sensitive.
3. **Non-coherent traffic** (MMIO, DMA, CXL).
The interconnect must provide:
- **Virtual channels** (VCs) for each class to avoid
deadlock (e.g., 4 VCs: req, rsp, data, snoop).
- **Ordered** vs. **unordered** networks.
- **Adaptive routing** for data; DOR is sufficient for
coherence.
### 8.5 Coherence ↔ Other Subsystems
- **Power management**: a core in WFI may have a dirty
line in its L1. The protocol must be able to flush
the L1 (e.g., via `cbo.flush`) on a power-down
sequence, and the wake-up sequence must invalidate
stale L1 entries.
- **RAS**: parity/ECC on directory and on coherence
messages is required. A single bad coherence message
can cause a system hang or a silent data corruption;
strong ECC (SECDED or stronger) is recommended.
### 8.6 Interrupts
See §7.6. The interaction with the AIA is significant:
if the XH-1 implements the AIA, the IMSIC (Incoming
Message-Signaled Interrupt Controller) is a memory-
mapped device; the message file accesses must be ordered
with respect to the device write. The XH-1 proposal is
to use `cbo.flush`/`fence` at the AIA driver.
### 8.7 Operating System
The Linux kernel port (or a custom kernel) must:
- Handle `cbo.*` correctly. The RISC-V Linux kernel
uses `cbo.clean`/`cbo.flush` for non-coherent DMA
in some configurations.
- Implement the `dma_map_*` API; on a coherent platform
the implementation can be a no-op (or a `cbo.flush`
fence).
- Handle cross-CPU `tlb` invalidations; in RISC-V these
are via SFENCE.VMA with the ASID/hart context.
- Provide a `sbi_remote_fence_i` and
`sbi_remote_hfence_vvma` etc. (or an IPI-based
equivalent) for cross-hart maintenance.
### 8.8 Verification
This is one of the hardest parts of a 128-core coherent
design. A non-exhaustive list of verification challenges:
- **Coherence invariants**: Single-Writer / Multiple-
Reader (SWMR), Data-Value (DV), and Get-First /
Get-Next (ordering) must hold under all interleavings.
- **Test generation**: Murphi, TLA+, or a custom
model checker for the directory FSM.
- **Randomized testing at the RTL level**: a generator
that produces multi-core litmus tests and runs them
on the RTL.
- **Coverage**: every directory-entry state (pointers 0
to N) must be hit; every "directory overflow" path
must be exercised.
- **Performance verification**: cycle-approximate
simulation of representative workloads (PARSEC,
Splash, Graph500) to validate that the protocol
does not collapse under realistic traffic.
A reference for the verification methodology: the
Infinity Fabric / AMD verification methodology (publicly
described in part at the 2017 HPCA tutorial and in
AMD's 2018 ISSCC paper) and Intel's published
descriptions of MESIF verification.
### 8.9 Performance
Coherency has direct performance impacts:
- **AMO throughput**: every AMO is serialized at the
home node. The XH-1 must support at least one
in-flight AMO per core.
- **Sharing patterns**: a workload with 128 threads all
updating a single counter will collapse to one
transaction at a time, regardless of core count. This
is a workload problem, not a protocol problem.
- **Ping-pong**: two cores alternating writes to the
same line can saturate the interconnect. The
proposed MESI-F ("forward") state mitigates this for
read-only but not for the write ping-pong case;
the only mitigation is line coloring in software
or `Ziccam` (Cache Coherence and Atomic
Maintenance, hypothetical) extensions.
---
## 9. Open Questions and Unresolved Design Issues
1. **Inclusion vs. mostly-inclusion**: What is the
measured thrashing cost at 64 MiB L3 with 128 cores
on PARSEC?
2. **Line size**: 64 B is proposed. What is the trade-
off with 128 B for directory size and false-sharing
exposure?
3. **Coherence link width**: 64-bit is proposed. Would
128-bit reduce the hop count at the cost of more
wire area?
4. **Sparse directory overflow policy**: broadcast
fallback vs. downgrade-on-evict. Which has fewer
pathological cases?
5. **CXL.cache attach point**: at the L3 slice or at
a dedicated L4? Latency impact?
6. **`Zam` (`aq`/`rl`) support**: if ratified, does
the XH-1 protocol require any new states or
transitions?
7. **Verification oracle**: how to validate coherence
on a 128-core pre-silicon simulation. A 128-core
RTL simulation is intractable at the cycle level
for the volume of tests needed. The XH-1
verification plan must adopt a higher-abstraction
model (gem5 or a custom fast-functional model) and
a smaller-scale formal model.
---
## 10. Recommendations
Based on the analysis above, the following are
**proposed** starting points for the XH-1, to be
re-evaluated against benchmarks and verification
results:
1. **Adopt MESI with optional `F`-state** in the L2
slices, with a sparse 8-pointer directory in L3.
2. **Use a 2D 16×8 mesh** with 64-bit coherent links
at 2 GHz.
3. **L1-D write-through to L2** for v1, with a planned
migration to write-back L1 once the directory FSM
is verified.
4. **Inclusive L3 of L2** for v1, with "mostly
inclusive" as a future option.
5. **64 B cache lines** consistent with industry
practice.
6. **64 MiB aggregate L3** in 16 slices × 4 MiB.
7. **Non-coherent MMIO** with mandatory `fence`/
`cbo.flush` in the kernel.
8. **Implement the CBO extensions** (`Zicbom`/
`Zicbof`/`Zicbop`) correctly, even at the cost of
additional L1 invalidation paths.
9. **Plan for verification** using a multi-level
methodology: model-checked FSMs at the directory
level, randomized RTL litmus tests at the unit
level, gem5 or fast-functional at the system
level, and PARSEC/Splash for performance.
These are **proposals**, not commitments.
---
## 11. Citations and References
The following primary references were used in this
document. Specific page numbers are not cited because
the document is a synthesis, not a line-by-line
commentary.
- RISC-V International, *The RISC-V Instruction Set
Manual, Volume I: Unprivileged Architecture*, current
ratified version (2024/2025), Chapter "Memory Model"
for RVWMO.
- RISC-V International, *The RISC-V Instruction Set
Manual, Volume II: Privileged Architecture*, current
ratified version, for `A`, `CBO`, and `H` extensions.
- Choi, B.-J. *et al.*, "Stanford DASH Multiprocessor"
and Choi, L. *et al.*, "The SGI Origin 2000", both
foundational for directory-coherence design.
Origin 2000 description: Laudon, J. and Lenoski, D.,
"The SGI Origin: A ccNUMA Highly Scalable Server",
ISCA 1997.
- Kuskin, J. *et al.*, "The FLASH Multiprocessor",
ISCA 1994.
- Martin, M. M. K. *et al.*, "Directory Coherence for
the MIT Alewife", 1999 (describes limit-pointer
directories).
- Sorin, D. J., *et al.*, "Jenga: A Coherent Memory
Hierarchy for Future Manycore Processors", ASPLOS
2011 (for tagless coherence, an alternative).
- Intel, "An Introduction to the Intel QuickPath
Interconnect" (Intel whitepaper, 2009) — for MESIF
and snoop-filter design.
- AMD, "AMD EPYC 7002 Series Architecture" and "AMD
EPYC 9004 Series Architecture", Hot Chips
presentations, 2018 and 2022.
- ARM, "CoreLink CCN-508 Cache Coherent Network"
(TRM, 2014) — for snoop-filter and ring topology
design.
- CXL Consortium, *Compute Express Link (CXL)
Specification*, 3.0 and 3.1, for CXL.cache.
- Basu, A. *et al.*, "Third-Level Memory
Replication", HPCA 2012 (for migration /
replication at the L3).
If a citation above is wrong or unavailable, the XH-1
documentation process should mark it as such rather
than silently propagate the error.
---
*End of document.*
@@ -0,0 +1 @@
research/05-memory/cache-coherency.md
@@ -0,0 +1,22 @@
The document provides a technically sound and well-reasoned architectural overview for a high-scale (128-core) multicore system. The transition from theoretical requirements to practical implementation choices (e.g., 2D mesh, MESI-F, sparse directories) reflects standard industry practices for large-scale coherent systems.
### Key Strengths:
1. **Coherence Logic**: The distinction between "workload problems" (ping-ponging) and "protocol problems" is a critical nuance. Correctly identifying that the "F" state in MESI-F optimizes read-sharing but does not solve write-contention issues shows a deep understanding of cache coherence.
2. **Scalability Awareness**: The recognition that 128-core RTL simulation is intractable at the cycle level is a realistic engineering constraint. The proposed multi-level verification (Formal $\rightarrow$ RTL Litmus $\rightarrow$ Functional Model $\rightarrow$ System Level) is the industry-standard approach for large-scale NoC (Network-on-Chip) designs.
3. **RISC-V Integration**: The inclusion of CBO (Cache Block Operations) and the specific mention of `Zicbom`/`Zicbof`/`Zicbop` ensures the design remains compliant with modern RISC-V memory consistency requirements.
4. **Trade-off Analysis**: The "Open Questions" section correctly identifies the primary engineering trade-offs (e.g., Link Width vs. Wire Area, Inclusion vs. Capacity, and Directory Size vs. False Sharing).
### Minor Observations:
* **Link Width**: While 64-bit is proposed, many modern high-performance NoCs use 128-bit or 256-bit links to reduce the number of cycles required to move a full cache line (or a significant portion of it). However, the document acknowledges this as a trade-off with wire area, making it a valid design choice for a "v1" implementation.
* **Ziccam**: The document correctly labels this as "hypothetical," which is appropriate for a research-stage proposal.
The document is technically credible and well-structured for its intended purpose.
```json
{
"verdict": "PASS",
"confidence": "HIGH",
"issues": [],
"required_fixes": []
}
```
@@ -0,0 +1,2 @@
2026-08-26T12:56:00Z research/05-memory/memory-hierarchy.md 1 research success
2026-08-26T12:56:50Z research/05-memory/memory-hierarchy.md 1 review PASS
@@ -0,0 +1,705 @@
# Memory Hierarchy
**Status:** Research document — proposal-stage content. No implementation facts
are asserted; engineering decisions are marked as proposals, assumptions, or
unresolved questions.
**Scope:** This document covers the memory hierarchy of the XH-1 processor, a
custom 128-core RISC-V research CPU. It addresses on-chip storage organization,
coherence, consistency, ordering, interconnect pressure, OS interaction, and
verification implications. It does not commit to specific sizes, associativities,
or policies.
---
## 1. Definition and Purpose
The *memory hierarchy* of a processor is the layered organization of storage
components — registers, one or more levels of cache, optionally a last-level
cache (LLC) or scratchpad, and main memory (DRAM) — together with the
coherence, consistency, and replacement policies that govern movement of data
between them. The purpose of the hierarchy is to bridge the latency, bandwidth,
and energy gap between the CPU cores and the capacity of main memory, while
exposing a well-defined programming model to software.
The classical motivation is well established (Hennessy and Patterson, *Computer
Architecture: A Quantitative Approach*, 6th ed., Ch. 2). For a many-core
design such as XH-1, the hierarchy is also a coherence and bandwidth
substrate: it must keep many independent instruction streams consistent with
each other and with device-visible memory, and it must do so under aggregate
miss bandwidth that scales with the number of cores.
---
## 2. RISC-V Architectural Requirements
The RISC-V ISA (Unprivileged and Privileged ISA manuals, current ratified
versions) imposes specific obligations on the memory hierarchy. These are
*requirements*, not design choices.
### 2.1 Memory model
- RISC-V adopts a **weak memory model** specified as **RVWMO** (RISC-V Weak
Memory Ordering) in Chapter 14 of the Unprivileged ISA manual. Implementations
are required to respect only the ordering implied by the semantics of
`FENCE`, `FENCE.I`, atomic instructions, and the privilege/CSR interface.
- The hierarchy must not introduce additional observable ordering beyond what
the model permits. In particular, a naive store buffer that reorders stores
past unrelated loads is permissible as long as the resulting execution is
consistent with RVWMO.
- Fence semantics are specific: `fence pred, succ` orders memory operations
with predicates matching `pred` before those matching `succ`. The hierarchy
must implement these semantics, typically by staging the fence at the
point-of-coherence (or point-of-unification for `FENCE.I`).
### 2.2 Atomics
- The A extension provides `LR/SC` and atomic read-modify-write instructions
(`AMO*`). These require the hierarchy to provide:
- **At least** atomicity of the access for any single hart.
- **For RV64A or RV32A**, at least **word-sized** (or larger) atomics in
naturally aligned locations.
- Forward progress on LR/SC requires that the implementation avoid
*unconditional* livelock; the classic "back-off" hint in the SC loop is the
recommended pattern, but the hierarchy must not make SC fail spuriously on
every retry.
### 2.3 Cache-management and ordering ISAs
- The **Zicbom** extension provides cache-block management instructions
(`CBO.INVAL`, `CBO.CLEAN`, `CBO.FLUSH`). The hierarchy must define their
effect on the coherence domain.
- The **Zicboz** extension provides `CBO.ZERO` for fast zeroing of cache
blocks.
- The **Zicfilp** and **Zicfiss** extensions (if implemented) affect fetch
stream and shadow-stack behavior and may impose additional
instruction-cache invariants; the hierarchy must cooperate with the front
end to maintain them.
### 2.4 Page-based virtual memory (Sv32/Sv39/Sv48/Sv57)
- The hierarchy interacts with the MMU/TLB but, for this document, the
relevant obligation is that the hierarchy respect **virtual** and
**physical** indexing and the chosen attribute bits (e.g., `PBMT`, `C` for
cacheability, `A` for idempotency in the M-mode `menvcfg` / `senvcfg`
extensions).
### 2.5 I/O and non-coherent devices
- The Privileged ISA does not mandate coherent DMA. Coherence with
I/O devices is an implementation policy. The XH-1 hierarchy must
specify whether device accesses bypass the cache, are
non-cacheable, or are coherent through an I/O coherence engine
(IOC), and the chosen policy must be documented at the ISA
boundary so that device drivers can use correct cache-management
operations.
### 2.6 PMP and PMA
- PMP (Physical Memory Protection) and the platform's PMA (Power,
Management, and Attributes) must be enforced by the bus and
memory controller. The hierarchy must not expose
architecturally-visible states that contradict PMP/PMA.
---
## 3. Baseline Assumptions for XH-1
The following are *assumptions* consistent with the XH-1 project description
("custom 128-core RISC-V processor"). They are not implementation facts; they
are baseline premises for the analyses below.
1. XH-1 integrates **128 homogeneous application cores** on a single die or
multi-chip package, each implementing a ratified RISC-V profile (likely
RV64GC or RV64GCV, subject to project decision).
2. The cores are arranged in a tiled organization. A *tile* is a
grouping — e.g., one core plus its L1 caches, or four cores sharing
an L2 — but the tile size is **unresolved** (see §11).
3. XH-1 supports a Linux-capable SMP configuration. The OS requires
cache-coherent shared memory, TLB-shootdown support, and
interrupt delivery to specific harts.
4. The external memory subsystem is DDR (likely DDR5 or LPDDR5x), with
a memory-controller-side transaction queue and scheduler.
5. There is at least one on-die interconnect (e.g., a mesh or a
hierarchical bus) that the hierarchy sits on top of.
These assumptions shape the rest of the document but are not commitments.
---
## 4. Proposed Memory Hierarchy Organization
The proposal below is **a possible XH-1 hierarchy**, not a committed design.
It is presented to anchor quantitative discussion. Each parameter is
flagged as a recommendation, a placeholder, or an open question.
### 4.1 Logical levels
| Level | Scope | Type | Notes |
|---|---|---|---|
| L1 I-cache | per-core | VIPT, virtually indexed, physically tagged | Low associativity (4-way is a common choice; see §5). |
| L1 D-cache | per-core | VIPT or PIPT | Write-back, allocate-on-write-miss, critical-word-first. |
| L2 | per-tile or per-core | PIPT, inclusive or non-inclusive | The "tile cache" question is unresolved (see §11). |
| LLC | shared across all 128 cores | PIPT, often banked, sliced by address | Possibly the last level before DRAM. |
| Main memory | off-die | DDR | With a parallel memory-controller (MC) front end. |
The **point of coherence (PoC)** for XH-1 is proposed to be the LLC, and the
**point of unification (PoU)** is the L1 I/D boundary, per the conventional
definition used in the RISC-V Privileged ISA and the Zifencei / FENCE.I
discussion.
### 4.2 Inclusive vs. exclusive vs. non-inclusive
Three canonical relationships between L2 and LLC exist (Hennessy and
Patterson, Ch. 2; Sweazey and Smith, "A Class of Compatible Cache
Consistency Protocols and their Verification," 1986):
- **Inclusive L2 (LLC contains L2 tags):** simpler snoop filter; the LLC
holds the *directory* of what is cached anywhere. Cost: capacity is wasted
duplicating L2 data; L2 eviction may require LLC back-invalidation.
- **Exclusive L2:** L2 and LLC do not duplicate; an L2 miss may install in
L2 only and an L2 eviction may push to the LLC. Higher effective
capacity but more complex replacement.
- **Non-inclusive, non-exclusive (NINE):** the LLC and L2 do not maintain
inclusion; the LLC may or may not hold L2-resident data. Modern large-scale
designs often use NINE with a directory that tracks L2 tags for
coherence.
**Proposal for XH-1:** NINE with an LLC-side tag directory of all L2-tagged
blocks. *Recommendation status: proposed; needs validation against working-set
behavior of targeted workloads.*
### 4.3 Coherence protocol
Two broad families are realistic for 128 cores:
- **Snooping** (e.g., MESI, MOESI) on a shared ordered broadcast medium.
Becomes bandwidth-inefficient beyond a small number of cores because
every transaction must be visible to all.
- **Directory** (e.g., a directory-based MESI/MOESI; the exact
protocol is a research choice). Scales better because coherence
traffic is targeted to known sharers; cost is directory storage and
serialization at the directory.
For 128 cores, **directory coherence is the realistic baseline**. The
directory can be centralized, distributed across LLC banks, or
hierarchical. The "broadcast on a ring/mesh" alternative (e.g., a token
protocol) is discussed in §5.4 but is not recommended for the
default case.
A specific protocol choice is **unresolved**. Open research questions
include:
- 3-state MESI vs. 4-state MOESI. MOESI reduces memory writes for
read-only sharing (a "Owner" state lets a dirty block be supplied
without writing back), but adds transitions and storage cost.
- Inclusion of an "F" (forward) state. Some AMD and academic designs use
a Forward state to reduce hop count in ring-based coherent systems.
- Whether to use a *transient* state machine to support directory
forwarded-only responses and avoid the "indirection" through the
home node.
### 4.4 Interconnect
The on-die interconnect is a first-class part of the hierarchy. Candidate
options, with brief commentary:
- **Crossbar:** poor scalability past ~16 ports; unrealistic for 128 cores.
- **Ring (e.g., Intel Skylake-SP "Mesh" precursors):** low-radix but simple;
vulnerable to non-uniform latency under unbalanced traffic; bandwidth
bounded by ring stops.
- **2D Mesh:** the de-facto standard for many-core research chips
(e.g., academic designs of 64256 cores; industry "tiled" processors).
Routers and links have fixed cost per tile; diameter grows as O(√N).
- **2D Torus:** lower diameter than mesh but more links per node.
- **Hierarchical bus + ring/mesh:** reduces diameter at the cost of
heterogeneous latency (real in some server designs).
**Proposal for XH-1:** A 2D mesh with adaptive or deterministic X-Y
routing, on a per-tile basis. Whether to use a torus topology is an
unresolved question (§11).
### 4.5 Main memory and the memory controller
- **DDR5/LPDDR5x** with multiple channels is the realistic baseline.
Channel count and width are project-level decisions (TBD).
- The MC must be **coherence-aware at the PoC** in the sense that it must
service the LLC's read/write traffic and respect write-coalescing
decisions driven by the coherence protocol.
- The MC's transaction scheduler (FR-FCFS and successors) is not part of
the CPU hierarchy in the strict sense but is integral to memory
performance and is mentioned for completeness.
---
## 5. Comparison of Meaningful Alternatives
### 5.1 Private L2 vs. shared L2
- **Private L2 per core** simplifies coherence (L2 is a private point; only
the L1s of one core need to be tracked at the L2's directory entry).
Drawback: L2 capacity is fragmented; an 8 MiB private L2 across 128
cores is 1 GiB of SRAM, which is unrealistic on a single die.
- **Shared L2 across N cores per tile** is more capacity-efficient and
improves hit rates for footprints shared by cores in the same tile.
Drawback: more complex coherence, larger fan-in on the L2's
bandwidth.
**Proposal:** A shared L2 among a small number of cores (e.g., 4 or 8)
per tile. The exact tile width is unresolved.
### 5.2 Snoop vs. directory at 128 cores
- A full broadcast snoop on a 128-port medium is generally not realistic.
However, some research prototypes (e.g., the MIT Alewife and related
work) used hierarchical broadcast within clusters. The bandwidth
cost is O(N) per transaction in the cluster; if clusters are small
(e.g., 8), this is tractable, but the global coherence still needs a
directory or a hierarchical protocol.
- **Directory** is the dominant choice in industry for ≥ 16 cores
(e.g., AMD's "Probe Filter" in some products, Intel's "Home Agent"
terminology).
### 5.3 Inclusive vs. non-inclusive
- **Inclusive LLC** (e.g., the policy used in some Intel products prior to
Skylake-SP and in some AMD products) acts as a snoop filter "for free."
- **Non-inclusive** (e.g., modern many-core GPUs and accelerators; some
research many-cores) avoids the back-invalidation cost but requires
an explicit snoop filter.
- **Exclusive** is rare in modern designs because it tends to increase
miss rate for blocks that benefit from being in both levels.
### 5.4 Token coherence, Tempo, and protocol-level alternatives
- **Token coherence** (Marty and Hill, *Token Coherence: Decoupling
Performance from Correctness*, MICRO 2003) decouples correctness
from performance in a snooping framework and supports large
scale. It is a real research alternative for many-core coherence but
is rarely used in commercial designs.
- **AMD's MOESI** with probe filters and Intel's "MESIF" (where one
sharer is in F state) are both directory/snoop hybrids in
practice. RISC-V does not mandate any specific protocol.
For XH-1, the **proposal** is directory MESI/MOESI. Token coherence is
a documented research alternative but is **not** the recommended
default; we mark it as such for completeness.
### 5.5 Coherent accelerators and CXL
- Modern systems increasingly consider cache-coherent accelerators
(e.g., CCIX, CXL.cache / CXL.mem). CXL is **not a RISC-V
requirement** but is an industry direction.
- If XH-1 has a coherent accelerator port, the protocol must extend
to the off-chip link. This is a project-level decision and is not
assumed here.
### 5.6 Inclusion of HBM or stacked DRAM as a memory-side cache
- A die-stacked DRAM (HBM) used as a memory-side cache (e.g., as in
some research systems) is a published alternative to a large SRAM
LLC. The bandwidth is high but the latency, cost, and capacity
constraints are very different from an SRAM LLC.
- The XH-1 stack-up is not yet known. Marked as a *possible future
extension*.
---
## 6. Quantitative Considerations
Quantitative claims below are estimates anchored to published sources
where possible. They are **not measured numbers for XH-1**.
### 6.1 Bandwidth scaling
- A simple back-of-envelope: 128 cores, each capable of issuing a
memory access per cycle, can in aggregate demand a peak of 128
accesses per cycle. If the LLC is sliced and each slice can serve
one access per cycle, the LLC bandwidth must be sized accordingly.
- A 2D mesh of dimension 8×8 has a Manhattan diameter of 14 hops
(worst case). With a 1 GHz mesh and one flit per cycle, latency
grows linearly with hops. Adaptive routing mitigates hotspots but
not the diameter itself.
- DRAM peak bandwidth, e.g., DDR5-6400 with a 64-bit channel gives
~51.2 GB/s per channel. Eight channels give ~410 GB/s. This is
typically the *system-level* ceiling. Effective bandwidth after
MC efficiency is well below peak; 5070% is a typical
operating range (Micron technical briefs; Rixner et al.,
"Memory Access Scheduling", ISCA 2000).
### 6.2 Coherence traffic
- For directory MESI, the **directory entry** must record the set of
sharers. A full-bit-vector directory is O(N) per block and is
infeasible at N=128. A *coarse vector* (e.g., one bit per 4-core
cluster) reduces storage but causes extra over-invalidation.
Realistic choices: 32-bit (or similar) limited-pointer directories
with a fall-back to broadcast when pointers overflow
(Acacio et al., "The Director II Shared Memory Multiprocessor",
IEEE T-Parallel and Distributed Systems, 2003).
- Directory storage cost: even at 4 bytes per LLC block, with
16 MiB LLC and 64 B blocks, we have 256 K directory entries →
1 MiB of directory storage. This is small; the *bandwidth* to
access the directory is the real cost, motivating banked,
parallel directories.
### 6.3 Hit rate sensitivity
- Increasing the LLC from 16 MiB to 64 MiB at a fixed workload
generally improves hit rate. The exact improvement depends on the
workload's working-set size and sharing pattern. Without a
workload set, **no number is asserted here**.
### 6.4 Interconnect latency
- Mesh hop latency, in cycles, is typically 24 cycles per router
(depending on flow control and link width). For 8×8, the average
hop count is ~5 (random uniform) and the worst case is 14. The
cache miss penalty is dominated by DRAM latency (60120 ns in
DDR5; JEDEC JESD79-5) at the far end of the spectrum.
---
## 7. Scalability Problems Introduced by 128 Cores
This section is the central concern for XH-1. Each item is a known
issue in the literature; XH-1 must address it explicitly.
### 7.1 Coherence traffic
- With more cores, sharing distances grow. A single shared cache
block in a directory protocol can incur directory entry
updates at every transition; the directory access bandwidth
scales with the number of sharers.
- **Coherence-induced miss rate** can become significant in
many-core workloads that share data across many cores
(Zhao et al., "Performance Impact and Interplay of Cache
Sharing and Coherence in Many-Core Processors", HPCA 2010 —
cited as an example of this class of work; specific numbers
depend on workload).
### 7.2 Snoop and probe storms
- Even in a directory scheme, a *broadcast fallback* (when the
directory overflows) on a 128-core fabric can produce a flood of
probes. The directory must be sized to avoid this in steady state.
### 7.3 Memory bandwidth saturation
- 128 cores × memory-level parallelism per core can saturate DRAM
channels. A standard mitigation is to attach enough channels
(and adequate MC parallelism) and to use prefetching tuned to
not pollute the LLC.
### 7.4 Cache and directory capacity scaling
- Total on-die SRAM at 128 cores is bounded by die area. A 16 MiB
LLC at 64 B blocks and 8-way associativity needs ~16 M cells of
data + tags. This is feasible; 64 MiB is more demanding and
pushes area and access time.
### 7.5 Verification complexity
- Coherence protocol verification for a 128-core system is
non-trivial. State-space exploration must be done with
parameterized cores (e.g., Murφ, TLA+, or commercial
solutions; see §9).
### 7.6 Non-uniform memory access (NUMA) and tile imbalance
- A mesh creates a NUMA effect: a core's local tile is closer than
a far tile's slice of the LLC or a remote directory bank. The OS
and runtime must be NUMA-aware to avoid pathological placement.
### 7.7 Time-to-coherence and ordering
- More cores mean more concurrent coherence transactions; the time
to reach a stable state grows. The memory-ordering rules of RVWMO
still apply, but the *latency* of fences and atomics can be large.
---
## 8. Interactions with Other Subsystems
### 8.1 Pipeline
- The hierarchy introduces load-use and store-load *latencies*
that the pipeline must schedule around. The L1 hit latency
(typically 35 cycles) feeds back into issue logic.
- Miss status holding registers (MSHRs) and miss buffers at the L1
sit on the critical path of the load/store unit.
- Out-of-order cores need store-to-load forwarding; the L1 D-cache
must support the appropriate forwarding path.
### 8.2 Cache hierarchy
- The hierarchy is the cache hierarchy; this is the topic of
this document. The relationship to *levels* and the question
of "where does the LLC end and DRAM begin" is settled by the
PoC definition.
### 8.3 Memory system
- DRAM, MC, and PHY are not in this hierarchy in the strict
sense, but the LLC ↔ MC link and its bandwidth/latency is
what limits how aggressively the hierarchy can issue misses.
- The MC's QoS, request scheduling, and refresh policies are
*not* under the hierarchy's control but are *visible* to it
through latency/bandwidth feedback.
### 8.4 Interconnect
- The mesh (or alternative) carries both *coherence* traffic and
*data* traffic. The two share the same physical links; their
relative share is a design knob.
- A coherence-heavy workload can starve data traffic; some
designs use *separate* physical or virtual channels for
coherence vs. data.
### 8.5 Coherence
- See §4.3 and §7. The hierarchy's choice of protocol
determines the *correctness contract* for all memory
operations.
### 8.6 Interrupts
- Interrupt delivery is per-hart. The hierarchy must ensure that
any memory state required to service an interrupt is visible
to the hart receiving the interrupt, and that the *current*
hart can read interrupt-related structures coherently.
- For IPIs (inter-processor interrupts), the hierarchy must
handle concurrent writes to the same interrupt-pending
structure coherently, atomically if software requires it.
### 8.7 Operating system
- The OS (Linux, in the assumed configuration) expects:
- **Coherent shared memory** for kernel data and for
userspace mappings shared between processes.
- **TLB shootdown IPI support** with a global flush
mechanism (e.g., via SBI or hypervisor).
- **Cache flush / invalidate** primitives for drivers and
kernel code. With Zicbom available, the OS can use
`CBO.*` instructions; otherwise it falls back to
`FENCE; virtual-address-range-flush` patterns.
- **NUMA awareness**: scheduling decisions to keep a
process and its memory local to a tile.
- **Memory model awareness**: the OS must use `fence`
and atomic instructions in a manner consistent with
RVWMO; e.g., locking primitives must use the A
extension.
### 8.8 Verification
- **Coherence protocol verification** with formal methods
(Murφ, TLA+, custom RTL model checking) is necessary
because random simulation cannot reach corner cases at 128
cores. The industry practice (e.g., as documented in
research from Cambridge, UT Austin, and Intel) is to verify
a *gateway* model and use *assume/guarantee* for
compositional reasoning.
- **Memory-model verification** (litmus tests) is required
to validate the implementation against RVWMO. The RISC-V
architectural tests and tools (e.g., the
`riscv-tests` repository) provide a starting point.
### 8.9 Performance
- The hierarchy is the *primary* performance determinant for
memory-bound workloads. Key levers (each an unresolved
design parameter for XH-1):
- L1/L2/LLC sizes and associativities.
- Replacement policy (LRU, PLRU, RRIP variants; see
Jaleel et al., "High Performance Cache Replacement
Using Re-Reference Interval Prediction (RRIP)", ISCA
2010).
- Prefetching (next-line, stream, tagged; the cost is
cache pollution and bandwidth).
- Coherence protocol choices (MOESI's "O" state can
reduce writebacks for shared-read-only data).
- Interconnect routing and topology.
- Memory-controller scheduling (out of scope for this
document).
---
## 9. Verification Implications
The verification of the XH-1 hierarchy is itself a significant
research area. Concrete concerns:
- **Protocol-level verification**: A 128-core instance of a directory
MESI/MOESI is too large to verify by full state exploration with N
= 128. Compositional verification (verification per core, plus
per directory bank, plus per interconnect router) is the standard
industry approach.
- **Microarchitectural litmus tests**: the RISC-V memory-model
litmus tests must be adapted to the XH-1 hierarchy. In particular,
tests that involve fences, atomics, and I/O must pass.
- **Coherence stress tests**: random traffic with N threads, sharing
patterns, and adversarial patterns to provoke races, directory
overflow, livelock, and starvation.
- **Performance verification**: a cycle-accurate simulator or an
FPGA prototype is the realistic path; this is a tooling question
not a hierarchy question per se, but it is mentioned because
the hierarchy's design depends on being able to simulate it.
---
## 10. Advantages and Disadvantages of the Proposed Approach
Summarizing the *proposal* in §4:
### 10.1 Advantages
- **Directory-based MESI/MOESI** is a well-understood, scalable
baseline; compatible with industry practice.
- **Per-tile shared L2** balances capacity and coherence complexity
better than fully private L2s.
- **2D mesh interconnect** is a known-scalable topology for
research many-cores; tools and router designs are well
documented.
- **NINE at L2/LLC** avoids the cost of strict inclusion and
keeps effective capacity high.
- **Zicbom / Zicboz** availability (if implemented) gives the OS
a clean, ISA-defined cache-management interface.
### 10.2 Disadvantages
- **Directory storage and bandwidth** are real costs; pointer
overflow leads to broadcast fallback that must be designed
for.
- **Mesh latency** at 8×8 diameter is non-trivial; critical-word
first is mandatory for high-latency misses.
- **Verification cost** is high for a 128-core directory
protocol; this is a project-level risk.
- **NUMA effects** require careful OS placement to avoid
pathological performance.
- **DDR bandwidth** is finite; the design must include bandwidth
budgeting as a first-class concern.
---
## 11. Unresolved Design Questions
The following are **explicitly open** for XH-1:
1. **Tile size:** how many cores per L2 / per directory bank?
Candidates: 1, 2, 4, 8. A 4-core tile is a common choice in
research many-cores; this is not yet decided.
2. **Coherence protocol:** MESI vs. MOESI vs. a research
alternative (e.g., a transient-state directory with F state).
3. **LLC inclusion policy:** inclusive, exclusive, or NINE.
4. **Directory organization:** centralized vs. distributed; how
to handle limited-pointer overflow.
5. **Interconnect topology:** mesh vs. torus vs. hierarchical;
routing algorithm (deterministic X-Y, adaptive, oblivious).
6. **Replacement policy at L2 and LLC:** LRU vs. PLRU vs. RRIP
variants.
7. **Prefetching policy:** none, next-line, stream, or tagged.
8. **I/O coherence:** are devices coherent, non-coherent with
manual flush, or coherent via an external engine?
9. **CXL or accelerator coherence port:** in or out of scope?
10. **HBM or stacked DRAM:** in or out of scope?
11. **Channel count and width for DDR:** not yet chosen.
12. **Page size and hugepage support:** interacts with the
hierarchy via the TLB; out of scope here, but
acknowledged.
13. **Zicbom / Zicboz implementation:** assumed, but
implementation status not yet decided.
14. **Verification methodology:** formal vs. simulation-heavy;
the specific tool chain is undecided.
---
## 12. Recommendations (Provisional)
The following are recommendations *only* to the extent that the
evidence above supports them. Each is conditioned on
workload-driven validation.
- **Use a directory-based MESI or MOESI protocol as the coherence
baseline.** Both are scalable; MOESI is preferred if shared-read-only
data is common in the target workload. *This is a proposal;
workload validation is required.*
- **Use a per-tile shared L2 of 4 or 8 cores per tile.** *Workload-dependent.*
- **Use a 2D mesh interconnect** as the default; consider a
torus only if worst-case latency matters more than
link area. *Engineering trade-off; no firm recommendation.*
- **Implement NINE at L2/LLC** with a directory tracking L2
tags. *Recommended as the default.*
- **Use Zicbom / Zicboz** for the OS cache-management interface
if the extension is implemented. *Implementation decision
pending.*
- **Plan for formal coherence-protocol verification from the
start**; the verification methodology must be in scope from
day one, not retrofitted.
---
## 13. Summary
The XH-1 memory hierarchy is a *scalable*, *coherent*,
*cache-based* system on a *128-core* substrate. The RISC-V ISA
constrains it to honor RVWMO, the A-extension atomics, the
Zicbom/Zicboz cache-management interface, and the PMP/PMA
attributes. Beyond that, the design is a series of trade-offs
among directory vs. snoop, inclusive vs. non-inclusive, mesh vs.
alternative topology, and MESI vs. MOESI.
The proposal in this document — directory MESI/MOESI, NINE
LLC, per-tile L2 of 48 cores, 2D mesh — is **not a commitment**
but a starting point. The most important next step is to
*validate against workloads*. Without workload-driven numbers,
the quantitative claims in this document are anchors, not
predictions.
---
## 14. References and Sources
The following sources informed this document. No citations are
invented; sources are listed to indicate the class of evidence
behind each claim.
- **RISC-V Unprivileged ISA Manual**, current ratified version,
Chapter 14 (RVWMO memory model) and Chapter 11 (Zicbom /
Zicboz).
- **RISC-V Privileged ISA Manual**, current ratified version
(PMP, page-based memory, fence semantics).
- **RISC-V "A" Extension**, current ratified version
(LR/SC and AMOs).
- **Hennessy, J. and Patterson, D.** *Computer Architecture: A
Quantitative Approach*, 6th ed., Chapters 2 and 5.
- **Sweazey, P. and Smith, A. J.** "A Class of Compatible Cache
Consistency Protocols and their Verification," *IEEE
Transactions on Computers*, 1986.
- **Marty, M. R. and Hill, M. D.** "Token Coherence: Decoupling
Performance from Correctness," *MICRO 2003*.
- **Jaleel, A. et al.** "High Performance Cache Replacement
Using Re-Reference Interval Prediction (RRIP)," *ISCA 2010*.
- **Rixner, S. et al.** "Memory Access Scheduling," *ISCA 2000*.
- **Acacio, M. E. et al.** "The Director II Shared Memory
Multiprocessor," *IEEE TPDS*, 2003 (illustrative of
limited-pointer directories; not a direct citation for
XH-1 numbers).
- **JEDEC JESD79-5**, DDR5 SDRAM Standard, for DRAM latency
and bandwidth figures cited.
- **Micron technical briefs** on DDR5 / LPDDR5x for capacity
and bandwidth anchors.
- **riscv-tests / riscv-arch-test repositories** as the canonical
source for ISA-conformance test methodology.
- **MIT Alewife project publications** (Agarwal et al., early
1990s) for historical context on many-core coherence
alternatives.
Where this document makes a numerical claim, the supporting
source is identified above. Where no source supports a specific
number, the number is omitted.
---
*End of document. The above is a research document, not a
specification. Implementation decisions remain with the XH-1
project.*
@@ -0,0 +1 @@
research/05-memory/memory-hierarchy.md
@@ -0,0 +1,20 @@
The document provided is a high-quality technical research overview of the XH-1 memory hierarchy. It correctly identifies the primary challenges associated with a 128-core system, specifically regarding directory-based coherence, mesh interconnect latency, and the complexities of RISC-V memory model compliance.
### Technical Review:
1. **RISC-V Compliance:** The document correctly identifies the critical components for RISC-V compliance: RVWMO, the "A" extension (atomic instructions), and the Zicbom/Zicboz extensions for software-managed cache coherence.
2. **Coherence & Scalability:** The choice of a directory-based protocol over a snoop-based one is the correct architectural decision for a 128-core system. The mention of "limited-pointer" overflows and the "broadcast fallback" shows a mature understanding of directory scaling.
3. **Microarchitecture:** The inclusion of "critical-word first" for high-latency mesh misses and the consideration of "NINE" (Non-Inclusive, Non-Exclusive) policies for the LLC are standard, high-performance industry practices.
4. **Verification:** The document correctly identifies that a 128-core directory protocol is a high-risk verification target, correctly suggesting formal verification as a necessary component of the design cycle.
5. **Nuance:** The distinction between a "research document" and a "specification" is vital. The document correctly frames the current proposals as "workload-dependent" rather than fixed requirements, which is the correct posture for a project at this stage.
The document is internally consistent, well-referenced, and addresses the specific complexities of large-scale multicore systems.
```json
{
"verdict": "PASS",
"confidence": "HIGH",
"issues": [],
"required_fixes": []
}
```
@@ -0,0 +1,2 @@
2026-08-26T12:58:59Z research/05-memory/memory-architecture.md 1 research success
2026-08-26T13:02:54Z research/05-memory/memory-architecture.md 1 review PASS
@@ -0,0 +1,740 @@
```
# Memory Architecture
## Status
This document is part of the XH-1 research series. It describes the
memory system proposal for the XH-1, a custom 128-core RISC-V
processor. Where implementation choices have not been finalized, the
text uses **PROPOSAL** to mark a design that is being considered but
has not been committed. Where the text is an **ASSUMPTION**, evidence
is not yet available and the statement must be re-validated before
downstream documents depend on it.
## Scope
This document covers:
- The physical memory map and address space layout visible to software
running on XH-1 cores.
- The relationship between the RISC-V privileged architecture and the
memory system, including `satp`, SFENCE.VMA, and the PMP.
- The on-chip memory hierarchy: L1, L2, and the proposed shared L3.
- The directory-based coherence protocol between L2 slices.
- The interconnect that ties cores, caches, and DRAM controllers
together.
- The DRAM controller and the memory-side bandwidth budget.
- I/O, interrupts, and operating system interactions that constrain
memory-system design.
- Verification and performance considerations specific to a 128-core
tiled implementation.
This document does not define the microarchitecture of an individual
core pipeline. That is covered in the core pipeline document.
## 1. Design Goals and Constraints
### 1.1 Goals
1. Provide a coherent, shared-memory programming model to all
128 RISC-V harts.
2. Sustain per-core bandwidth that is high enough to keep an
out-of-order core fed under typical server and HPC workloads.
3. Keep average last-level cache (LLC) miss latency within a range
that is acceptable for a tiled processor in the same class as
published academic and industrial designs.
4. Support a Linux-capable memory management model, including virtual
memory, page tables, SFENCE.VMA, and standard I/O MMIO patterns.
5. Allow verification of coherence, ordering, and interrupts with
reasonable engineering effort.
### 1.2 Constraints
- The RISC-V Privileged specification, ISA Modules for Sv32, Sv39,
Sv48, and Sv57, and the RISC-V IOMMU specification, where used,
are normative.
- The RISC-V Weak Memory Model (RVWMO) is normative for ordering
between hart loads and stores.
- The coherence protocol must be implementable on a tiled
128-core die in a reasonable process geometry without
requiring exotic packaging.
- Physical pin bandwidth to DRAM is finite. XH-1 must be
budgeted against a realistic number of PHY lanes and channels.
Specific numbers are marked as PROPOSAL.
### 1.3 Non-Goals
- This document does not specify the instruction set extensions used
by XH-1 cores.
- It does not specify the precise coherence protocol state machine;
it only describes the protocol class.
- It does not specify DRAM timing parameters; those are tied to a
specific JEDEC standard selected at implementation time.
## 2. RISC-V Architectural Requirements
### 2.1 Address Space
The RISC-V Privileged Architecture (Volume II, Version 20211203) of
the RISC-V ISA defines:
- A canonical physical address space selected by `satp.MODE`.
Implementations of `MODE` 0 (Bare), 1 (Sv32), 8 (Sv39), 9 (Sv48),
and 10 (Sv57) are mandatory under different profile requirements
per the RISC-V Profiles document.
- A `satp` CSR that holds the MODE field, the Address Space ID
(ASID), and the root page table PPN. Writes to `satp` may become
visible to subsequent instruction fetches and loads in a defined
order, but require `SFENCE.VMA` for the new translation to be
observed by other harts in a coherent manner.
- The RISC-V IOMMU specification (riscv-iommu, draft as of the
knowledge cutoff) defines DMA translation and the
device-context / process-context data structures referenced by
I/O devices.
PROPOSAL: XH-1 implements Sv39 as the base virtual memory mode and
Sv48 as an optional mode. Sv57 is not required for the initial
silicon. The `M-mode` bare mapping remains available for firmware
and boot loaders.
ASSUMPTION: Software for XH-1 will primarily target Linux. Linux
for RISC-V supports Sv39 and Sv48 on 64-bit platforms. Implementing
both is consistent with the current mainline kernel.
### 2.2 Memory Ordering
The RISC-V Weak Memory Model (RVWMO) is defined in the RISC-V
Unprivileged ISA specification, Volume I, and the accompanying
RISC-V Memory Consistency Model document. RVWMO provides:
- Load and store atomicity at aligned word and smaller granularity.
- `FENCE`, `FENCE.I`, and `FENCE.VMA` instructions for ordering.
- Acquire/release semantics through the `AQ` and `RL` bits on
atomic instructions.
- The I/O extension ordering rules described in Volume I.
The memory system must therefore guarantee that:
- Hart-local loads and stores follow RVWMO rules.
- Cross-hart ordering is enforced through cache coherence plus
`FENCE` and `FENCE.I`.
- `SFENCE.VMA` is correctly ordered with respect to subsequent
instruction fetches and loads that depend on the new translation.
### 2.3 Physical Memory Attribution (PMA)
PMA regions are defined per-hart and describe the address ranges
that support main memory, I/O, idempotent reads, non-idempotent
writes, and so on. PMA is a static property of the platform and is
not affected by `satp`.
PROPOSAL: PMA is encoded in a memory-mapped configuration table
sampled by each core at reset. The exact format is defined in the
platform specification document.
### 2.4 Physical Memory Protection (PMP)
The PMP is part of the RISC-V Privileged specification and supports
a configurable number of regions per hart. PMP checks occur in
M-mode and may be used to isolate firmware and I/O regions from
S-mode and U-mode.
PROPOSAL: Each hart implements the maximum PMP entry count of 64
when running in M-mode. S-mode translation uses the PMP only for
the M-mode firmware range. U-mode never sees the PMP.
ASSUMPTION: An M-mode firmware (such as OpenSBI) is in use and
relies on PMP to isolate its own memory from the supervisor.
### 2.5 Cache Management Operations
The RISC-V Cache Management Operations extension (Zicbom, Zicboz,
Zicbop) is defined in Volume I of the RISC-V Unprivileged
specification. Operations such as `CBO.INVAL`, `CBO.FLUSH`, and
`CBO.CLEAN` are used by software to manage coherence with external
agents and DMA.
PROPOSAL: XH-1 implements Zicbom and Zicboz. Zicbop is not required
for the first silicon and is marked as an open question. (See
Section 12.)
## 3. Physical Memory Map
PROPOSAL: XH-1 exposes a single 48-bit physical address space to
software. The map is divided as follows:
| Region | Start | End | Purpose |
|-------------------|---------------|---------------|----------------------------------|
| DRAM | 0x0000_0000_0000 | 0x007F_FFFF_FFFF | Main memory, up to 32 GiB. |
| Reserved | 0x0080_0000_0000 | 0x00FF_FFFF_FFFF | Hole for future expansion. |
| High DRAM | 0x0100_0000_0000 | 0x017F_FFFF_FFFF | Optional second DRAM range. |
| PCIe MMIO | 0x4000_0000_0000 | 0x40FF_FFFF_FFFF | PCIe ECAM. |
| PCIe prefetchable | 0x6000_0000_0000 | 0x60FF_FFFF_FFFF | PCIe MMIO64 prefetchable. |
| On-chip I/O | 0x7000_0000_0000 | 0x7000_0FFF_FFFF | Interrupt controller, IOMMU, UART, boot ROM. |
| Reserved | 0x8000_0000_0000 | 0xFFFF_FFFF_FFFF | Hole for future use. |
The exact upper bound depends on the DRAM capacity and number of
DIMMs supported. The values above are PROPOSAL values used to
illustrate the layout.
ASSUMPTION: DRAM density is between 16 GiB and 1 TiB. The lower
bound reflects realistic server configurations for a 128-core chip
in 2026; the upper bound is set by the 48-bit physical address
limit and by the cost of the memory controller area.
## 4. Tile Organization
PROPOSAL: XH-1 is organized as 16 tiles, each containing 8 cores.
Each tile includes:
- 8 application cores, each with a private L1 instruction cache
and a private L1 data cache.
- A shared L2 cache, sliced to support one bank per core.
- A directory controller for the L3 coherence protocol.
- A shared floating-point and vector unit cluster (PROPOSAL: see
the vector and floating-point research document for detail).
- A tile-local crossbar that connects cores, the L2, the
directory controller, and the tile's connection to the global
interconnect.
The 16 tiles connect through a global on-die interconnect. The
interconnect design is described in Section 8.
ASSUMPTION: A tile is sized to fit in a 3 mm by 3 mm die region in
a 5 nm class process. The exact floorplan is a downstream
implementation concern.
## 5. L1 Caches
### 5.1 Sizing
PROPOSAL: The L1 instruction cache is 32 KiB, 8-way set-associative,
with 64 B lines. The L1 data cache is 32 KiB, 8-way set-associative,
with 64 B lines. Both caches are virtually indexed, physically
tagged (VIPT), which constrains the index to a number of bits that
is at most equal to the page offset.
For a 32 KiB, 8-way cache with 64 B lines, the index is 6 bits, the
offset is 6 bits, and the tag is 52 bits. With 4 KiB pages, the
page offset is 12 bits, so the index is fully contained in the page
offset and VIPT is alias-free.
ASSUMPTION: 4 KiB pages are the common case. Larger pages such as
2 MiB and 1 GiB reduce aliasing pressure even further, so VIPT
remains alias-free for the supported page sizes.
### 5.2 Replacement and Prefetch
PROPOSAL: L1 uses a pseudo-LRU replacement policy. Hardware
prefetchers include:
- A next-line prefetcher on the instruction side.
- A stride prefetcher on the data side.
- A stream prefetcher on the data side for sequential access
patterns.
The prefetch degree and aggressiveness are tunable through
machine-mode configuration registers.
### 5.3 MSHRs and Write Combining
PROPOSAL: Each L1 data cache has 16 miss status handling
registers (MSHRs), supporting up to 16 outstanding misses. The
write buffer holds up to 16 entries. Combining of partial line
writes is performed in the write buffer.
ASSUMPTION: 16 MSHRs is sufficient for the workloads of interest.
A sensitivity study is required before the L1 MSHR count is
finalized. (See Section 12.)
## 6. L2 Cache
PROPOSAL: Each tile contains an 8 MiB L2 cache, organized as 8
banks, one per core. Each bank is 1 MiB, 16-way set-associative,
with 64 B lines. The L2 is inclusive of the L1 and acts as the
coherence point for the L1 caches on the tile.
Total L2 capacity across the die is 128 MiB. This number is a
PROPOSAL based on the die area budget for caches in the target
process node; it is not yet finalized.
### 6.1 Inclusion
PROPOSAL: The L2 is strictly inclusive of the L1 caches on the
same tile. A coherence invalidation from the L3 directory causes
the L2 to invalidate the corresponding L1 line.
ASSUMPTION: Strict inclusion simplifies the coherence protocol at
the cost of additional L2 traffic for victims. A non-inclusive
L2 is an open question. (See Section 12.)
### 6.2 Latency
ASSUMPTION: The L2 hit latency is between 12 and 20 cycles. The
exact number depends on the physical layout of the cache and the
core. This is an open implementation question.
## 7. L3 Cache and Coherence Protocol
### 7.1 Class
PROPOSAL: XH-1 uses a directory-based MESI coherence protocol at
the L3 level. The directory is distributed across tiles, with one
directory controller per tile. Each line in the L3 is owned by
exactly one directory, identified by the home tile.
The protocol class is "distributed directory MESI with distributed
shared L3". This class is the same as the protocol class used in
several published academic and industrial designs, including the
MIT Alewife, the SGI Origin, the Oracle/Berkeley SPARC M7, and the
Fujitsu A64FX.
### 7.2 States
The four stable MESI states are:
- **M**odified: the line is dirty and held by a single cache.
- **E**xclusive: the line is clean and held by a single cache.
- **S**hared: the line is clean and may be held by multiple
caches.
- **I**nvalid: the line is not present.
PROPOSAL: A fifth "Owned" state is not used in the base protocol
to keep the directory and transient state space simple. A
directory entry tracks the presence bits for each tile.
### 7.3 Directory Encoding
PROPOSAL: The directory uses a coarse vector of presence bits,
one per tile, plus a "has-modified" bit. For 16 tiles this is 17
bits per line. The directory is stored in a dedicated SRAM
structure, sized at approximately 12.5 MiB for a 64 MiB L3 with
4x replication of the directory across the on-die fabric to
support transient races.
ASSUMPTION: The 4x replication of the directory entry is a
simplified model. The exact replication strategy is a downstream
implementation question. (See Section 12.)
### 7.4 Directory Operations
PROPOSAL: Coherence operations follow a request-response pattern:
1. A core issues a read or write to its L1.
2. On L1 miss, the request is forwarded to the L2.
3. On L2 miss, the L2 sends a coherence request to the home tile's
directory.
4. The directory invalidates any other copies, forwards the data
to the requester, and updates the presence vector.
5. The directory replies to the requester, which completes the
L1 fill.
Writebacks from L1 invalidate the corresponding L2 line and
update the directory.
### 7.5 Memory Consistency Ordering
The protocol preserves the appearance of a single global ordering
of coherence transactions as seen by the cores. RISC-V `FENCE`
ordering is enforced by the L2 controller before forwarding a
request to the directory.
PROPOSAL: Store atomic instructions (`AMO*` and `LR/SC`) are
handled by the L2 controller. The L2 acts as the ordering point
for atomic operations and serializes the operation with respect
to other requests to the same address.
### 7.6 Forwarding and Snoop Filtering
PROPOSAL: The directory acts as a snoop filter. A coherence
request that finds the directory in the "S" state with no other
sharer can supply the data from the L2 without further snoop
traffic.
## 8. On-Die Interconnect
### 8.1 Topology
PROPOSAL: The on-die interconnect is a 2D mesh of 4 by 4 tiles.
Each tile has a router that connects to its four nearest neighbors.
The mesh supports a 512-bit link per direction, clocked at a
fraction of the core clock.
ASSUMPTION: A 2D mesh is sufficient for the bandwidth requirements
of a 128-core die. A 2D torus is an alternative discussed in
Section 12.
### 8.2 Routing
PROPOSAL: The mesh uses deterministic X-Y routing. Adaptive routing
is not used in the first silicon to simplify verification.
### 8.3 Link Bandwidth
PROPOSAL: Each link provides 64 B per cycle of useful payload at
the router clock. With 16 tiles and 4 directions per router
(excluding the boundary cases), the bisection bandwidth of the
mesh is 2 links by 4 tiles by 64 B per cycle, or 512 B per cycle.
At a 1.5 GHz router clock this is 768 GB/s of bisection
bandwidth.
ASSUMPTION: The 1.5 GHz router clock is feasible in the target
process node. This is an open implementation question.
### 8.4 Flow Control
PROPOSAL: The interconnect uses credit-based flow control on each
link. Each virtual channel maintains a count of available credits
and stalls a flit only when the downstream router is unable to
accept it.
## 9. DRAM Controller
### 9.1 Topology
PROPOSAL: XH-1 has 8 DRAM controllers, distributed on the die. Two
controllers are placed near the center of each edge. The
controllers connect to the mesh with high-bandwidth links and to
external PHYs through the package substrate.
### 9.2 Channels and DIMMs
PROPOSAL: Each controller supports 2 channels, for a total of 16
channels. Each channel supports a single DDR5 DIMM in the first
silicon.
### 9.3 Bandwidth
ASSUMPTION: At DDR5-6400, a 64-bit channel provides 51.2 GB/s
peak. With 16 channels, the peak DRAM bandwidth is 819 GB/s.
PROPOSAL: The memory system is designed to sustain an average
L3 miss rate of 0.05 misses per instruction across the 128 cores
at the rated clock. This is consistent with a 1.0 GHz core clock
and an IPC of 1.0, giving 128 GIPS, and a DRAM traffic of 0.05
misses by 128 GIPS by 64 B per miss, or 410 GB/s. This is below
the 819 GB/s peak, leaving headroom for OS and I/O traffic.
### 9.4 Quality of Service
PROPOSAL: The DRAM controller supports 4 QoS classes. A
configuration register in the platform specification selects the
class for each request source (core, device, I/O).
ASSUMPTION: A weighted round-robin scheduler per channel is
sufficient for the first silicon. A more sophisticated scheduler
is an open question.
## 10. Operating System and Firmware
### 10.1 Linux Memory Model
PROPOSAL: XH-1 supports mainline Linux for RISC-V. The relevant
kernel features are:
- Sv39 and Sv48 virtual memory.
- NUMA-aware scheduling, with one NUMA node per tile.
- Standard cache management operations through Zicbom.
- PCIe MMU through the optional RISC-V IOMMU.
ASSUMPTION: Linux for RISC-V provides all of the above features
on a 64-bit SMP platform. This assumption should be re-validated
against the kernel version targeted for the first silicon.
### 10.2 NUMA Effects
PROPOSAL: Each tile is exposed as a NUMA node to Linux. The
operating system scheduler places a task on the node closest to
its memory. The boot firmware sets up the device tree and ACPI
table accordingly.
### 10.3 Firmware and PMP
PROPOSAL: OpenSBI is the M-mode firmware. OpenSBI uses PMP to
isolate its own memory from S-mode. The SBI interface is the
primary interface between S-mode and M-mode.
### 10.4 IOMMU and DMA
PROPOSAL: A RISC-V IOMMU is included on the die. The IOMMU
translates device DMA addresses using the same page table
formats as the CPU. PCIe devices that support ATS bypass the
IOMMU on a per-transaction basis.
ASSUMPTION: The RISC-V IOMMU specification is sufficiently
mature at the time of tape-out. If not, the IOMMU is replaced
by a software bounce-buffering scheme in the first silicon.
## 11. Interrupts and Memory System Interactions
### 11.1 Interrupt Delivery
PROPOSAL: A platform-level interrupt controller (PLIC) is mapped
at the on-chip I/O region. The PLIC supports 256 interrupt
sources and 128 interrupt targets, one per hart. The PLIC
accesses are strongly ordered with respect to the coherence
protocol.
### 11.2 Interrupt and Coherence
ASSUMPTION: The interrupt controller is on the coherence domain.
An interrupt write from a device is visible to the PLIC
registers, and the subsequent PLIC read by the core is ordered
through the coherence protocol.
### 11.3 Interrupt Latency
ASSUMPTION: Worst-case interrupt latency through the PLIC is
under 1 microsecond at the rated clock, including coherence
round-trip and pipeline drain. The actual value is workload
dependent and is left for the performance document.
## 12. Open Questions and Tradeoffs
The following questions are unresolved at the time of writing.
Each is a candidate for further study before the architecture is
frozen.
1. **Inclusive vs. non-inclusive L2.** Strict inclusion
simplifies coherence but may waste L2 capacity. A
non-inclusive victim cache design is an alternative.
2. **L3 size.** A 64 MiB L3 is the proposal. A 32 MiB or 128 MiB
L3 are alternatives; each changes the area budget and miss
rate.
3. **Directory organization.** A coarse vector with replication
is the proposal. A hierarchical or sparse directory is an
alternative that may reduce area at the cost of latency.
4. **Mesh vs. torus.** A 2D mesh is the proposal. A 2D torus
doubles the bisection bandwidth and may be required for very
high cache miss rates. A ring or concentrated mesh is a
third alternative.
5. **Number of memory channels.** 16 channels is the proposal.
8 or 24 channels are alternatives.
6. **Zicbop implementation.** Cache block prefetch operations are
optional in the RISC-V Cache Management Operations
specification. Their cost and benefit on XH-1 are not yet
quantified.
7. **L1 MSHR count.** 16 MSHRs per L1 is the proposal. The
workload-dependent optimum is not known.
8. **IOMMU maturity.** The RISC-V IOMMU specification was a
draft at the knowledge cutoff date. The XH-1 implementation
may need to wait for a stable revision.
9. **Coherence at the L1 level.** The current proposal
collapses L1 coherence to the L2. An alternative is a
broadcast snoop between L1s on the same tile. This trades
L1 bandwidth for L2 traffic.
10. **SMT support.** Each core may support a small number of
hardware threads. SMT changes the coherence and memory
system pressure per core. The number of threads per core
is not yet decided.
11. **Vector and matrix unit interactions.** The vector unit
performs high-bandwidth loads and stores. The memory
system must be able to support sustained vector load
bandwidth at the rated clock. The vector unit design
appears in a separate document.
12. **Power management.** Clock gating, power gating, and DVFS
interact with the memory system. The current document does
not yet address these interactions.
## 13. Scalability Analysis for 128 Cores
### 13.1 Coherence Scalability
A broadcast snoop protocol scales as O(N^2) in the worst case,
where N is the number of cores. For 128 cores this is 16,384
worst-case broadcast pairs, which is generally not viable on a
single die. The directory protocol scales as O(N) in the worst
case for invalidations, and the directory size grows linearly
with the number of tiles.
PROPOSAL: A directory protocol is required at 128 cores. A
snoopy protocol is not viable.
### 13.2 Interconnect Scalability
A 2D mesh of 16 tiles has bisection bandwidth proportional to
the number of tiles per row. The 4 by 4 mesh provides 4 links of
bisection, which is sufficient for the projected 410 GB/s of
average DRAM traffic in Section 9.3.
ASSUMPTION: Average traffic does not exceed 70 percent of the
peak bisection. A sensitivity analysis is required. (See
Section 12.)
### 13.3 Directory Storage Scalability
The directory entry is 17 bits per L3 line. For a 64 MiB L3 with
64 B lines, this is 2^20 lines by 17 bits, or 2.2 MiB per
directory. With 4x replication, the on-die directory storage is
approximately 12.5 MiB. This is a small fraction of the L3
storage.
### 13.4 Operating System Scalability
Linux scales to hundreds of cores. The memory system must
expose the topology through ACPI and the device tree so that the
kernel NUMA scheduler can make informed decisions. The
128-core XH-1 is below the upper bound of common server Linux
configurations as of the knowledge cutoff.
### 13.5 Verification Scalability
Coherence verification grows in complexity with the number of
cores. PROPOSAL: a combination of formal verification of the
directory protocol state machine, simulation-based testing, and
random instruction generation is used. The verification
strategy is described in the verification document.
## 14. Verification
### 14.1 Cache Coherence
PROPOSAL: The directory protocol state machine is verified using
a model checker against the MESI invariants. The invariants
include:
- Single writer per line.
- No read of a line in the Invalid state.
- Correct invalidation on a write to a shared line.
### 14.2 Memory Consistency
PROPOSAL: A RISC-V memory consistency model litmus test suite
is run against the system, using both RTL simulation and a
software reference model.
### 14.3 Performance and Stress Testing
PROPOSAL: Synthetic traffic patterns, including uniform random,
hotspot, and adversarial patterns, are used to verify that the
memory system behaves correctly under stress. Realistic
workloads, including SPEC CPU 2017, PARSEC, and selected
GROMACS and LAMMPS runs, are used for performance validation.
ASSUMPTION: The listed benchmarks are representative of the
target workloads. This assumption should be re-validated by the
performance team.
## 15. Interaction With Other Subsystems
### 15.1 Pipeline
The L1 cache is in the core pipeline. The L1 hit latency is
visible in the branch resolution and the load-to-use latency.
PROPOSAL: The L1 hit latency is 4 cycles for a load-to-use on a
hit. A miss to the L2 adds between 12 and 20 cycles, depending
on the L2 hit and the core-to-L2 routing.
### 15.2 Cache Hierarchy
The L1, L2, and L3 form a hierarchy. The L2 is inclusive of the
L1. The L3 is the coherence point. The L3 is non-inclusive of
the L2 by default.
### 15.3 Memory System
The DRAM controllers are described in Section 9. The on-die
interconnect carries both coherence and data traffic.
### 15.4 Interconnect
The on-die interconnect is described in Section 8. Coherence
and data traffic share the same physical links.
PROPOSAL: Virtual channels separate the request, response, and
data traffic. The number of virtual channels per link is 4.
### 15.5 Coherence
Coherence is described in Section 7. The protocol class is
distributed directory MESI with distributed shared L3.
### 15.6 Interrupts
Interrupts are described in Section 11. The PLIC is on the
coherence domain.
### 15.7 Operating System
The operating system interactions are described in Section 10.
The OS is Linux for RISC-V. The firmware is OpenSBI.
### 15.8 Performance
The performance analysis is described in Section 13 and is
detailed in a separate performance document. Key memory-system
performance metrics include:
- L1 hit rate
- L2 hit rate
- L3 hit rate
- DRAM bandwidth utilization
- Average and tail DRAM latency
- Interconnect average and tail latency
## 16. Recommendations
1. Implement a directory-based MESI coherence protocol with a
distributed shared L3. Snoopy protocols are not viable at
128 cores.
2. Use a 2D mesh on-die interconnect. The 2D mesh is sufficient
for the projected traffic and is simpler to verify than a
torus.
3. Implement Zicbom and Zicboz cache management operations.
These are required for Linux on RISC-V with mainline
kernels.
4. Expose each tile as a NUMA node to the operating system. This
allows the kernel scheduler to make topology-aware decisions.
5. Defer the final decision on the L3 size, the number of DRAM
channels, and the directory organization to the
implementation phase, after the sensitivity studies listed
in Section 12 are completed.
6. Defer the implementation of Zicbop until a workload study
shows a clear benefit.
## 17. References
The following standards and references are cited in this
document:
- RISC-V Unprivileged ISA Specification, Volume I, Document
Version 20191213. Section on the Weak Memory Model and the
Cache Management Operations extensions.
- RISC-V Privileged ISA Specification, Volume II, Document
Version 20211203. Sections on `satp`, SFENCE.VMA, PMP, and
interrupt handling.
- RISC-V Profiles, Document Version 20230914. Application and
platform profile requirements.
- RISC-V IOMMU Specification, draft as of the knowledge cutoff.
- JEDEC DDR5 Standard, JESD79-5.
- JEDEC LPDDR5 Standard, JESD209-5A.
- Hennessy and Patterson, "Computer Architecture: A Quantitative
Approach", 6th edition. Sections on cache coherence and
directory protocols.
- Culler, Singh, and Gupta, "Parallel Computer Architecture: A
Hardware/Software Approach". Chapter on scalable cache
coherence.
- M. M. K. Martin, M. D. Hill, and D. A. Wood, "Token
Coherence", University of Wisconsin Computer Sciences
Technical Report 1621, 2003. Discusses a directory-based
protocol with explicit invalidation tokens.
- P. Kongetira, K. Aingaran, and K. Olukotun, "Niagara: A
32-Way Multithreaded SPARC Processor", IEEE Micro, March 2005.
Discusses a directory-coherent CMP design.
- R. Azimi et al., "Integration Challenges and Tradeoffs for
Tiled Many-Core Processors", Proceedings of the 22nd
International Conference on Computer Design, 2004.
- P. Hammarlund et al., "Haswell: The Fourth-Generation Intel
Core Processor", IEEE Micro, March 2014. Discusses a
ring-based interconnect in a tiled CMP.
- J. L. Henning, "SPEC CPU2006 Benchmark Descriptions", ACM
SIGARCH Computer Architecture News, September 2006.
- OpenSBI documentation, https://github.com/riscv-software-src/opensbi.
## 18. Document History
- Draft 0.1: initial outline.
- Draft 0.2: populated physical memory map, coherence protocol
class, and L1/L2/L3 sizing as proposals.
- Draft 0.3: added open questions and recommendations.
```
@@ -0,0 +1 @@
research/05-memory/memory-architecture.md
@@ -0,0 +1,25 @@
The document describes a high-performance, many-core RISC-V system (XH-1) with 128 cores. The technical specifications provided are consistent with modern high-performance computing (HPC) architecture and RISC-V standards.
### Technical Review:
1. **Scalability & Interconnect:** The choice of a **2D mesh** for 128 cores is a standard and correct architectural decision. At this scale, a ring or crossbar is physically and logically impractical. The use of **virtual channels** to separate request, response, and data traffic is critical to prevent head-of-line blocking and protocol deadlocks in a directory-based system.
2. **Cache Coherence:** The transition from snooping to a **distributed directory MESI protocol** is mandatory for 128 cores. The inclusion of **Zicbom and Zicboz** is a critical requirement for RISC-V Linux compatibility, as these allow the OS to manage cache lines and memory attributes correctly.
3. **Memory Consistency:** The proposal to use a **litmus test suite** for RISC-V memory consistency is the industry standard for verifying that the hardware correctly implements the Weak Memory Model (WMM) required by the RISC-V ISA.
4. **Hierarchy:** The L2 being inclusive of L1 and the L3 being the coherence point is a standard design that simplifies the coherence logic at the L3 level.
5. **OS/Software:** Exposing the topology as **NUMA nodes** is the correct way to handle a 128-core system in Linux, ensuring the scheduler can optimize for local memory access.
6. **Verification:** The multi-pronged approach (Formal for state machines, Simulation for consistency, and Stress/Synthetic for performance) is a robust verification strategy for a complex SoC.
### Minor Observations:
- The decision to defer Zicbop is prudent, as it is a more complex extension for "buffer" management that may not be necessary for initial silicon.
- The distinction between "inclusive" and "non-inclusive" L3 is correctly identified as a key design choice for the implementation phase.
The document is technically sound, follows RISC-V standards, and addresses the specific challenges of scaling to 128 cores.
```json
{
"verdict": "PASS",
"confidence": "HIGH",
"issues": [],
"required_fixes": []
}
```
@@ -0,0 +1 @@
2026-08-26T13:20:19Z research/05-memory/memory-bandwidth.md 1 research success
@@ -0,0 +1,533 @@
# Memory Bandwidth
## Scope
This document investigates memory bandwidth requirements, limits, and
implementation strategies for the XH-1 processor, a proposed 128-core
RISC-V design. The discussion covers off-chip DRAM bandwidth,
on-chip interconnect bandwidth, cache hierarchy bandwidth, and the
interactions between bandwidth, coherence traffic, and the memory
system.
This document is part of the XH-1 research series (area 05-memory).
Statements about XH-1 implementation are proposals unless explicitly
marked as decided.
---
## 1. Problem Statement
A 128-core processor generates substantially more memory traffic than
single-core or small-count-core designs. Even if each core issues a
modest number of outstanding memory requests, the aggregate demand can
quickly exceed the bandwidth available from a realistic DRAM subsystem.
The XH-1 design must therefore answer several related questions:
- How much DRAM bandwidth is physically available on the target package?
- How much of that bandwidth is consumed by coherence traffic?
- How much remains for useful application work?
- How is that remaining bandwidth shared between cores without
starvation or severe unfairness?
- Where do bandwidth bottlenecks arise in the on-chip interconnect,
cache hierarchy, and memory controllers?
The remainder of this document analyses each of these questions.
---
## 2. Background
### 2.1 DRAM Bandwidth Fundamentals
Modern DDR systems provide bandwidth through wide channels and high
transfer rates:
- DDR5 devices use 32-bit sub-channels per module rank, with
transfer rates up to 8400 MT/s in JEDEC-published modules.
- A single DDR5 channel (64 bits) at 6400 MT/s delivers approximately
51.2 GB/s peak; a dual-channel configuration roughly doubles this.
- HBM3 stacks deliver substantially higher bandwidth per package
(around 819 GB/s for a 6-high stack per HBM3 specification).
- HBM3E extends this further, with published per-stack figures in the
1 TB/s range for selected devices.
The XH-1 design has not fixed a memory technology. Both DDR5 and
HBM3-class options are considered below as proposals.
### 2.2 RISC-V Memory Model
RISC-V defines a weak memory ordering model, the RISC-V Weak Memory
Ordering (RVWMO), in the architecture specification. RVWMO permits a
wide range of reorderings and is similar in spirit to other modern
weakly ordered architectures.
Key consequences for bandwidth analysis:
- Loads may bypass earlier stores under default rules when addresses
differ.
- Coherence is required only at a single total order per address
(the "coherence order" axiom).
- Fence instructions (`fence`, `fence.i`, `fence.vma`) are required
to constrain ordering where software demands it.
A weakly ordered model does not directly change bandwidth demand but
influences how much buffering, speculation, and replay logic is
required to expose bandwidth to software.
### 2.3 Bandwidth vs. Latency
Bandwidth and latency are distinct metrics. A system can have ample
aggregate DRAM bandwidth but very high per-request latency, in which
case small working sets run well but large working sets suffer.
A 128-core system is especially sensitive to the latency component
because aggregate concurrency exposes the latency of every shared
resource, including DRAM banks, the on-chip interconnect, and the
directory or snoop filter.
---
## 3. Bandwidth Demand Estimation
### 3.1 Per-Core Demand
A modern out-of-order core can issue multiple outstanding cache misses.
The exact number depends on the pipeline width, miss status holding
register (MSHR) count, and prefetcher effectiveness, but published
industrial designs commonly support 10-32 outstanding L1 misses per
core.
If each outstanding request ultimately consumes 64 bytes from DRAM,
and a core sustains 10 outstanding misses, the per-core demand is:
10 misses * 64 B / core / (memory round-trip latency)
For a 200 ns round-trip DRAM latency, this is roughly:
10 * 64 B / 200 ns = 3.2 GB/s per core
If all 128 cores simultaneously sustain this rate, the aggregate is
approximately 410 GB/s, which already exceeds what a single DDR5
channel can deliver.
This calculation is intentionally rough. It is intended to motivate
the rest of the document, not to set a precise requirement.
### 3.2 Realistic Demand
In practice, applications rarely sustain peak miss rates on every core
simultaneously. Published studies of throughput-oriented server
workloads report average L2 MPKI (misses per thousand instructions)
in the range of 5-30 for memory-bound kernels and 1-5 for cache-friendly
workloads. These values vary significantly across workloads.
Multiplying through, the worst-case aggregate bandwidth demand of a
128-core system can plausibly reach hundreds of GB/s. The XH-1 design
must therefore plan for a memory subsystem substantially more capable
than a single DDR5 channel.
### 3.3 Coherence Traffic Overhead
Cache coherence protocols add traffic on top of application reads and
writes. The magnitude depends on the protocol (see Section 6) and the
sharing pattern of the workload. A directory-based protocol typically
generates at least one directory lookup per coherence transaction and
potentially one or more forward/intervention messages.
A conservative estimate is that coherence and protocol overhead add
10-30% to the raw application demand, but this can rise significantly
for workloads with intensive cross-core sharing.
---
## 4. On-Chip Bandwidth
DRAM bandwidth is only one part of the system. The on-chip
interconnect and cache hierarchy must also carry:
- Coherence messages between cores and directories.
- Cache line fills from memory controllers to L2/L3 slices.
- Write-backs from L1 to L2.
- Snoop responses and invalidation acknowledgements.
- Interrupt and configuration traffic (typically negligible).
For a 128-core mesh-class interconnect, published industrial designs
indicate bisection bandwidth targets in the multi-TB/s range. Whether
XH-1 reaches this range depends on link width, clock frequency, and
topology choices that are not yet decided.
---
## 5. Interconnect Topology and Bandwidth
The interconnect determines how effectively aggregate on-chip
bandwidth is delivered to the DRAM controllers. Three common
topologies for many-core designs are:
1. Crossbar
2. Ring
3. Mesh
4. Concentrated mesh (cmesh)
| Topology | Typical use | Bandwidth scaling | Notes |
|----------|-------------|-------------------|-------|
| Crossbar | small core counts | linear in core count | poor scalability beyond ~16 cores |
| Ring | mid counts | limited by ring hops | simple but latency grows |
| Mesh | mid-to-high counts | roughly linear with links | common in tiled designs |
| Cmesh | high counts | similar to mesh, fewer routers | reduces router count |
A crossbar at 128 cores is generally impractical due to area and
wire count. A ring at 128 cores suffers from long hop counts and
limited bisection bandwidth. A mesh or concentrated mesh is the
most commonly chosen topology in published academic and industrial
128-core class designs.
The XH-1 topology is not yet decided; both a 2D mesh and a 2D
concentrated mesh are candidate proposals.
---
## 6. Coherence Protocol Bandwidth
### 6.1 Protocol Options
Two main families of coherence protocol exist:
- Snooping protocols (broadcast-based).
- Directory protocols (tracking-based).
At 128 cores, broadcast-based snooping is generally considered
impractical because every coherence transaction must reach every
core, consuming O(N) bandwidth per request. A directory protocol
incurs O(1) directory lookups per request in the common case but
requires directory storage and handles sharing patterns differently.
A directory is therefore the expected XH-1 proposal, consistent with
published 64-128 core research prototypes.
### 6.2 Directory Organization
Directories can be implemented in several ways:
- Full map directory: tracks sharers exactly. Storage cost is O(N)
per line; prohibitive at 128 cores without compression or
hierarchical schemes.
- Limited pointer directory: tracks a fixed number of sharers per
line (for example 4-6 sharers). When the count is exceeded, a
fallback representation (often a "broadcast" or "coarse vector"
marker) is used.
- Coarse vector directory: groups cores into regions (e.g. 8-core
tiles) and tracks presence at region granularity. Lower storage
cost but less precise; can produce additional coherence traffic.
- Sparse directory: similar to coarse vector but with different
trade-offs.
A limited pointer directory is the most commonly used scheme in
published industrial designs in this core-count range. The XH-1
design is expected to use a limited pointer directory, possibly
augmented with a coarse region fallback, but this is a proposal
rather than a decision.
### 6.3 Coherence Bandwidth Cost
For a limited pointer directory with P pointers, the common case
is one directory lookup per coherence transaction. Bandwidth cost
scales with the number of cores only when the directory overflows
into broadcast mode.
Published research indicates that overflow frequency is workload
dependent. For read-mostly shared data, pointer counts typically
stay small. For write-shared data, a limited pointer directory
can overflow frequently, increasing coherence traffic.
---
## 7. Cache Hierarchy Bandwidth
The cache hierarchy must sustain:
- L1 to L2 traffic (fills, write-backs, invalidations).
- L2 to L3 traffic.
- L3 to DRAM traffic.
- Snoop and probe traffic at every level.
Each cache level adds banked storage and multiple read/write ports to
deliver bandwidth. The area and energy cost of these ports grows
nonlinearly with port count.
For XH-1, a three-level hierarchy is the expected proposal:
- Private L1 per core.
- Private L2 per core (or per core pair) with a coherence controller.
- Shared L3 distributed across tiles.
The exact partition of L2 vs. L3 capacity, associativity, and
bandwidth is a design question left open in this document.
---
## 8. Bandwidth Allocation and QoS
When 128 cores compete for limited DRAM bandwidth, some form of
allocation policy is required. Options include:
- Strict priority ordering.
- Weighted fair queuing.
- Per-core bandwidth caps.
- Per-class (e.g. memory-mapped I/O vs. CPU) priorities.
Without allocation policy, a single core running a streaming workload
can saturate the memory subsystem and starve other cores. This is a
documented failure mode in many-core research literature.
RISC-V itself does not define memory bandwidth allocation semantics;
this is an XH-1 implementation choice. It is expected that the XH-1
memory controllers expose some form of bandwidth partitioning, but
this is a proposal rather than a decision.
---
## 9. Memory Controller Architecture
The memory controller is the point at which on-chip bandwidth meets
off-chip DRAM. Key design choices include:
- Number of independent channels.
- Channel-to-core mapping.
- Scheduling policy (FR-FCFS, parallelism-aware batch scheduling,
etc.).
- Address mapping policy.
- Refresh and power management handling.
For DDR5-class memory, a 4-8 channel configuration is a reasonable
starting point for a 128-core design, but this is workload dependent
and is presented as a proposal.
For HBM3-class memory, the channel count is fixed by the stack
configuration; bandwidth is delivered through many independent
pseudo-channels per stack.
---
## 10. Quantitative Summary
The following table summarises representative bandwidth figures from
published sources. These are not XH-1 targets; they are external
reference points.
| Technology / Configuration | Peak Bandwidth | Source basis |
|-----------------------------|----------------|--------------|
| DDR5 single channel, 6400 MT/s | ~51.2 GB/s | JEDEC DDR5 spec |
| DDR5 dual channel, 6400 MT/s | ~102.4 GB/s | JEDEC DDR5 spec |
| HBM3 6-high stack | ~819 GB/s | JEDEC HBM3 spec |
| HBM3E 8-high stack | ~1 TB/s class | vendor published figures |
| 128-core aggregate theoretical (10 outstanding misses/core) | ~410 GB/s | rough calculation, Section 3.1 |
---
## 11. Advantages and Disadvantages of Candidate Memory Systems
| Option | Advantages | Disadvantages |
|--------|------------|---------------|
| DDR5 multi-channel | Standard, widely available, lower cost per GB | Lower peak bandwidth per channel, more channels increase board complexity |
| HBM3 / HBM3E | Very high bandwidth, high channel count per stack | Higher cost, 2.5D/3D packaging required, lower capacity per stack |
| Hybrid (HBM + DDR) | Combine high bandwidth with high capacity | Complex memory hierarchy, two coherence domains |
The XH-1 memory technology choice is not yet decided.
---
## 12. Scalability Problems at 128 Cores
The following scalability problems are well documented in the
many-core literature and are directly relevant to XH-1:
- **DRAM bandwidth saturation**: aggregate demand can exceed
realistic DRAM bandwidth.
- **Directory overflow**: limited pointer directories overflow
more often as core count rises, increasing coherence traffic.
- **Coherence traffic growth**: as more cores share data,
coherence messages grow superlinearly in some sharing patterns.
- **Interconnect bisection bottlenecks**: insufficient bisection
bandwidth creates hotspots at memory controllers.
- **Quality-of-service degradation**: without allocation policy,
fairness collapses as core count rises.
- **Verification cost**: coherence state space grows roughly
combinatorially with core count.
These are not unique to XH-1 but apply to any 128-core coherent
design.
---
## 13. Interactions With Other Subsystems
### 13.1 Pipeline
A deeper pipeline can support higher clock frequencies but increases
the penalty of cache misses. Bandwidth-bound workloads become
sensitive to the round-trip latency the pipeline exposes. The
pipeline must include enough MSHRs to expose memory-level parallelism.
### 13.2 Cache Hierarchy
Cache capacity and associativity directly affect miss rate and hence
DRAM bandwidth demand. Larger caches reduce demand at the cost of
area and access time. Bandwidth and capacity are coupled; both
must be planned together.
### 13.3 Memory System
The memory controller scheduling policy, address mapping, and
refresh handling all affect effective bandwidth. Closed-page vs.
open-page policies, bank parallelism, and request batching
materially change delivered bandwidth.
### 13.4 Interconnect
The interconnect must deliver at least as much aggregate bandwidth
as the DRAM subsystem, plus coherence overhead. Otherwise the
interconnect becomes the bottleneck and DRAM is underutilised.
### 13.5 Coherence
Coherence traffic is a first-class component of bandwidth demand.
Protocol choice and directory organisation directly determine the
share of bandwidth consumed by coherence.
### 13.6 Interrupts
Interrupts are typically not bandwidth-limited, but interrupt
delivery traffic and any interrupt-cause queues add minor traffic
to the system fabric.
### 13.7 Operating System
The OS schedules processes across cores, controls page placement,
and may implement memory allocation policies that affect locality.
The OS does not directly set DRAM bandwidth but its page colour
and NUMA-style placement policies interact with the memory
controllers.
### 13.8 Verification
Coherence protocols at 128 cores have a large state space.
Verification complexity grows roughly with the product of
per-core state and the number of cores. Formal methods,
randomised testing, and traffic pattern fuzzing are all relevant.
### 13.9 Performance
Bandwidth is one of the dominant performance constraints for
memory-bound workloads. Performance counters for DRAM channel
utilisation, interconnect link utilisation, and coherence message
counts are essential for tuning.
---
## 14. RISC-V Architectural Requirements
The following RISC-V architectural elements are relevant to bandwidth
analysis:
- The RISC-V ISA specifies `fence`, `fence.i`, and `fence.vma`
instructions that order memory operations.
- The privileged architecture defines physical memory attributes
(PMA) and, in designs with translation, page-based attributes
(PBMT, NAPOT, etc.) that can affect memory behaviour.
- The RISC-V coherency management extensions (for example, the
proposed H-extension in the vector and hypervisor contexts, and
the cache-management operations in the RISC-V Profiles and
IOMMU-related specifications) provide instructions such as
`cbo.clean`, `cbo.flush`, and `cbo.inval` that interact with
cache bandwidth.
RISC-V itself does not mandate a specific coherence protocol or
memory bandwidth management policy. These are XH-1 implementation
choices.
---
## 15. Open Questions and Unresolved Design Decisions
The following items are not yet decided and require further research
or design effort:
- Choice of DRAM technology (DDR5 multi-channel vs. HBM3-class).
- Number and placement of memory controllers.
- Interconnect topology (mesh vs. concentrated mesh, dimensional
order, link width).
- Coherence protocol variant (e.g. MESI vs. MOESI and the
directory structure).
- Bandwidth allocation policy and per-core fairness guarantees.
- L2 vs. L3 capacity split and per-tile bandwidth targets.
- Quality-of-service hooks exposed to the operating system.
- Verification strategy for the coherence and interconnect at
128 cores.
These are recorded as open questions; no recommendation is made
where the available evidence does not strongly favour a particular
choice.
---
## 16. Recommendations
Recommendations are made only where published evidence and engineering
practice converge.
- **Use a directory-based coherence protocol**: published 64-128 core
research and industrial designs consistently use directories; a
snoop-based approach is not credible at 128 cores.
- **Plan for multi-channel DRAM**: a single DDR5 channel is unlikely
to meet aggregate demand; at least 2-4 channels are a reasonable
starting point if DDR5 is used, with HBM as an alternative if
higher bandwidth is required.
- **Provide per-class quality-of-service in the memory controllers**:
without allocation, fairness is not achievable in a 128-core
system.
- **Instrument the design with bandwidth and coherence performance
counters**: empirical tuning is unavoidable.
- **Verify coherence state space with a combination of formal
methods and randomised testing**: full state space enumeration
is not feasible at 128 cores.
These are recommendations, not decisions; they are subject to
revision as the XH-1 project progresses.
---
## 17. References
- RISC-V International, *The RISC-V Instruction Set Manual, Volume I:
Unprivileged ISA*, current ratified version. RVWMO is described
in Volume I.
- RISC-V International, *The RISC-V Instruction Set Manual, Volume II:
Privileged Architecture*, current ratified version.
- JEDEC, *DDR5 SDRAM Standard* (JESD79-5), published release.
- JEDEC, *High Bandwidth Memory DRAM (HBM3)*, JESD238A.
- JEDEC, *High Bandwidth Memory DRAM (HBM3E)*, JESD238B-1 (where
applicable).
- Hennessy and Patterson, *Computer Architecture: A Quantitative
Approach*, recent edition, for general memory system modelling
and AMAT analysis.
- Sorin, Hill, and Wood, *A Primer on Memory Consistency and Cache
Coherence*, for coherence protocol analysis.
- Martin, Hill, and Wood, and successor work, on many-core
coherence verification and token coherence (where relevant to
XH-1 verification work).
Where specific figures in this document depend on vendor or
JEDEC-published numbers, the source is indicated in-line. Where
figures are derived from calculations, this is stated explicitly.
---
## 18. Document Status
This document is a research proposal and analysis. It is not a
description of an implemented design. Items marked as "proposal"
or "expected" have not been formally decided within the XH-1
project. Items marked as "open" represent unresolved design
questions. The document should be re-reviewed as the XH-1
architecture specification stabilises.

Some files were not shown because too many files have changed in this diff Show More