mirror of
https://github.com/allexanderbergmns/xh1-research.git
synced 2026-08-26 17:47:02 +00:00
PASSED: 4x | FAILED: 1X
We have continued with the testing/research. We have seen a gigantic improvement with the `PASSING` tests. When we changed the objective agents/models to more careful model, we had a failed test after 3 continue attempts. We will continue with this test, the failed test was: `research/05-memory/atomics.md`. We also published the `issues` which are both capable of displaying failed and passed tests: https://git.revivesparc.com/riscvcxh1/xh1-research-fork/issues. We are going to continue more tests, and make sure that the `atomics` research will pass.
This commit is contained in:
@@ -320,3 +320,432 @@
|
|||||||
[2026-08-26T10:34:46Z] Maximum research rounds reached.
|
[2026-08-26T10:34:46Z] Maximum research rounds reached.
|
||||||
[2026-08-26T10:34:46Z] Leaving original document unchanged.
|
[2026-08-26T10:34:46Z] Leaving original document unchanged.
|
||||||
[2026-08-26T10:34:46Z] Marking topic failed for this run.
|
[2026-08-26T10:34:46Z] Marking topic failed for this run.
|
||||||
|
[2026-08-26T12:11:25Z] Started run: 20260826T121125Z
|
||||||
|
[2026-08-26T12:11:25Z] ==================================================
|
||||||
|
[2026-08-26T12:11:25Z] Researching: research/05-memory/memory-ordering.md
|
||||||
|
[2026-08-26T12:11:25Z] ==================================================
|
||||||
|
[2026-08-26T12:11:25Z] Research round 1/3
|
||||||
|
[2026-08-26T12:11:25Z] Running researcher.
|
||||||
|
[2026-08-26T12:11:25Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=1
|
||||||
|
[2026-08-26T12:11:25Z] API request failed.
|
||||||
|
[2026-08-26T12:11:26Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=2
|
||||||
|
[2026-08-26T12:11:26Z] API request failed.
|
||||||
|
[2026-08-26T12:11:58Z] Started run: 20260826T121158Z
|
||||||
|
[2026-08-26T12:11:58Z] ==================================================
|
||||||
|
[2026-08-26T12:11:58Z] Researching: research/05-memory/memory-ordering.md
|
||||||
|
[2026-08-26T12:11:58Z] ==================================================
|
||||||
|
[2026-08-26T12:11:58Z] Research round 1/3
|
||||||
|
[2026-08-26T12:11:58Z] Running researcher.
|
||||||
|
[2026-08-26T12:11:58Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=1
|
||||||
|
[2026-08-26T12:12:47Z] Response received: 19485 bytes
|
||||||
|
[2026-08-26T12:12:47Z] Running reviewer.
|
||||||
|
[2026-08-26T12:12:47Z] API: provider=ollama model=gemma4:12b attempt=1
|
||||||
|
[2026-08-26T12:15:45Z] Response received: 1622 bytes
|
||||||
|
[2026-08-26T12:15:45Z] Research rejected by reviewer.
|
||||||
|
[2026-08-26T12:15:45Z] Preparing revision round 2.
|
||||||
|
[2026-08-26T12:15:50Z] Research round 2/3
|
||||||
|
[2026-08-26T12:15:50Z] Running revision agent.
|
||||||
|
[2026-08-26T12:15:50Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=1
|
||||||
|
[2026-08-26T12:15:51Z] API request failed.
|
||||||
|
[2026-08-26T12:15:52Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=2
|
||||||
|
[2026-08-26T12:16:25Z] Response received: 23244 bytes
|
||||||
|
[2026-08-26T12:16:25Z] Running reviewer.
|
||||||
|
[2026-08-26T12:16:25Z] API: provider=ollama model=gemma4:12b attempt=1
|
||||||
|
[2026-08-26T12:19:32Z] Response received: 1 bytes
|
||||||
|
[2026-08-26T12:19:32Z] Research rejected by reviewer.
|
||||||
|
[2026-08-26T12:19:32Z] Preparing revision round 3.
|
||||||
|
[2026-08-26T12:19:37Z] Research round 3/3
|
||||||
|
[2026-08-26T12:19:37Z] Running revision agent.
|
||||||
|
[2026-08-26T12:19:37Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=1
|
||||||
|
[2026-08-26T12:19:38Z] API request failed.
|
||||||
|
[2026-08-26T12:19:39Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=2
|
||||||
|
[2026-08-26T12:19:44Z] API request failed.
|
||||||
|
[2026-08-26T12:19:48Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=3
|
||||||
|
[2026-08-26T12:20:25Z] Response received: 25475 bytes
|
||||||
|
[2026-08-26T12:20:25Z] Running reviewer.
|
||||||
|
[2026-08-26T12:20:25Z] API: provider=ollama model=gemma4:12b attempt=1
|
||||||
|
[2026-08-26T12:23:32Z] Response received: 1 bytes
|
||||||
|
[2026-08-26T12:23:32Z] Research rejected by reviewer.
|
||||||
|
[2026-08-26T12:23:32Z] Maximum research rounds reached.
|
||||||
|
[2026-08-26T12:23:32Z] Leaving original document unchanged.
|
||||||
|
[2026-08-26T12:23:32Z] Marking topic failed for this run.
|
||||||
|
[2026-08-26T12:25:59Z] Started run: 20260826T122559Z
|
||||||
|
[2026-08-26T12:25:59Z] ==================================================
|
||||||
|
[2026-08-26T12:25:59Z] Researching: research/05-memory/memory-ordering.md
|
||||||
|
[2026-08-26T12:25:59Z] ==================================================
|
||||||
|
[2026-08-26T12:25:59Z] Research round 1/3
|
||||||
|
[2026-08-26T12:25:59Z] Running researcher.
|
||||||
|
[2026-08-26T12:25:59Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=1
|
||||||
|
[2026-08-26T12:26:00Z] API HTTP status: 429
|
||||||
|
[2026-08-26T12:26:00Z] WARNING: Rate limited by openrouter.
|
||||||
|
[2026-08-26T12:26:00Z] Waiting 5s before retry.
|
||||||
|
[2026-08-26T12:26:05Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=2
|
||||||
|
[2026-08-26T12:26:06Z] API HTTP status: 429
|
||||||
|
[2026-08-26T12:26:06Z] WARNING: Rate limited by openrouter.
|
||||||
|
[2026-08-26T12:26:06Z] Waiting 20s before retry.
|
||||||
|
[2026-08-26T12:26:26Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=3
|
||||||
|
[2026-08-26T12:27:36Z] Response received: 23159 bytes
|
||||||
|
[2026-08-26T12:27:36Z] Running reviewer.
|
||||||
|
[2026-08-26T12:27:37Z] API: provider=ollama model=gemma4:12b attempt=1
|
||||||
|
[2026-08-26T12:30:29Z] Response received: 4699 bytes
|
||||||
|
[2026-08-26T12:30:29Z] WARNING: Reviewer returned no valid VERDICT.
|
||||||
|
[2026-08-26T12:30:29Z] WARNING: This is a reviewer failure, not a research failure.
|
||||||
|
[2026-08-26T12:33:03Z] Started run: 20260826T123303Z
|
||||||
|
[2026-08-26T12:33:03Z] ==================================================
|
||||||
|
[2026-08-26T12:33:03Z] Researching: research/05-memory/memory-ordering.md
|
||||||
|
[2026-08-26T12:33:03Z] ==================================================
|
||||||
|
[2026-08-26T12:33:03Z] Research round 1/3
|
||||||
|
[2026-08-26T12:33:03Z] Running researcher.
|
||||||
|
[2026-08-26T12:33:03Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=1
|
||||||
|
[2026-08-26T12:33:04Z] API HTTP status: 429
|
||||||
|
[2026-08-26T12:33:04Z] WARNING: Rate limited by openrouter.
|
||||||
|
[2026-08-26T12:33:04Z] Waiting 5s before retry.
|
||||||
|
[2026-08-26T12:33:09Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=2
|
||||||
|
[2026-08-26T12:35:05Z] Response received: 41239 bytes
|
||||||
|
[2026-08-26T12:35:05Z] Running reviewer.
|
||||||
|
[2026-08-26T12:35:05Z] API: provider=ollama model=gemma4:12b attempt=1
|
||||||
|
[2026-08-26T12:38:10Z] Response received: 2000 bytes
|
||||||
|
[2026-08-26T12:38:10Z] Reviewer did not provide a machine-readable verdict.
|
||||||
|
[2026-08-26T12:38:10Z] Requesting verdict-format repair.
|
||||||
|
[2026-08-26T12:38:10Z] API: provider=ollama model=gemma4:12b attempt=1
|
||||||
|
[2026-08-26T12:38:52Z] Response received: 87 bytes
|
||||||
|
[2026-08-26T12:38:52Z] WARNING: Reviewer returned no valid VERDICT.
|
||||||
|
[2026-08-26T12:38:52Z] WARNING: This is a reviewer failure, not a research failure.
|
||||||
|
[2026-08-26T12:39:54Z] Started run: 20260826T123954Z
|
||||||
|
[2026-08-26T12:39:54Z] ==================================================
|
||||||
|
[2026-08-26T12:39:54Z] Researching: research/05-memory/memory-ordering.md
|
||||||
|
[2026-08-26T12:39:54Z] ==================================================
|
||||||
|
[2026-08-26T12:39:54Z] Research round 1/3
|
||||||
|
[2026-08-26T12:39:54Z] Running researcher.
|
||||||
|
[2026-08-26T12:39:54Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=1
|
||||||
|
[2026-08-26T12:41:33Z] Response received: 38665 bytes
|
||||||
|
[2026-08-26T12:41:33Z] Running reviewer.
|
||||||
|
[2026-08-26T12:41:33Z] API: provider=ollama model=gemma4:12b attempt=1
|
||||||
|
[2026-08-26T12:44:38Z] Response received: 2266 bytes
|
||||||
|
[2026-08-26T12:44:38Z] Reviewer did not provide a machine-readable verdict.
|
||||||
|
[2026-08-26T12:44:38Z] Requesting verdict-format repair.
|
||||||
|
[2026-08-26T12:44:38Z] API: provider=ollama model=gemma4:12b attempt=1
|
||||||
|
[2026-08-26T12:45:19Z] Response received: 87 bytes
|
||||||
|
[2026-08-26T12:45:19Z] Recovered reviewer verdict: PASS
|
||||||
|
[2026-08-26T12:45:19Z] Research passed review.
|
||||||
|
[2026-08-26T12:45:19Z] Accepted research document: research/05-memory/memory-ordering.md
|
||||||
|
[2026-08-26T12:48:21Z] Started run: 20260826T124821Z
|
||||||
|
[2026-08-26T12:48:21Z] ==================================================
|
||||||
|
[2026-08-26T12:48:21Z] Researching: research/05-memory/cache-coherency.md
|
||||||
|
[2026-08-26T12:48:21Z] ==================================================
|
||||||
|
[2026-08-26T12:48:21Z] Research round 1/3
|
||||||
|
[2026-08-26T12:48:21Z] Running researcher.
|
||||||
|
[2026-08-26T12:48:21Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=1
|
||||||
|
[2026-08-26T12:49:42Z] Response received: 27828 bytes
|
||||||
|
[2026-08-26T12:49:42Z] Running reviewer.
|
||||||
|
[2026-08-26T12:49:42Z] API: provider=ollama model=gemma4:12b attempt=1
|
||||||
|
[2026-08-26T12:52:25Z] Response received: 2096 bytes
|
||||||
|
[2026-08-26T12:52:25Z] Research passed review.
|
||||||
|
[2026-08-26T12:52:25Z] Accepted research document: research/05-memory/cache-coherency.md
|
||||||
|
[2026-08-26T12:54:26Z] Started run: 20260826T125426Z
|
||||||
|
[2026-08-26T12:54:26Z] ==================================================
|
||||||
|
[2026-08-26T12:54:26Z] Researching: research/05-memory/memory-hierarchy.md
|
||||||
|
[2026-08-26T12:54:26Z] ==================================================
|
||||||
|
[2026-08-26T12:54:26Z] Research round 1/3
|
||||||
|
[2026-08-26T12:54:26Z] Running researcher.
|
||||||
|
[2026-08-26T12:54:26Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=1
|
||||||
|
[2026-08-26T12:54:28Z] API HTTP status: 429
|
||||||
|
[2026-08-26T12:54:28Z] WARNING: Rate limited by openrouter.
|
||||||
|
[2026-08-26T12:54:28Z] Waiting 5s before retry.
|
||||||
|
[2026-08-26T12:54:33Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=2
|
||||||
|
[2026-08-26T12:56:00Z] Response received: 30493 bytes
|
||||||
|
[2026-08-26T12:56:00Z] Running reviewer.
|
||||||
|
[2026-08-26T12:56:00Z] API: provider=ollama model=gemma4:12b attempt=1
|
||||||
|
[2026-08-26T12:56:50Z] Response received: 1780 bytes
|
||||||
|
[2026-08-26T12:56:50Z] Research passed review.
|
||||||
|
[2026-08-26T12:56:50Z] Accepted research document: research/05-memory/memory-hierarchy.md
|
||||||
|
[2026-08-26T12:57:40Z] Started run: 20260826T125740Z
|
||||||
|
[2026-08-26T12:57:40Z] ==================================================
|
||||||
|
[2026-08-26T12:57:40Z] Researching: research/05-memory/memory-architecture.md
|
||||||
|
[2026-08-26T12:57:40Z] ==================================================
|
||||||
|
[2026-08-26T12:57:40Z] Research round 1/3
|
||||||
|
[2026-08-26T12:57:40Z] Running researcher.
|
||||||
|
[2026-08-26T12:57:40Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=1
|
||||||
|
[2026-08-26T12:58:59Z] Response received: 28405 bytes
|
||||||
|
[2026-08-26T12:58:59Z] Running reviewer.
|
||||||
|
[2026-08-26T12:58:59Z] API: provider=ollama model=gemma4:12b attempt=1
|
||||||
|
[2026-08-26T13:02:07Z] WARNING: Provider returned an empty/near-empty response (0 bytes).
|
||||||
|
[2026-08-26T13:02:09Z] API: provider=ollama model=gemma4:12b attempt=2
|
||||||
|
[2026-08-26T13:02:54Z] Response received: 2234 bytes
|
||||||
|
[2026-08-26T13:02:54Z] Research passed review.
|
||||||
|
[2026-08-26T13:02:54Z] Accepted research document: research/05-memory/memory-architecture.md
|
||||||
|
[2026-08-26T13:19:26Z] Started run: 20260826T131926Z
|
||||||
|
[2026-08-26T13:19:26Z] ==================================================
|
||||||
|
[2026-08-26T13:19:26Z] Researching: research/05-memory/memory-bandwidth.md
|
||||||
|
[2026-08-26T13:19:26Z] ==================================================
|
||||||
|
[2026-08-26T13:19:26Z] Research round 1/3
|
||||||
|
[2026-08-26T13:19:26Z] Running researcher.
|
||||||
|
[2026-08-26T13:19:27Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=1
|
||||||
|
[2026-08-26T13:20:19Z] Response received: 20125 bytes
|
||||||
|
[2026-08-26T13:20:19Z] Running reviewer.
|
||||||
|
[2026-08-26T13:20:19Z] API: provider=qwen model=qwen3.7-plus attempt=1
|
||||||
|
[2026-08-26T13:20:19Z] API HTTP status: 000
|
||||||
|
[2026-08-26T13:20:19Z] WARNING: Network/API connection failure.
|
||||||
|
[2026-08-26T13:20:22Z] API: provider=qwen model=qwen3.7-plus attempt=2
|
||||||
|
[2026-08-26T13:20:22Z] API HTTP status: 000
|
||||||
|
[2026-08-26T13:20:22Z] WARNING: Network/API connection failure.
|
||||||
|
[2026-08-26T13:20:28Z] API: provider=qwen model=qwen3.7-plus attempt=3
|
||||||
|
[2026-08-26T13:20:28Z] API HTTP status: 000
|
||||||
|
[2026-08-26T13:20:28Z] WARNING: Network/API connection failure.
|
||||||
|
[2026-08-26T13:21:34Z] Started run: 20260826T132134Z
|
||||||
|
[2026-08-26T13:21:34Z] ==================================================
|
||||||
|
[2026-08-26T13:21:34Z] Researching: research/05-memory/memory-bandwidth.md
|
||||||
|
[2026-08-26T13:21:34Z] ==================================================
|
||||||
|
[2026-08-26T13:21:34Z] Research round 1/3
|
||||||
|
[2026-08-26T13:21:34Z] Running researcher.
|
||||||
|
[2026-08-26T13:21:34Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=1
|
||||||
|
[2026-08-26T13:21:34Z] API HTTP status: 401
|
||||||
|
[2026-08-26T13:21:34Z] WARNING: Client/API error from openrouter.
|
||||||
|
[2026-08-26T13:21:37Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=2
|
||||||
|
[2026-08-26T13:21:38Z] API HTTP status: 401
|
||||||
|
[2026-08-26T13:21:38Z] WARNING: Client/API error from openrouter.
|
||||||
|
[2026-08-26T13:21:44Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=3
|
||||||
|
[2026-08-26T13:21:44Z] API HTTP status: 401
|
||||||
|
[2026-08-26T13:21:44Z] WARNING: Client/API error from openrouter.
|
||||||
|
[2026-08-26T13:21:44Z] WARNING: API request exhausted all retries.
|
||||||
|
[2026-08-26T13:21:44Z] WARNING: Research API failed.
|
||||||
|
[2026-08-26T13:21:44Z] WARNING: Topic will NOT be marked as failed.
|
||||||
|
[2026-08-26T13:21:46Z] Started run: 20260826T132146Z
|
||||||
|
[2026-08-26T13:21:46Z] ==================================================
|
||||||
|
[2026-08-26T13:21:46Z] Researching: research/05-memory/memory-bandwidth.md
|
||||||
|
[2026-08-26T13:21:46Z] ==================================================
|
||||||
|
[2026-08-26T13:21:46Z] Research round 1/3
|
||||||
|
[2026-08-26T13:21:46Z] Running researcher.
|
||||||
|
[2026-08-26T13:21:46Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=1
|
||||||
|
[2026-08-26T13:21:47Z] API HTTP status: 401
|
||||||
|
[2026-08-26T13:21:47Z] WARNING: Client/API error from openrouter.
|
||||||
|
[2026-08-26T13:23:02Z] Started run: 20260826T132302Z
|
||||||
|
[2026-08-26T13:23:02Z] ==================================================
|
||||||
|
[2026-08-26T13:23:02Z] Researching: research/05-memory/memory-bandwidth.md
|
||||||
|
[2026-08-26T13:23:02Z] ==================================================
|
||||||
|
[2026-08-26T13:23:02Z] Research round 1/3
|
||||||
|
[2026-08-26T13:23:02Z] Running researcher.
|
||||||
|
[2026-08-26T13:23:02Z] API: provider=qwen model=qwen3.8-max attempt=1
|
||||||
|
[2026-08-26T13:23:02Z] API HTTP status: 000
|
||||||
|
[2026-08-26T13:23:02Z] WARNING: Network/API connection failure.
|
||||||
|
[2026-08-26T13:23:05Z] API: provider=qwen model=qwen3.8-max attempt=2
|
||||||
|
[2026-08-26T13:23:05Z] API HTTP status: 000
|
||||||
|
[2026-08-26T13:23:05Z] WARNING: Network/API connection failure.
|
||||||
|
[2026-08-26T13:23:29Z] Started run: 20260826T132329Z
|
||||||
|
[2026-08-26T13:23:29Z] ==================================================
|
||||||
|
[2026-08-26T13:23:29Z] Researching: research/05-memory/memory-bandwidth.md
|
||||||
|
[2026-08-26T13:23:29Z] ==================================================
|
||||||
|
[2026-08-26T13:23:29Z] Research round 1/3
|
||||||
|
[2026-08-26T13:23:29Z] Running researcher.
|
||||||
|
[2026-08-26T13:23:29Z] API: provider=qwen model=qwen3.8-max attempt=1
|
||||||
|
[2026-08-26T13:23:29Z] API HTTP status: 000
|
||||||
|
[2026-08-26T13:23:29Z] WARNING: Network/API connection failure.
|
||||||
|
[2026-08-26T13:25:56Z] Started run: 20260826T132556Z
|
||||||
|
[2026-08-26T13:25:56Z] ==================================================
|
||||||
|
[2026-08-26T13:25:56Z] Researching: research/05-memory/memory-bandwidth.md
|
||||||
|
[2026-08-26T13:25:56Z] ==================================================
|
||||||
|
[2026-08-26T13:25:56Z] Research round 1/3
|
||||||
|
[2026-08-26T13:25:56Z] Running researcher.
|
||||||
|
[2026-08-26T13:25:56Z] API: provider=qwen model=qwen3.8-max attempt=1
|
||||||
|
[2026-08-26T13:25:56Z] API HTTP status: 000
|
||||||
|
[2026-08-26T13:25:56Z] WARNING: Network/API connection failure.
|
||||||
|
[2026-08-26T13:26:07Z] Started run: 20260826T132607Z
|
||||||
|
[2026-08-26T13:26:07Z] ==================================================
|
||||||
|
[2026-08-26T13:26:07Z] Researching: research/05-memory/memory-bandwidth.md
|
||||||
|
[2026-08-26T13:26:07Z] ==================================================
|
||||||
|
[2026-08-26T13:26:07Z] Research round 1/3
|
||||||
|
[2026-08-26T13:26:07Z] Running researcher.
|
||||||
|
[2026-08-26T13:26:07Z] API: provider=qwen model=qwen3.8-max attempt=1
|
||||||
|
[2026-08-26T13:26:07Z] API HTTP status: 000
|
||||||
|
[2026-08-26T13:26:07Z] WARNING: Network/API connection failure.
|
||||||
|
[2026-08-26T13:26:13Z] Started run: 20260826T132613Z
|
||||||
|
[2026-08-26T13:26:13Z] ==================================================
|
||||||
|
[2026-08-26T13:26:13Z] Researching: research/05-memory/memory-bandwidth.md
|
||||||
|
[2026-08-26T13:26:13Z] ==================================================
|
||||||
|
[2026-08-26T13:26:13Z] Research round 1/3
|
||||||
|
[2026-08-26T13:26:13Z] Running researcher.
|
||||||
|
[2026-08-26T13:26:13Z] API: provider=qwen model=qwen3.8-max attempt=1
|
||||||
|
[2026-08-26T13:26:13Z] API HTTP status: 000
|
||||||
|
[2026-08-26T13:26:13Z] WARNING: Network/API connection failure.
|
||||||
|
[2026-08-26T13:26:50Z] Started run: 20260826T132650Z
|
||||||
|
[2026-08-26T13:26:50Z] ==================================================
|
||||||
|
[2026-08-26T13:26:50Z] Researching: research/05-memory/memory-bandwidth.md
|
||||||
|
[2026-08-26T13:26:50Z] ==================================================
|
||||||
|
[2026-08-26T13:26:50Z] Research round 1/3
|
||||||
|
[2026-08-26T13:26:50Z] Running researcher.
|
||||||
|
[2026-08-26T13:26:50Z] API: provider=qwen model=qwen3.8-max attempt=1
|
||||||
|
[2026-08-26T13:26:50Z] API HTTP status: 000
|
||||||
|
[2026-08-26T13:26:50Z] WARNING: Network/API connection failure.
|
||||||
|
[2026-08-26T13:27:40Z] Started run: 20260826T132740Z
|
||||||
|
[2026-08-26T13:27:40Z] ==================================================
|
||||||
|
[2026-08-26T13:27:40Z] Researching: research/05-memory/memory-bandwidth.md
|
||||||
|
[2026-08-26T13:27:40Z] ==================================================
|
||||||
|
[2026-08-26T13:27:40Z] Research round 1/3
|
||||||
|
[2026-08-26T13:27:40Z] Running researcher.
|
||||||
|
[2026-08-26T13:27:40Z] API: provider=qwen model=qwen3.8-max attempt=1
|
||||||
|
[2026-08-26T13:27:40Z] API HTTP status: 000
|
||||||
|
[2026-08-26T13:27:40Z] WARNING: Network/API connection failure.
|
||||||
|
[2026-08-26T13:27:57Z] Started run: 20260826T132757Z
|
||||||
|
[2026-08-26T13:27:57Z] ==================================================
|
||||||
|
[2026-08-26T13:27:57Z] Researching: research/05-memory/memory-bandwidth.md
|
||||||
|
[2026-08-26T13:27:57Z] ==================================================
|
||||||
|
[2026-08-26T13:27:57Z] Research round 1/3
|
||||||
|
[2026-08-26T13:27:57Z] Running researcher.
|
||||||
|
[2026-08-26T13:27:57Z] API: provider=qwen model=qwen3.8-max attempt=1
|
||||||
|
[2026-08-26T13:27:57Z] API HTTP status: 000
|
||||||
|
[2026-08-26T13:27:57Z] WARNING: Network/API connection failure.
|
||||||
|
[2026-08-26T13:28:46Z] Started run: 20260826T132846Z
|
||||||
|
[2026-08-26T13:28:46Z] ==================================================
|
||||||
|
[2026-08-26T13:28:46Z] Researching: research/05-memory/memory-bandwidth.md
|
||||||
|
[2026-08-26T13:28:46Z] ==================================================
|
||||||
|
[2026-08-26T13:28:46Z] Research round 1/3
|
||||||
|
[2026-08-26T13:28:46Z] Running researcher.
|
||||||
|
[2026-08-26T13:28:46Z] API: provider=qwen model=qwen3.8-max attempt=1
|
||||||
|
[2026-08-26T13:29:00Z] API HTTP status: 401
|
||||||
|
[2026-08-26T13:29:00Z] WARNING: Client/API error from qwen.
|
||||||
|
[2026-08-26T13:29:03Z] API: provider=qwen model=qwen3.8-max attempt=2
|
||||||
|
[2026-08-26T13:29:14Z] Started run: 20260826T132914Z
|
||||||
|
[2026-08-26T13:29:14Z] ==================================================
|
||||||
|
[2026-08-26T13:29:14Z] Researching: research/05-memory/memory-bandwidth.md
|
||||||
|
[2026-08-26T13:29:14Z] ==================================================
|
||||||
|
[2026-08-26T13:29:14Z] Research round 1/3
|
||||||
|
[2026-08-26T13:29:14Z] Running researcher.
|
||||||
|
[2026-08-26T13:29:14Z] API: provider=qwen model=qwen3.8-max attempt=1
|
||||||
|
[2026-08-26T13:29:15Z] API HTTP status: 401
|
||||||
|
[2026-08-26T13:29:15Z] WARNING: Client/API error from qwen.
|
||||||
|
[2026-08-26T13:29:59Z] Started run: 20260826T132959Z
|
||||||
|
[2026-08-26T13:29:59Z] ==================================================
|
||||||
|
[2026-08-26T13:29:59Z] Researching: research/05-memory/memory-bandwidth.md
|
||||||
|
[2026-08-26T13:29:59Z] ==================================================
|
||||||
|
[2026-08-26T13:29:59Z] Research round 1/3
|
||||||
|
[2026-08-26T13:29:59Z] Running researcher.
|
||||||
|
[2026-08-26T13:29:59Z] API: provider=qwen model=qwen3.8-max attempt=1
|
||||||
|
[2026-08-26T13:30:00Z] API HTTP status: 401
|
||||||
|
[2026-08-26T13:30:00Z] WARNING: Client/API error from qwen.
|
||||||
|
[2026-08-26T13:30:40Z] Started run: 20260826T133040Z
|
||||||
|
[2026-08-26T13:30:40Z] ==================================================
|
||||||
|
[2026-08-26T13:30:40Z] Researching: research/05-memory/memory-bandwidth.md
|
||||||
|
[2026-08-26T13:30:40Z] ==================================================
|
||||||
|
[2026-08-26T13:30:40Z] Research round 1/3
|
||||||
|
[2026-08-26T13:30:40Z] Running researcher.
|
||||||
|
[2026-08-26T13:30:40Z] API: provider=qwen model=qwen3.8-max attempt=1
|
||||||
|
[2026-08-26T13:30:42Z] API HTTP status: 401
|
||||||
|
[2026-08-26T13:30:42Z] WARNING: Client/API error from qwen.
|
||||||
|
[2026-08-26T13:30:59Z] Started run: 20260826T133059Z
|
||||||
|
[2026-08-26T13:30:59Z] ==================================================
|
||||||
|
[2026-08-26T13:30:59Z] Researching: research/05-memory/memory-bandwidth.md
|
||||||
|
[2026-08-26T13:30:59Z] ==================================================
|
||||||
|
[2026-08-26T13:30:59Z] Research round 1/3
|
||||||
|
[2026-08-26T13:30:59Z] Running researcher.
|
||||||
|
[2026-08-26T13:30:59Z] API: provider=qwen model=qwen3.8-max attempt=1
|
||||||
|
[2026-08-26T13:31:00Z] API HTTP status: 401
|
||||||
|
[2026-08-26T13:31:00Z] WARNING: Client/API error from qwen.
|
||||||
|
[2026-08-26T13:31:32Z] Started run: 20260826T133132Z
|
||||||
|
[2026-08-26T13:31:32Z] ==================================================
|
||||||
|
[2026-08-26T13:31:32Z] Researching: research/05-memory/memory-bandwidth.md
|
||||||
|
[2026-08-26T13:31:32Z] ==================================================
|
||||||
|
[2026-08-26T13:31:32Z] Research round 1/3
|
||||||
|
[2026-08-26T13:31:32Z] Running researcher.
|
||||||
|
[2026-08-26T13:31:32Z] API: provider=qwen model=qwen3.8-max attempt=1
|
||||||
|
[2026-08-26T13:31:33Z] API HTTP status: 401
|
||||||
|
[2026-08-26T13:31:33Z] WARNING: Client/API error from qwen.
|
||||||
|
[2026-08-26T13:31:55Z] Started run: 20260826T133155Z
|
||||||
|
[2026-08-26T13:31:55Z] ==================================================
|
||||||
|
[2026-08-26T13:31:55Z] Researching: research/05-memory/memory-bandwidth.md
|
||||||
|
[2026-08-26T13:31:55Z] ==================================================
|
||||||
|
[2026-08-26T13:31:55Z] Research round 1/3
|
||||||
|
[2026-08-26T13:31:55Z] Running researcher.
|
||||||
|
[2026-08-26T13:31:55Z] API: provider=qwen model=qwen3.8-max attempt=1
|
||||||
|
[2026-08-26T13:39:15Z] Started run: 20260826T133915Z
|
||||||
|
[2026-08-26T13:39:15Z] ==================================================
|
||||||
|
[2026-08-26T13:39:15Z] Researching: research/05-memory/memory-bandwidth.md
|
||||||
|
[2026-08-26T13:39:15Z] ==================================================
|
||||||
|
[2026-08-26T13:39:15Z] Research round 1/3
|
||||||
|
[2026-08-26T13:39:15Z] Running researcher.
|
||||||
|
[2026-08-26T13:39:15Z] API: provider=qwen model=qwen3.8-max attempt=1
|
||||||
|
[2026-08-26T13:41:06Z] Started run: 20260826T134106Z
|
||||||
|
[2026-08-26T13:41:06Z] ==================================================
|
||||||
|
[2026-08-26T13:41:06Z] Researching: research/05-memory/memory-bandwidth.md
|
||||||
|
[2026-08-26T13:41:06Z] ==================================================
|
||||||
|
[2026-08-26T13:41:06Z] Research round 1/3
|
||||||
|
[2026-08-26T13:41:06Z] Running researcher.
|
||||||
|
[2026-08-26T13:41:06Z] API: provider=qwen model=qwen3.7-plus attempt=1
|
||||||
|
[2026-08-26T13:43:49Z] Started run: 20260826T134349Z
|
||||||
|
[2026-08-26T13:43:49Z] ==================================================
|
||||||
|
[2026-08-26T13:43:49Z] Researching: research/05-memory/memory-bandwidth.md
|
||||||
|
[2026-08-26T13:43:49Z] ==================================================
|
||||||
|
[2026-08-26T13:43:49Z] Research round 1/3
|
||||||
|
[2026-08-26T13:43:49Z] Running researcher.
|
||||||
|
[2026-08-26T13:43:49Z] API: provider=qwen model=qwen3.7-plus attempt=1
|
||||||
|
[2026-08-26T13:45:21Z] Response received: 11713 bytes
|
||||||
|
[2026-08-26T13:45:21Z] Running reviewer.
|
||||||
|
[2026-08-26T13:45:21Z] API: provider=qwen model=deepseek-v4-pro attempt=1
|
||||||
|
[2026-08-26T13:47:18Z] Response received: 1142 bytes
|
||||||
|
[2026-08-26T13:47:18Z] Research passed review.
|
||||||
|
[2026-08-26T13:47:18Z] Accepted research document: research/05-memory/memory-bandwidth.md
|
||||||
|
[2026-08-26T13:47:58Z] Started run: 20260826T134758Z
|
||||||
|
[2026-08-26T13:47:58Z] ==================================================
|
||||||
|
[2026-08-26T13:47:58Z] Researching: research/05-memory/atomics.md
|
||||||
|
[2026-08-26T13:47:58Z] ==================================================
|
||||||
|
[2026-08-26T13:47:58Z] Research round 1/3
|
||||||
|
[2026-08-26T13:47:58Z] Running researcher.
|
||||||
|
[2026-08-26T13:47:58Z] API: provider=qwen model=qwen3.7-plus attempt=1
|
||||||
|
[2026-08-26T13:49:00Z] Started run: 20260826T134900Z
|
||||||
|
[2026-08-26T13:49:00Z] ==================================================
|
||||||
|
[2026-08-26T13:49:00Z] Researching: research/05-memory/atomics.md
|
||||||
|
[2026-08-26T13:49:00Z] ==================================================
|
||||||
|
[2026-08-26T13:49:00Z] Research round 1/3
|
||||||
|
[2026-08-26T13:49:00Z] Running researcher.
|
||||||
|
[2026-08-26T13:49:00Z] API: provider=qwen model=qwen3.7-plus attempt=1
|
||||||
|
[2026-08-26T13:50:43Z] Response received: 12853 bytes
|
||||||
|
[2026-08-26T13:50:43Z] Running reviewer.
|
||||||
|
[2026-08-26T13:50:43Z] API: provider=qwen model=deepseek-v4-pro attempt=1
|
||||||
|
[2026-08-26T13:52:53Z] Response received: 3245 bytes
|
||||||
|
[2026-08-26T13:52:53Z] Research rejected by reviewer.
|
||||||
|
[2026-08-26T13:52:53Z] Preparing revision round 2.
|
||||||
|
[2026-08-26T13:52:58Z] Research round 2/3
|
||||||
|
[2026-08-26T13:52:58Z] Running revision agent.
|
||||||
|
[2026-08-26T13:52:58Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=1
|
||||||
|
[2026-08-26T13:52:59Z] API HTTP status: 401
|
||||||
|
[2026-08-26T13:52:59Z] WARNING: Client/API error from openrouter.
|
||||||
|
[2026-08-26T13:53:02Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=2
|
||||||
|
[2026-08-26T13:53:02Z] API HTTP status: 401
|
||||||
|
[2026-08-26T13:53:02Z] WARNING: Client/API error from openrouter.
|
||||||
|
[2026-08-26T13:53:08Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=3
|
||||||
|
[2026-08-26T13:53:08Z] API HTTP status: 401
|
||||||
|
[2026-08-26T13:53:08Z] WARNING: Client/API error from openrouter.
|
||||||
|
[2026-08-26T13:53:08Z] WARNING: API request exhausted all retries.
|
||||||
|
[2026-08-26T13:53:08Z] WARNING: Revision API failed.
|
||||||
|
[2026-08-26T13:53:08Z] WARNING: Research round was NOT counted as a valid revision.
|
||||||
|
[2026-08-26T13:53:36Z] Started run: 20260826T135336Z
|
||||||
|
[2026-08-26T13:53:36Z] ==================================================
|
||||||
|
[2026-08-26T13:53:36Z] Researching: research/05-memory/atomics.md
|
||||||
|
[2026-08-26T13:53:36Z] ==================================================
|
||||||
|
[2026-08-26T13:53:36Z] Research round 1/3
|
||||||
|
[2026-08-26T13:53:36Z] Running researcher.
|
||||||
|
[2026-08-26T13:53:36Z] API: provider=qwen model=qwen3.7-plus attempt=1
|
||||||
|
[2026-08-26T13:55:19Z] Response received: 11530 bytes
|
||||||
|
[2026-08-26T13:55:19Z] Running reviewer.
|
||||||
|
[2026-08-26T13:55:19Z] API: provider=qwen model=deepseek-v4-pro attempt=1
|
||||||
|
[2026-08-26T13:57:10Z] Response received: 3825 bytes
|
||||||
|
[2026-08-26T13:57:10Z] Research rejected by reviewer.
|
||||||
|
[2026-08-26T13:57:10Z] Preparing revision round 2.
|
||||||
|
[2026-08-26T13:57:15Z] Research round 2/3
|
||||||
|
[2026-08-26T13:57:15Z] Running revision agent.
|
||||||
|
[2026-08-26T13:57:15Z] API: provider=qwen model=qwen3.7-flash-2026-07-15 attempt=1
|
||||||
|
[2026-08-26T13:58:00Z] Response received: 10054 bytes
|
||||||
|
[2026-08-26T13:58:00Z] Running reviewer.
|
||||||
|
[2026-08-26T13:58:00Z] API: provider=qwen model=deepseek-v4-pro attempt=1
|
||||||
|
[2026-08-26T14:00:15Z] Response received: 3571 bytes
|
||||||
|
[2026-08-26T14:00:15Z] Research rejected by reviewer.
|
||||||
|
[2026-08-26T14:00:15Z] Preparing revision round 3.
|
||||||
|
[2026-08-26T14:00:20Z] Research round 3/3
|
||||||
|
[2026-08-26T14:00:20Z] Running revision agent.
|
||||||
|
[2026-08-26T14:00:20Z] API: provider=qwen model=qwen3.7-flash-2026-07-15 attempt=1
|
||||||
|
[2026-08-26T14:01:09Z] Response received: 9606 bytes
|
||||||
|
[2026-08-26T14:01:09Z] Running reviewer.
|
||||||
|
[2026-08-26T14:01:09Z] API: provider=qwen model=deepseek-v4-pro attempt=1
|
||||||
|
[2026-08-26T14:03:31Z] Response received: 4773 bytes
|
||||||
|
[2026-08-26T14:03:31Z] Research rejected by reviewer.
|
||||||
|
[2026-08-26T14:03:31Z] Maximum research rounds reached.
|
||||||
|
[2026-08-26T14:03:31Z] Leaving original document unchanged.
|
||||||
|
|||||||
@@ -0,0 +1 @@
|
|||||||
|
{"error":{"message":"Provider returned error","code":429,"metadata":{"raw":"minimax/minimax-m3:free is temporarily rate-limited upstream. Please retry shortly, or add your own key to accumulate your rate limits: https://openrouter.ai/settings/integrations","provider_name":"GMICloud","is_byok":false,"provider_error_code":"rate_limit_exceeded","limit_source":"upstream_provider_shared_pool","remedy_hint":"Retry shortly, add your own provider key (https://openrouter.ai/settings/integrations), or route to another provider with provider routing: https://openrouter.ai/docs/features/provider-routing","retry_after_seconds":60,"retry_after_seconds_raw":60,"headers":{"Retry-After":"60"}}},"user_id":"user_2tRWhLFlPACV1q6VcCZVeaq8af4"}
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
{"error":{"message":"Provider returned error","code":429,"metadata":{"raw":"minimax/minimax-m3:free is temporarily rate-limited upstream. Please retry shortly, or add your own key to accumulate your rate limits: https://openrouter.ai/settings/integrations","provider_name":"GMICloud","is_byok":false,"provider_error_code":"rate_limit_exceeded","limit_source":"upstream_provider_shared_pool","remedy_hint":"Retry shortly, add your own provider key (https://openrouter.ai/settings/integrations), or route to another provider with provider routing: https://openrouter.ai/docs/features/provider-routing","retry_after_seconds":60,"retry_after_seconds_raw":60,"headers":{"Retry-After":"60"}}},"user_id":"user_2tRWhLFlPACV1q6VcCZVeaq8af4"}
|
||||||
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
@@ -0,0 +1 @@
|
|||||||
|
{"error":{"message":"Provider returned error","code":429,"metadata":{"raw":"minimax/minimax-m3:free is temporarily rate-limited upstream. Please retry shortly, or add your own key to accumulate your rate limits: https://openrouter.ai/settings/integrations","provider_name":"GMICloud","is_byok":false,"provider_error_code":"rate_limit_exceeded","limit_source":"upstream_provider_shared_pool","remedy_hint":"Retry shortly, add your own provider key (https://openrouter.ai/settings/integrations), or route to another provider with provider routing: https://openrouter.ai/docs/features/provider-routing","retry_after_seconds":60,"retry_after_seconds_raw":60,"headers":{"Retry-After":"60"}}},"user_id":"user_2tRWhLFlPACV1q6VcCZVeaq8af4"}
|
||||||
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
@@ -0,0 +1 @@
|
|||||||
|
{"id":"chatcmpl-948","object":"chat.completion","created":1787747932,"model":"gemma4:12b","system_fingerprint":"fp_ollama","choices":[{"index":0,"message":{"role":"assistant","content":"{\n \"verdict\": \"PASS\",\n \"confidence\": \"HIGH\",\n \"issues\": [],\n \"required_fixes\": []\n}","reasoning":"* Input: A technical review of an XH-1 RISC-V CPU research document.\n * Task: Determine if the reviewer intended \"PASS\" or \"FAIL\".\n * Output Format: Strict JSON (no markdown, no explanation, no code fences).\n * JSON Structure:\n * PASS: `{\"verdict\": \"PASS\", \"confidence\": \"HIGH\", \"issues\": [], \"required_fixes\": []}`\n * FAIL: `{\"verdict\": \"FAIL\", \"confidence\": \"HIGH\", \"issues\": [\"...\"], \"required_fixes\": [\"...\"]}`\n\n * \"The document is a well-structured and technically sound research proposal...\" (Positive)\n * \"It correctly identifies the primary challenges...\" (Positive)\n * \"The proposal aligns with the RISC-V ecosystem...\" (Positive)\n * \"demonstrates a sophisticated understanding...\" (Positive)\n * \"The choice of a directory-based MESI protocol... is the industry standard...\" (Positive)\n * \"The recommendation to use a mesh topology is also appropriate...\" (Positive)\n * \"The inclusion of herd7, diy7, and formal tools... indicates a rigorous approach...\" (Positive)\n * \"The document correctly references the relevant parts of the RISC-V ISA manuals...\" (Positive)\n\n * The reviewer uses exclusively positive language.\n * There are no criticisms, \"but\" clauses, or \"needs improvement\" sections.\n * The verdict is clearly a \"PASS\".\n\n * Verdict: PASS\n * Confidence: HIGH\n * Issues: []\n * Required Fixes: []\n\n * No Markdown? Yes.\n * No Explanation? Yes.\n * No Code Fences? Yes.\n * Valid JSON? Yes."},"finish_reason":"stop"}],"usage":{"prompt_tokens":635,"completion_tokens":445,"total_tokens":1080}}
|
||||||
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
@@ -0,0 +1 @@
|
|||||||
|
{"id":"chatcmpl-997","object":"chat.completion","created":1787748319,"model":"gemma4:12b","system_fingerprint":"fp_ollama","choices":[{"index":0,"message":{"role":"assistant","content":"{\n \"verdict\": \"PASS\",\n \"confidence\": \"HIGH\",\n \"issues\": [],\n \"required_fixes\": []\n}","reasoning":"* Input: A technical review of an XH-1 RISC-V CPU research document.\n * Task: Determine if the reviewer intended \"PASS\" or \"FAIL\".\n * Output Format: Strict JSON (no markdown, no explanation, no code fences).\n * JSON Structure:\n * PASS: `{\"verdict\": \"PASS\", \"confidence\": \"HIGH\", \"issues\": [], \"required_fixes\": []}`\n * FAIL: `{\"verdict\": \"FAIL\", \"confidence\": \"HIGH\", \"issues\": [\"...\"], \"required_fixes\": [\"...\"]}`\n\n * \"The document is technically sound and well-structured...\" (Positive)\n * \"The document correctly identifies...\" (Positive)\n * \"This demonstrates a high level of adherence...\" (Positive)\n * \"The choice of... is the industry standard.\" (Positive)\n * \"The inclusion of... is essential...\" (Positive)\n * \"The 'Unresolved' section shows a mature understanding...\" (Positive)\n * \"This is a valid concern...\" (Positive)\n * There are no \"Critical\" or \"Major\" issues mentioned. The \"Unresolved\" section is praised for showing maturity, not as a failure.\n\n * The reviewer is very positive.\n * Verdict: PASS.\n * Confidence: HIGH.\n * Issues: [].\n * Required Fixes: [].\n\n * No markdown.\n * No code fences.\n * No explanation.\n\n ```json\n {\n \"verdict\": \"PASS\",\n \"confidence\": \"HIGH\",\n \"issues\": [],\n \"required_fixes\": []\n }\n ```"},"finish_reason":"stop"}],"usage":{"prompt_tokens":730,"completion_tokens":427,"total_tokens":1157}}
|
||||||
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
@@ -0,0 +1 @@
|
|||||||
|
{"error":{"message":"Provider returned error","code":429,"metadata":{"raw":"minimax/minimax-m3:free is temporarily rate-limited upstream. Please retry shortly, or add your own key to accumulate your rate limits: https://openrouter.ai/settings/integrations","provider_name":"GMICloud","is_byok":false,"provider_error_code":"rate_limit_exceeded","limit_source":"upstream_provider_shared_pool","remedy_hint":"Retry shortly, add your own provider key (https://openrouter.ai/settings/integrations), or route to another provider with provider routing: https://openrouter.ai/docs/features/provider-routing","retry_after_seconds":60,"retry_after_seconds_raw":60,"headers":{"Retry-After":"60"}}},"user_id":"user_2tRWhLFlPACV1q6VcCZVeaq8af4"}
|
||||||
File diff suppressed because one or more lines are too long
@@ -0,0 +1 @@
|
|||||||
|
{"id":"chatcmpl-481","object":"chat.completion","created":1787749010,"model":"gemma4:12b","system_fingerprint":"fp_ollama","choices":[{"index":0,"message":{"role":"assistant","content":"The document provided is a high-quality technical research overview of the XH-1 memory hierarchy. It correctly identifies the primary challenges associated with a 128-core system, specifically regarding directory-based coherence, mesh interconnect latency, and the complexities of RISC-V memory model compliance.\n\n### Technical Review:\n\n1. **RISC-V Compliance:** The document correctly identifies the critical components for RISC-V compliance: RVWMO, the \"A\" extension (atomic instructions), and the Zicbom/Zicboz extensions for software-managed cache coherence.\n2. **Coherence \u0026 Scalability:** The choice of a directory-based protocol over a snoop-based one is the correct architectural decision for a 128-core system. The mention of \"limited-pointer\" overflows and the \"broadcast fallback\" shows a mature understanding of directory scaling.\n3. **Microarchitecture:** The inclusion of \"critical-word first\" for high-latency mesh misses and the consideration of \"NINE\" (Non-Inclusive, Non-Exclusive) policies for the LLC are standard, high-performance industry practices.\n4. **Verification:** The document correctly identifies that a 128-core directory protocol is a high-risk verification target, correctly suggesting formal verification as a necessary component of the design cycle.\n5. **Nuance:** The distinction between a \"research document\" and a \"specification\" is vital. The document correctly frames the current proposals as \"workload-dependent\" rather than fixed requirements, which is the correct posture for a project at this stage.\n\nThe document is internally consistent, well-referenced, and addresses the specific complexities of large-scale multicore systems.\n\n```json\n{\n \"verdict\": \"PASS\",\n \"confidence\": \"HIGH\",\n \"issues\": [],\n \"required_fixes\": []\n}\n```"},"finish_reason":"stop"}],"usage":{"prompt_tokens":2051,"completion_tokens":404,"total_tokens":2455}}
|
||||||
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
@@ -0,0 +1 @@
|
|||||||
|
{"id":"chatcmpl-921","object":"chat.completion","created":1787749374,"model":"gemma4:12b","system_fingerprint":"fp_ollama","choices":[{"index":0,"message":{"role":"assistant","content":"The document describes a high-performance, many-core RISC-V system (XH-1) with 128 cores. The technical specifications provided are consistent with modern high-performance computing (HPC) architecture and RISC-V standards.\n\n### Technical Review:\n\n1. **Scalability \u0026 Interconnect:** The choice of a **2D mesh** for 128 cores is a standard and correct architectural decision. At this scale, a ring or crossbar is physically and logically impractical. The use of **virtual channels** to separate request, response, and data traffic is critical to prevent head-of-line blocking and protocol deadlocks in a directory-based system.\n2. **Cache Coherence:** The transition from snooping to a **distributed directory MESI protocol** is mandatory for 128 cores. The inclusion of **Zicbom and Zicboz** is a critical requirement for RISC-V Linux compatibility, as these allow the OS to manage cache lines and memory attributes correctly.\n3. **Memory Consistency:** The proposal to use a **litmus test suite** for RISC-V memory consistency is the industry standard for verifying that the hardware correctly implements the Weak Memory Model (WMM) required by the RISC-V ISA.\n4. **Hierarchy:** The L2 being inclusive of L1 and the L3 being the coherence point is a standard design that simplifies the coherence logic at the L3 level.\n5. **OS/Software:** Exposing the topology as **NUMA nodes** is the correct way to handle a 128-core system in Linux, ensuring the scheduler can optimize for local memory access.\n6. **Verification:** The multi-pronged approach (Formal for state machines, Simulation for consistency, and Stress/Synthetic for performance) is a robust verification strategy for a complex SoC.\n\n### Minor Observations:\n- The decision to defer Zicbop is prudent, as it is a more complex extension for \"buffer\" management that may not be necessary for initial silicon.\n- The distinction between \"inclusive\" and \"non-inclusive\" L3 is correctly identified as a key design choice for the implementation phase.\n\nThe document is technically sound, follows RISC-V standards, and addresses the specific challenges of scaling to 128 cores.\n\n```json\n{\n \"verdict\": \"PASS\",\n \"confidence\": \"HIGH\",\n \"issues\": [],\n \"required_fixes\": []\n}\n```"},"finish_reason":"stop"}],"usage":{"prompt_tokens":2051,"completion_tokens":532,"total_tokens":2583}}
|
||||||
File diff suppressed because one or more lines are too long
@@ -0,0 +1 @@
|
|||||||
|
{"error":{"message":"No cookie auth credentials found","code":401}}
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
{"error":{"message":"No cookie auth credentials found","code":401}}
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
{"error":{"message":"No cookie auth credentials found","code":401}}
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
{"error":{"message":"No cookie auth credentials found","code":401}}
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
{"error":{"message":"You didn't provide an API key. You need to provide your API key in an Authorization header using Bearer auth (i.e. Authorization: Bearer YOUR_KEY). ","type":"invalid_request_error","param":null,"code":null},"request_id":"5b968769-a7b3-9d60-8195-7de88e17e1a3"}
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
{"error":{"message":"You didn't provide an API key. You need to provide your API key in an Authorization header using Bearer auth (i.e. Authorization: Bearer YOUR_KEY). ","type":"invalid_request_error","param":null,"code":null},"request_id":"66fd0f2b-1ac5-9b0c-9966-75b3d773b23b"}
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
{"error":{"message":"You didn't provide an API key. You need to provide your API key in an Authorization header using Bearer auth (i.e. Authorization: Bearer YOUR_KEY). ","type":"invalid_request_error","param":null,"code":null},"request_id":"3a23149c-14e0-98ba-9207-bb436ec6d749"}
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
{"error":{"message":"You didn't provide an API key. You need to provide your API key in an Authorization header using Bearer auth (i.e. Authorization: Bearer YOUR_KEY). ","type":"invalid_request_error","param":null,"code":null},"request_id":"48d53b24-2073-9696-8bc0-651a6215025c"}
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
{"error":{"message":"You didn't provide an API key. You need to provide your API key in an Authorization header using Bearer auth (i.e. Authorization: Bearer YOUR_KEY). ","type":"invalid_request_error","param":null,"code":null},"request_id":"3071b8b5-52b6-93b0-bb3a-13392a8ccf4e"}
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
{"error":{"message":"You didn't provide an API key. You need to provide your API key in an Authorization header using Bearer auth (i.e. Authorization: Bearer YOUR_KEY). ","type":"invalid_request_error","param":null,"code":null},"request_id":"f167a86f-f816-9ebf-96d1-5f1d6ba405a6"}
|
||||||
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
@@ -0,0 +1 @@
|
|||||||
|
{"error":{"message":"No cookie auth credentials found","code":401}}
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
{"error":{"message":"No cookie auth credentials found","code":401}}
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
{"error":{"message":"No cookie auth credentials found","code":401}}
|
||||||
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
@@ -0,0 +1,6 @@
|
|||||||
|
2026-08-26T12:12:47Z research/05-memory/memory-ordering.md 1 research completed
|
||||||
|
2026-08-26T12:15:45Z research/05-memory/memory-ordering.md 1 review
|
||||||
|
2026-08-26T12:16:25Z research/05-memory/memory-ordering.md 2 revision completed
|
||||||
|
2026-08-26T12:19:32Z research/05-memory/memory-ordering.md 2 review
|
||||||
|
2026-08-26T12:20:25Z research/05-memory/memory-ordering.md 3 revision completed
|
||||||
|
2026-08-26T12:23:32Z research/05-memory/memory-ordering.md 3 review
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
research/05-memory/memory-ordering.md
|
||||||
@@ -0,0 +1,333 @@
|
|||||||
|
# Memory Ordering
|
||||||
|
|
||||||
|
## Status
|
||||||
|
|
||||||
|
Stub document. XH-1 memory ordering model has not yet been defined. The repository contains no source artifacts, RTL, ISA extension proposals, or reference implementations that establish a memory consistency model for XH-1. All quantitative or XH-1-specific claims in this document are flagged INSUFFICIENT EVIDENCE where applicable.
|
||||||
|
|
||||||
|
## Abstract
|
||||||
|
|
||||||
|
Memory ordering defines the rules by which loads and stores issued by one or more cores become visible to other cores in the system. For a 128-core RISC-V processor, the choice of memory model — and the hardware mechanisms that enforce it — is one of the most consequential architectural decisions, affecting correctness, performance, area, power, scalability, and software portability. This document surveys the design space, identifies the candidate models available within the RISC-V ISA framework, examines alternative hardware enforcement strategies, and frames the open questions for XH-1. No final recommendation is made because the repository does not establish pipeline style, coherency protocol, or interconnect topology, all of which constrain the memory ordering decision. A conditional tentative default is stated, contingent on the assumed constraints being correct.
|
||||||
|
|
||||||
|
## Research Question
|
||||||
|
|
||||||
|
What memory consistency model and hardware ordering mechanism should XH-1 adopt, given:
|
||||||
|
|
||||||
|
1. A 128-core replication target
|
||||||
|
2. The RISC-V ISA baseline
|
||||||
|
3. The need to balance software portability against hardware implementation cost
|
||||||
|
4. Unknown microarchitecture style (in-order vs. out-of-order is not established in the repository)
|
||||||
|
5. The relationship to the cache coherency protocol (defined in `cache-coherency.md`, currently SOON)
|
||||||
|
|
||||||
|
## Background
|
||||||
|
|
||||||
|
### Memory Consistency Models — Terminology
|
||||||
|
|
||||||
|
PROPOSAL — Use the following working taxonomy in this document:
|
||||||
|
|
||||||
|
- **Sequential Consistency (SC)**: All cores observe a single total order of memory operations that respects each core's program order.
|
||||||
|
- **Total Store Order (TSO)**: Loads may bypass stores from the same core; stores from one core are seen in order by all cores. Used by x86.
|
||||||
|
- **Relaxed / Weak Memory Models**: Many orderings are allowed; fences (barriers) are required to enforce constraints. Used by ARM, RISC-V, and PowerPC.
|
||||||
|
- **Release Consistency (RC) / RCsc / RCpc / RCO**: A family of models with different fence semantics. Used by academic designs and the basis for much of the C11/C++ memory model formalization.
|
||||||
|
|
||||||
|
ASSUMPTION — This document uses "RC variants" as an umbrella term for RCsc, RCpc, and RCO. The specific semantics of each variant are not redefined here; readers are referred to the literature.
|
||||||
|
|
||||||
|
### The RISC-V "RVWMO" Model
|
||||||
|
|
||||||
|
FACT — The RISC-V ISA specification defines the **RISC-V Weak Memory Ordering (RVWMO)** model as the baseline. Under RVWMO:
|
||||||
|
|
||||||
|
- Loads and stores to the same address from the same hart are ordered.
|
||||||
|
- A store followed by a load to a *different* address may be reordered.
|
||||||
|
- The `FENCE` instruction is used to enforce orderings. The base FENCE provides `pred` (predecessor) and `succ` (successor) bits covering device I/O (I), memory reads (R), and memory writes (W). Device ordering follows a stronger "Preserved over Higher Privilege and PMA" model (PostModel).
|
||||||
|
- The optional `Ztso` extension provides **RVTSO**, which gives x86-like TSO semantics with a relaxed form of the FENCE.
|
||||||
|
|
||||||
|
FACT — RVWMO and the `Ztso` extension are documented in the ratified RISC-V Unprivileged ISA specification. The repository does not state which revision of the specification is the reference for XH-1; this is an open question.
|
||||||
|
|
||||||
|
ASSUMPTION — RVWMO + the Ztso optional extension is the available ISA-level model space that XH-1 must select from. (The repository does not document any custom ISA extensions, so XH-1 is assumed to stay within ratified RISC-V.)
|
||||||
|
|
||||||
|
### Why Memory Ordering is Hard at 128 Cores
|
||||||
|
|
||||||
|
OPEN QUESTION — The repository does not establish:
|
||||||
|
- Whether XH-1 is a single die, multi-die package, or chiplet-based system
|
||||||
|
- The interconnect topology (mesh, ring, crossbar, NoC, etc.)
|
||||||
|
- Whether coherence is directory-based, broadcast-based, or hybrid
|
||||||
|
- The private-vs-shared L2/L3 split
|
||||||
|
- The memory-side directory or snoop filter organization
|
||||||
|
|
||||||
|
These decisions have direct bearing on which ordering mechanisms are feasible. Any hardware ordering scheme (invalidation queues, store buffers, fence acknowledgments, coherence-based ordering) must be analyzed in the context of these unknowns.
|
||||||
|
|
||||||
|
## Existing Approaches
|
||||||
|
|
||||||
|
### Approach A — RISC-V RVWMO (Baseline)
|
||||||
|
|
||||||
|
The standard weak memory model provided by the RISC-V ISA. Requires implementation of `FENCE`, `FENCE.I`, and the I/O ordering semantics.
|
||||||
|
|
||||||
|
**Mechanism**: Architecture permits aggressive reordering. Hardware implementations commonly use:
|
||||||
|
- Store buffers with forwarding
|
||||||
|
- Load-load reordering queues
|
||||||
|
- Coherence-based ordering: stores become visible at coherence invalidation/upgrade points
|
||||||
|
- Fence implementation via global acknowledgment counters or by draining reorder queues
|
||||||
|
|
||||||
|
INSUFFICIENT EVIDENCE — Which of these mechanisms XH-1 will use is not established by the repository.
|
||||||
|
|
||||||
|
### Approach B — RVWMO + Ztso Extension
|
||||||
|
|
||||||
|
Adds RVTSO. Software that requires x86-like ordering can avoid fences in many cases. Still RISC-V compliant.
|
||||||
|
|
||||||
|
**Mechanism**: Stores from one hart are serialized in the interconnect; loads may still pass earlier stores to different addresses.
|
||||||
|
|
||||||
|
### Approach C — Sequential Consistency (SC)
|
||||||
|
|
||||||
|
All operations observed in a single global order respecting program order. Simplest software model; most expensive hardware in general.
|
||||||
|
|
||||||
|
**Mechanism**: Commonly implemented via store-buffer draining on loads, broadcast-based coherence, and conservative replay mechanisms. The qualitative characterization of "very costly" is a general literature claim, not an XH-1 measurement; whether SC is specifically costly at 128 cores is workload- and topology-dependent and is treated here as a hypothesis, not a known result.
|
||||||
|
|
||||||
|
### Approach D — Release Consistency variants (RCsc, RCpc, RCO)
|
||||||
|
|
||||||
|
Academic models where release/acquire ordering is cheap but stronger fences are required for full memory ordering. Basis of C11/C++ model.
|
||||||
|
|
||||||
|
**Mechanism**: Only release/acquire fences require acknowledgment; acquire-only loads are cheap. Alignment with C11 is a strong software benefit. These variants are not part of the ratified RISC-V ISA, so adoption would require a custom ISA extension.
|
||||||
|
|
||||||
|
## Alternative Designs
|
||||||
|
|
||||||
|
### Hardware Ordering Mechanisms
|
||||||
|
|
||||||
|
PROPOSAL — The following mechanisms should be considered for fence/ordering enforcement once coherency and interconnect are known:
|
||||||
|
|
||||||
|
1. **Invalidation-acknowledgment fences** — A fence blocks until all outstanding coherence requests have been acknowledged. Simple and common; scaling impact depends on directory hop count, aggregate invalidation traffic, and invalidation filter structure.
|
||||||
|
|
||||||
|
2. **Per-core store-buffer drain fences** — A fence stalls the issuing core until its store buffer is empty and all stores have been globally observed. Predictable latency; serializes through the coherence fabric.
|
||||||
|
|
||||||
|
3. **Tournament / Token-based fences** — Fences acquire a global token or wait for a "fence epoch" counter. Avoids worst-case drain, but adds global state.
|
||||||
|
|
||||||
|
4. **FIFO coherence ordering** — Leverages the in-order completion of coherence transactions to provide ordering without explicit fences. Requires in-order interconnect, which conflicts with typical latency-optimized NoC designs.
|
||||||
|
|
||||||
|
5. **Time-to-live / epoch schemes** — Each coherence request carries an epoch tag; ordering is enforced at the L2/L3 directory by queueing requests by epoch.
|
||||||
|
|
||||||
|
6. **Speculative load reordering with rollback** — Loads can execute speculatively past stores; on conflict, the load is replayed. Common in high-performance out-of-order cores (e.g., reported in the IBM POWER literature).
|
||||||
|
|
||||||
|
ASSUMPTION — XH-1's microarchitecture style is not yet established. If out-of-order, the "speculative load reordering with rollback" category becomes a candidate. If in-order, mechanism choices narrow considerably.
|
||||||
|
|
||||||
|
## Comparison
|
||||||
|
|
||||||
|
| Property | SC | RVWMO | RVWMO + Ztso | RC variants |
|
||||||
|
|----------|----|----|----|----|
|
||||||
|
| Software ease | Highest | Medium | Medium-High | Low-Medium |
|
||||||
|
| Hardware cost (general) | Highest | Low | Low-Medium | Low |
|
||||||
|
| Performance headroom (general) | Lowest | Highest | High | High |
|
||||||
|
| Fence latency (general) | N/A | Variable | Variable | Variable |
|
||||||
|
| Verification cost (general) | Lowest in principle | High | High | High |
|
||||||
|
| C11/C++ alignment | Approximate | Approximate | Weak | Direct |
|
||||||
|
| 128-core scalability | Poor (hypothesis) | Unmeasured | Unmeasured | Unmeasured |
|
||||||
|
|
||||||
|
INSUFFICIENT EVIDENCE — Quantitative fence-latency, coherence-traffic, and scalability numbers for XH-1 are unavailable. Several rows above contain qualitative judgments rather than measured results, and the "128-core scalability" and "verification cost" rows reflect general literature consensus, not XH-1 measurements.
|
||||||
|
|
||||||
|
## Advantages
|
||||||
|
|
||||||
|
### RVWMO (Baseline)
|
||||||
|
- **Lowest hardware cost** of all candidates for the base case, as a general property of the model — most orderings are simply not enforced.
|
||||||
|
- **Maximum performance** potential, as a general property — reordering is unconstrained.
|
||||||
|
- **Familiar** to the RISC-V software ecosystem (Linux, RISC-V GCC, LLVM), subject to the open question of whether Linux is the target OS.
|
||||||
|
- **Standardized** — concrete compliance test suites exist externally.
|
||||||
|
|
||||||
|
### RVWMO + Ztso
|
||||||
|
- **x86-like TSO** for software that benefits from stronger ordering.
|
||||||
|
- **Backwards compatible** — software not using Ztso still runs correctly under RVWMO.
|
||||||
|
- **Modest hardware cost** — primarily requires store-store ordering at the coherence layer (qualitative claim, not measured for XH-1).
|
||||||
|
|
||||||
|
### Release Consistency Variants
|
||||||
|
- **Excellent fit for C11/C++** atomics — acquire/release are the most-used fence types in real code, per the C/C++ standards.
|
||||||
|
- **Cheapest common case** — acquire and release can be implemented with lightweight mechanisms, as a general property.
|
||||||
|
- **Hardware cost is concentrated on the rare case** of full seq_cst fences, as a general property.
|
||||||
|
|
||||||
|
## Disadvantages
|
||||||
|
|
||||||
|
### RVWMO (Baseline)
|
||||||
|
- **Fence latency depends on coherence round-trip** — at 128 cores, fence latency is expected to be high because global acknowledgment must traverse the coherence fabric. Specific cycle counts: INSUFFICIENT EVIDENCE.
|
||||||
|
- **Verification complexity is high** — many legal reorderings; the litmus-test failure surface is large. This is a general property of weak models, not an XH-1 measurement.
|
||||||
|
- **Software burden** — kernel and runtime code must insert fences correctly.
|
||||||
|
|
||||||
|
### RVWMO + Ztso
|
||||||
|
- **Two consistency models in one chip** — adds documentation, validation, and software education cost.
|
||||||
|
- **In-flight mixing** of ordering regimes in the same software is hard to reason about.
|
||||||
|
|
||||||
|
### Release Consistency Variants
|
||||||
|
- **Non-standard for RISC-V** — would require a custom ISA extension. Cannot be recommended without a strong software-side driver.
|
||||||
|
- **Verification cost is high** in general because of subtle fence semantics; whether it is the "highest" among candidates is a claim about the literature, not a measured XH-1 result.
|
||||||
|
|
||||||
|
### Sequential Consistency
|
||||||
|
- **Performance cost is severe in the general case** — every load may stall on store-buffer drain; whether this is "severe at 128 cores" specifically is workload- and topology-dependent and is not established by the repository.
|
||||||
|
- **No selective escape valve** — software cannot opt out of the strongest model.
|
||||||
|
- **Does not match RISC-V ecosystem expectations** — surprising to RISC-V software developers, as a general observation.
|
||||||
|
|
||||||
|
## XH-1 Considerations
|
||||||
|
|
||||||
|
OPEN QUESTION — The following XH-1-internal questions are unresolved by the repository:
|
||||||
|
|
||||||
|
1. Is XH-1's pipeline in-order or out-of-order?
|
||||||
|
2. What is the coherence protocol (MOESI, MESI, directory-based, broadcast)?
|
||||||
|
3. Is the interconnect a NoC, ring, or crossbar? Does it preserve in-order delivery of coherence responses?
|
||||||
|
4. What is the expected working set of a typical hart — does streaming-store optimization matter?
|
||||||
|
5. Does XH-1 target HPC, server, embedded, or mixed workloads? (Workload affects fence frequency.)
|
||||||
|
6. Is XH-1 a research vehicle (where SC is acceptable for simplicity) or a product (where RVWMO compliance is mandatory)?
|
||||||
|
7. Does the memory map use non-coherent regions (e.g., DMA, I/O) that interact with the ordering model?
|
||||||
|
|
||||||
|
ASSUMPTION — If XH-1 targets RISC-V ecosystem compatibility (Linux, RISC-V GCC, RISC-V LLVM), then RVWMO is mandatory. Ztso may optionally be added for x86-software porting convenience.
|
||||||
|
|
||||||
|
## 128-Core Scaling Considerations
|
||||||
|
|
||||||
|
INSUFFICIENT EVIDENCE — Quantitative figures for XH-1 at 128 cores are unavailable. The qualitative observations below are hypotheses grounded in general literature, not XH-1 measurements.
|
||||||
|
|
||||||
|
### Fence Latency
|
||||||
|
- **Fence global acknowledgment** latency is bounded below by the worst-case coherence round-trip time. In a 128-core system, this can plausibly reach many tens of cycles in a NoC and higher in a ring or multi-hop topology; specific XH-1 figures: INSUFFICIENT EVIDENCE.
|
||||||
|
- **Store-buffer drain** latency is approximately proportional to the store buffer depth and the time to invalidate all sharers. Specific XH-1 figures: INSUFFICIENT EVIDENCE.
|
||||||
|
- **Token / epoch fences** may scale better because they avoid draining the entire coherence fabric — they synchronize at a logical epoch boundary. This is a general property, not an XH-1 measurement.
|
||||||
|
|
||||||
|
### Coherence Traffic
|
||||||
|
- **RVWMO fences** generate invalidation or acknowledgment traffic that competes with normal coherence traffic. The magnitude depends on directory organization, snoop-filter effectiveness, and share-set size.
|
||||||
|
- **SC at 128 cores** can plausibly saturate a directory-based interconnect because every load's coherence transaction may stall pending store ordering; this is a hypothesis, not a measured XH-1 result.
|
||||||
|
- **RC variants** concentrate traffic on the fence path only, as a general property.
|
||||||
|
|
||||||
|
### Verification Scalability
|
||||||
|
- **State-space explosion**: at 128 cores, the number of concurrently observable memory operations grows combinatorially in the unmitigated model. Memory-ordering verification is a known scaling bottleneck in the literature, and the repository does not establish what reduction or abstraction techniques XH-1 will employ.
|
||||||
|
- **RC variants** are commonly cited as among the hardest to verify because of subtle fence interactions; whether they are hardest in the absolute sense is a literature claim, not an XH-1 measurement.
|
||||||
|
- **SC is easiest to verify in principle** because it admits fewer legal behaviors, but worst in performance.
|
||||||
|
|
||||||
|
INSUFFICIENT EVIDENCE — No quantitative fence-latency, coherence-traffic, or verification-scaling data is available for XH-1. Estimates cannot be made without knowing the interconnect, coherence protocol, and verification methodology.
|
||||||
|
|
||||||
|
## Performance Considerations
|
||||||
|
|
||||||
|
PROPOSAL — Performance-impact dimensions to evaluate once workload and microarchitecture are defined:
|
||||||
|
|
||||||
|
1. **Steady-state IPC** for memory-bound workloads — qualitatively, RVWMO permits the most reordering and SC the least; XH-1-specific IPC: INSUFFICIENT EVIDENCE.
|
||||||
|
2. **Fence latency** — measured as cycles from fence issue to fence completion. Dominant for synchronization-heavy code. XH-1-specific: INSUFFICIENT EVIDENCE.
|
||||||
|
3. **Scalable synchronization** — how the model interacts with lock implementations, RCU, and atomic primitives.
|
||||||
|
4. **Load-use latency** — whether loads can bypass earlier stores (allowed in RVWMO and TSO, not in SC).
|
||||||
|
5. **Coherence miss rate** interaction — does the ordering model inflate or deflate cache-line ping-pong? Model-dependent; XH-1-specific: INSUFFICIENT EVIDENCE.
|
||||||
|
|
||||||
|
ASSUMPTION — For HPC workloads with high-fence frequency (e.g., MPI, OpenMP, lock-heavy code), even small differences in fence latency will dominate. For server workloads with mixed locking, the cost of fences is amortized over longer critical sections. No XH-1 workload profile is documented in the repository.
|
||||||
|
|
||||||
|
## Area Considerations
|
||||||
|
|
||||||
|
INSUFFICIENT EVIDENCE — All area descriptions in this section are qualitative. Precise XH-1 area figures (in gates, µm², or mm²) are unavailable; no synthesis, layout, or RTL data exists in the repository.
|
||||||
|
|
||||||
|
PROPOSAL — Qualitative area cost of ordering mechanisms, as a function of microarchitectural choices:
|
||||||
|
|
||||||
|
| Mechanism | Approximate area cost per core (qualitative) |
|
||||||
|
|-----------|----------------------------------------------|
|
||||||
|
| Store buffer with forwarding | Small (few entries × cache-line width) |
|
||||||
|
| Load-load reordering queue | Small to medium (grows with IPC and reorder depth) |
|
||||||
|
| Fence-acknowledgment counter | Negligible (per-core counter) |
|
||||||
|
| Global epoch / token state | Medium (one per chip, not per core) |
|
||||||
|
| Speculative load reorder + rollback | Large (load-load queue, replay logic) |
|
||||||
|
|
||||||
|
PROPOSAL — At 128 cores, per-core area costs are replicated 128 times. Global state (epoch counters, fence acknowledgers) is O(1) per chip and therefore amortizes favorably on a per-core basis, but the critical-path and clock-distribution impact of global state is not established by the repository.
|
||||||
|
|
||||||
|
## Power and Energy Considerations
|
||||||
|
|
||||||
|
INSUFFICIENT EVIDENCE — No per-fence, per-core, or aggregate energy figures for XH-1 are available. The earlier version of this document included a "100 pJ-per-core fence" figure that was not substantiated by any repository artifact or external citation; that figure has been removed.
|
||||||
|
|
||||||
|
PROPOSAL — Qualitative energy implications, as general properties of the model classes:
|
||||||
|
|
||||||
|
- **Fences** generate global coherence traffic — energy cost scales with the number of cores that must acknowledge. The aggregate chip-level energy per fence is therefore expected to grow with core count. Quantitative XH-1 values: INSUFFICIENT EVIDENCE.
|
||||||
|
- **Speculative load reordering** wastes energy on rollback when a violation is detected, as a general property.
|
||||||
|
- **Strong ordering models** keep more coherence state in flight (more invalidations, more retries), as a general property.
|
||||||
|
- **RVWMO** with lightweight fences (when software permits) is the lowest-energy regime in the general case.
|
||||||
|
|
||||||
|
## Implementation Considerations
|
||||||
|
|
||||||
|
PROPOSAL — Implementation order of dependencies (the repository does not establish these, but the research area suggests the following):
|
||||||
|
|
||||||
|
1. Define coherency protocol → `cache-coherency.md`
|
||||||
|
2. Define interconnect topology → not yet in repository
|
||||||
|
3. Define memory hierarchy → `memory-hierarchy.md`
|
||||||
|
4. Define memory ordering model → this document
|
||||||
|
5. Define atomic primitives → `atomics.md`
|
||||||
|
|
||||||
|
ASSUMPTION — The Ztso extension is a small RTL delta over RVWMO. The dominant implementation cost is verification, not area. No RTL exists in the repository to substantiate even a relative area claim, so this is treated as a plausible qualitative statement rather than a measured result.
|
||||||
|
|
||||||
|
OPEN QUESTION — Does XH-1 have a coherent accelerator fabric or non-coherent IO (e.g., CXL, DMA engines)? If so, the memory ordering model must define how these agents interact, and the repository does not currently do so.
|
||||||
|
|
||||||
|
## Verification Considerations
|
||||||
|
|
||||||
|
PROPOSAL — Verification challenges:
|
||||||
|
|
||||||
|
- **Litmus tests**: A standard RISC-V litmus test suite is available externally; whether XH-1 will adopt it as-is or define a custom suite is OPEN.
|
||||||
|
- **State-space explosion**: 128 cores × N outstanding operations produces a combinatorial unmitigated state space. The repository does not establish whether XH-1 will use state-space reduction, abstraction, bisimulation, or formal methods. Any specific XH-1 claim: INSUFFICIENT EVIDENCE.
|
||||||
|
- **Fence semantics**: Each fence variant must be exhaustively tested at corner cases (interrupts, exceptions, MMIO).
|
||||||
|
- **Coherence-ordering interaction**: The ordering model must be checked against the chosen coherence protocol for consistency. This check cannot be performed until the coherence protocol is defined.
|
||||||
|
|
||||||
|
INSUFFICIENT EVIDENCE — No XH-1 verification infrastructure, formal spec, or litmus-test set is documented in the repository.
|
||||||
|
|
||||||
|
ASSUMPTION — RVWMO is the most documented and tested model for RISC-V, per the ratified specification and ecosystem practice. Ztso is a small addition. SC and RC variants would require custom formal infrastructure. These are qualitative claims about ecosystem maturity, not XH-1-specific measurements.
|
||||||
|
|
||||||
|
## Software Considerations
|
||||||
|
|
||||||
|
PROPOSAL — Software-side impact, as a general property of each model:
|
||||||
|
|
||||||
|
- **Linux kernel**: Designed for weak memory models with appropriate fences. Compatible with RVWMO.
|
||||||
|
- **C11/C++ atomics**: Map cleanly onto release/acquire semantics. Compatible with RVWMO and with RC variants.
|
||||||
|
- **x86 software ports**: Benefit from Ztso for fewer fences.
|
||||||
|
- **OpenSHMEM, MPI shmem**: Often assume TSO-like ordering; Ztso helps.
|
||||||
|
- **HPC codes with hand-rolled atomics**: Highly sensitive to fence latency.
|
||||||
|
|
||||||
|
ASSUMPTION — XH-1 will run RISC-V Linux. The kernel's memory model expectations must be satisfied by XH-1's hardware ordering rules. The repository does not document whether Linux is in fact the target.
|
||||||
|
|
||||||
|
## Recommendation
|
||||||
|
|
||||||
|
INSUFFICIENT EVIDENCE to make a final recommendation.
|
||||||
|
|
||||||
|
The repository does not yet establish:
|
||||||
|
- Pipeline microarchitecture
|
||||||
|
- Coherence protocol
|
||||||
|
- Interconnect topology
|
||||||
|
- Workload target
|
||||||
|
|
||||||
|
A recommendation can be made only after the following documents in the memory research area are filled in: `memory-architecture.md`, `cache-coherency.md`, `memory-hierarchy.md`. Without these, any model selection is premature.
|
||||||
|
|
||||||
|
PROPOSAL — Tentative default, contingent on the stated assumptions holding and explicitly subject to revision once repository context is available:
|
||||||
|
|
||||||
|
1. **Baseline**: RVWMO (assumed mandatory for RISC-V ecosystem compatibility, if Linux is the target OS).
|
||||||
|
2. **Optional add-on**: Ztso extension for x86-software porting convenience and modest hardware cost, contingent on verification cost being acceptable.
|
||||||
|
3. **Avoid**: SC, custom RC variants, unless a strong workload or research driver emerges.
|
||||||
|
4. **Defer to**: `cache-coherency.md` for the coherence-side ordering mechanism.
|
||||||
|
|
||||||
|
ASSUMPTION — The tentative default above depends on (a) Linux being the target OS, (b) ecosystem compatibility being prioritized over x86-port convenience, and (c) the coherence and interconnect unknowns being resolvable without invalidating this default. None of these are established by the repository.
|
||||||
|
|
||||||
|
## Confidence
|
||||||
|
|
||||||
|
| Topic | Confidence | Reason |
|
||||||
|
|-------|------------|--------|
|
||||||
|
| RVWMO is the RISC-V baseline | High | Ratified ISA specification |
|
||||||
|
| Ztso is a small optional extension | High | Ratified extension in the ISA specification |
|
||||||
|
| SC has high implementation cost in general | High | General literature consensus |
|
||||||
|
| SC specifically impractical at 128 cores | Low | General claim, not XH-1-measured; workload- and topology-dependent |
|
||||||
|
| Per-fence energy scales with acknowledging-core count | Medium | Logical consequence; no measurement |
|
||||||
|
| Quantitative per-fence energy at 128 cores | None | INSUFFICIENT EVIDENCE |
|
||||||
|
| Verification is hardest for RC variants | Low–Medium | Literature claim, not XH-1-measured |
|
||||||
|
| XH-1 should adopt RVWMO | Low | Repository context not yet sufficient |
|
||||||
|
| XH-1 should adopt RVWMO + Ztso | Low | Same |
|
||||||
|
| Ztso is a small RTL delta over RVWMO | Low–Medium | Plausible qualitative claim; no RTL to substantiate |
|
||||||
|
|
||||||
|
## Open Questions
|
||||||
|
|
||||||
|
1. What is XH-1's microarchitecture style (in-order vs. out-of-order)?
|
||||||
|
2. What coherence protocol and directory organization will XH-1 use?
|
||||||
|
3. What is the interconnect topology, and does it preserve coherence-response order?
|
||||||
|
4. What workloads are targeted — HPC, server, embedded?
|
||||||
|
5. Does XH-1 include non-coherent agents (DMA, CXL, accelerators)?
|
||||||
|
6. Is XH-1 a research vehicle, a prototype, or a product?
|
||||||
|
7. What is the expected fence frequency in target software?
|
||||||
|
8. What is the verification methodology — formal, simulation, or hybrid?
|
||||||
|
9. Will XH-1 use the standard RISC-V litmus-test suite or define a custom one?
|
||||||
|
10. Does XH-1 require compatibility with x86-software-only ports? (Drives Ztso inclusion.)
|
||||||
|
11. What fence-latency, area, and energy budgets are acceptable for the target envelope?
|
||||||
|
12. What is the state-space reduction or formal-verification strategy, if any?
|
||||||
|
13. Which revision of the RISC-V Unprivileged ISA specification is the reference for XH-1?
|
||||||
|
|
||||||
|
## Sources
|
||||||
|
|
||||||
|
INSUFFICIENT EVIDENCE — The XH-1 Research repository does not yet contain source material specific to memory ordering. No citations, papers, or measurements can be honestly attributed to XH-1.
|
||||||
|
|
||||||
|
External ISA-level facts cited above (RVWMO, Ztso, FENCE instruction semantics) are drawn from the publicly ratified RISC-V Unprivileged ISA specification, which is the source of truth for the ISA-level memory model. The RISC-V ISA specification is the canonical reference; no specific paper, version, or page reference is provided because the XH-1 repository does not document which revision is being targeted.
|
||||||
|
|
||||||
|
INSUFFICIENT EVIDENCE — All quantitative claims about fence latency, area, power, energy, and verification cost at the XH-1 level have been removed or flagged. No benchmarks, no measurements, no RTL data, no simulation results are available in the repository at this time.
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
|
||||||
+1
File diff suppressed because one or more lines are too long
@@ -0,0 +1,319 @@
|
|||||||
|
# Memory Ordering
|
||||||
|
|
||||||
|
## Status
|
||||||
|
|
||||||
|
Stub document. XH-1 memory ordering model has not yet been defined. The repository contains no source artifacts, RTL, ISA extension proposals, or reference implementations that establish a memory consistency model for XH-1.
|
||||||
|
|
||||||
|
## Abstract
|
||||||
|
|
||||||
|
Memory ordering defines the rules by which loads and stores issued by one or more cores become visible to other cores in the system. For a 128-core RISC-V processor, the choice of memory model — and the hardware mechanisms that enforce it — is one of the most consequential architectural decisions, affecting correctness, performance, area, power, scalability, and software portability. This document surveys the design space, identifies the candidate models available within the RISC-V ISA framework, examines alternative hardware enforcement strategies, and frames the open questions for XH-1. No recommendation is made because the repository context does not yet establish pipeline style, coherency protocol, or interconnect topology — all of which constrain the memory ordering decision.
|
||||||
|
|
||||||
|
## Research Question
|
||||||
|
|
||||||
|
What memory consistency model and hardware ordering mechanism should XH-1 adopt, given:
|
||||||
|
|
||||||
|
1. A 128-core replication target
|
||||||
|
2. The RISC-V ISA baseline
|
||||||
|
3. The need to balance software portability against hardware implementation cost
|
||||||
|
4. Unknown microarchitecture style (in-order vs. out-of-order is not established in the repository)
|
||||||
|
5. The relationship to the cache coherency protocol (defined in `cache-coherency.md`, currently SOON)
|
||||||
|
|
||||||
|
## Background
|
||||||
|
|
||||||
|
### Memory Consistency Models — Terminology
|
||||||
|
|
||||||
|
PROPOSAL — Use the following working taxonomy in this document:
|
||||||
|
|
||||||
|
- **Sequential Consistency (SC)**: All cores observe a single total order of memory operations that respects each core's program order.
|
||||||
|
- **Total Store Order (TSO)**: Loads may bypass stores; stores from one core are seen in order by all cores. Used by x86.
|
||||||
|
- **Relaxed / Weak Memory Models**: Many orderings are allowed; fences (barriers) are required to enforce constraints. Used by ARM, RISC-V, and PowerPC.
|
||||||
|
- **Release Consistency (RC) / RCO / RCpc / RCsc**: A family of models with different fence semantics. Used by academic designs and the basis for much of C11/C++ memory model formalization.
|
||||||
|
|
||||||
|
### The RISC-V "RVWMO" Model
|
||||||
|
|
||||||
|
FACT — The RISC-V ISA specification defines the **RISC-V Weak Memory Ordering (RVWMO)** model as the baseline. Under RVWMO:
|
||||||
|
|
||||||
|
- Loads and stores to the same address from the same hart are ordered.
|
||||||
|
- A store followed by a load to a *different* address may be reordered.
|
||||||
|
- The `FENCE` instruction is used to enforce orderings. The base FENCE provides `pred` (predecessor) and `succ` (successor) bits covering device I/O (I), memory reads (R), and memory writes (W). Device ordering follows a stronger PostModel.
|
||||||
|
- The optional `Ztso` extension provides **RVTSO**, which gives x86-like TSO semantics with a relaxed form of the FENCE.
|
||||||
|
|
||||||
|
ASSUMPTION — RVWMO + the Ztso optional extension is the available ISA-level model space that XH-1 must select from. (The repository does not document any custom ISA extensions, so XH-1 is assumed to stay within ratified RISC-V.)
|
||||||
|
|
||||||
|
### Why Memory Ordering is Hard at 128 Cores
|
||||||
|
|
||||||
|
OPEN QUESTION — The repository does not establish:
|
||||||
|
- Whether XH-1 is a single die, multi-die package, or chiplet-based system
|
||||||
|
- The interconnect topology (mesh, ring, crossbar, NoC, etc.)
|
||||||
|
- Whether coherence is directory-based, broadcast-based, or hybrid
|
||||||
|
- The private-vs-shared L2/L3 split
|
||||||
|
- The memory-side directory or snoop filter organization
|
||||||
|
|
||||||
|
These decisions have direct bearing on which ordering mechanisms are feasible. Any hardware ordering scheme (invalidation queues, store buffers, fence acknowledgments, coherence-based ordering) must be analyzed in the context of these unknowns.
|
||||||
|
|
||||||
|
## Existing Approaches
|
||||||
|
|
||||||
|
### Approach A — RISC-V RVWMO (Baseline)
|
||||||
|
|
||||||
|
The standard weak memory model provided by the RISC-V ISA. Requires implementation of `FENCE`, `FENCE.I`, and the I/O ordering semantics.
|
||||||
|
|
||||||
|
**Mechanism**: Architecture permits aggressive reordering. Hardware typically uses:
|
||||||
|
- Store buffers with forwarding
|
||||||
|
- Load-load reordering queues
|
||||||
|
- Coherence-based ordering: stores become visible at coherence invalidation/upgrade points
|
||||||
|
- Fence implementation via global acknowledgment counters or by draining reorder queues
|
||||||
|
|
||||||
|
### Approach B — RVWMO + Ztso Extension
|
||||||
|
|
||||||
|
Adds RVTSO. Software that requires x86-like ordering can avoid fences in many cases. Still RISC-V compliant.
|
||||||
|
|
||||||
|
**Mechanism**: Stores from one hart are serialized in the interconnect; loads may still pass earlier stores to different addresses.
|
||||||
|
|
||||||
|
### Approach C — Sequential Consistency (SC)
|
||||||
|
|
||||||
|
All operations observed in a single global order respecting program order. Simplest software model; most expensive hardware.
|
||||||
|
|
||||||
|
**Mechanism**: Typically requires store-buffer draining on loads, broadcast-based coherence, and conservative replay mechanisms. Very costly for 128 cores.
|
||||||
|
|
||||||
|
### Approach D — Release Consistency variants (RCsc, RCpc, RCO)
|
||||||
|
|
||||||
|
Academic models where release/acquire ordering is cheap but stronger fences are required for full memory ordering. Basis of C11/C++ model.
|
||||||
|
|
||||||
|
**Mechanism**: Only release/acquire fences require acknowledgment; acquire-only loads are cheap. Alignment with C11 is a strong software benefit.
|
||||||
|
|
||||||
|
## Alternative Designs
|
||||||
|
|
||||||
|
### Hardware Ordering Mechanisms (to be evaluated once coherency and interconnect are known)
|
||||||
|
|
||||||
|
PROPOSAL — The following mechanisms should be considered for fence/ordering enforcement:
|
||||||
|
|
||||||
|
1. **Invalidation-acknowledgment fences** — A fence blocks until all outstanding coherence requests have been acknowledged. Simple, common, but scales poorly with directory hop count and aggregate invalidation traffic at 128 cores.
|
||||||
|
|
||||||
|
2. **Per-core store-buffer drain fences** — A fence stalls the issuing core until its store buffer is empty and all stores have been globally observed. Predictable latency; serializes through the coherence fabric.
|
||||||
|
|
||||||
|
3. **Tournament / Token-based fences** — Fences acquire a global token (Coherence-Next-Line style) or wait for a "fence epoch" counter. Avoids worst-case drain, but adds global state.
|
||||||
|
|
||||||
|
4. **FIFO coherence ordering** — Leverages the in-order completion of coherence transactions to provide ordering without explicit fences. Requires in-order interconnect, which conflicts with typical latency-optimized NoC designs.
|
||||||
|
|
||||||
|
5. **Time-to-live / epoch schemes** — Each coherence request carries an epoch tag; ordering is enforced at the L2/L3 directory by queueing requests by epoch.
|
||||||
|
|
||||||
|
6. **Speculative load reordering with rollback** — Loads can execute speculatively past stores; on conflict, the load is replayed. Common in high-performance out-of-order cores (e.g., IBM POWER).
|
||||||
|
|
||||||
|
ASSUMPTION — XH-1's microarchitecture style is not yet established. If out-of-order, the "speculative load reordering with rollback" category becomes a candidate. If in-order, mechanism choices narrow considerably.
|
||||||
|
|
||||||
|
## Comparison
|
||||||
|
|
||||||
|
| Property | SC | RVWMO | RVWMO + Ztso | RC variants |
|
||||||
|
|----------|----|----|----|----|
|
||||||
|
| Software ease | Highest | Medium | Medium-High | Low-Medium |
|
||||||
|
| Hardware cost | Highest | Low | Low-Medium | Low |
|
||||||
|
| Performance headroom | Lowest | Highest | High | High |
|
||||||
|
| Fence latency | N/A | High | Medium | Variable |
|
||||||
|
| Verification cost | Lowest | High | High | Highest |
|
||||||
|
| C11/C++ alignment | Approximate | Approximate | Weak | Direct |
|
||||||
|
| 128-core scalability | Poor | Good | Good | Good |
|
||||||
|
|
||||||
|
Quantitative figures: INSUFFICIENT EVIDENCE. All numbers above are relative qualitative ratings, not benchmarked measurements.
|
||||||
|
|
||||||
|
## Advantages
|
||||||
|
|
||||||
|
### RVWMO (Baseline)
|
||||||
|
- **Lowest hardware cost** of all candidates for the base case — most orderings are simply not enforced.
|
||||||
|
- **Maximum performance** potential — reordering is unconstrained.
|
||||||
|
- **Familiar** to RISC-V software ecosystem (Linux, RISC-V GCC, LLVM).
|
||||||
|
- **Standardized** — concrete compliance test suite exists.
|
||||||
|
|
||||||
|
### RVWMO + Ztso
|
||||||
|
- **x86-like TSO** for software that benefits from stronger ordering.
|
||||||
|
- **Backwards compatible** — software not using Ztso still runs correctly under RVWMO.
|
||||||
|
- **Modest hardware cost** — primarily requires store-store ordering at the coherence layer.
|
||||||
|
|
||||||
|
### Release Consistency Variants
|
||||||
|
- **Excellent fit for C11/C++** atomics — acquire/release are the most-used fence types in real code.
|
||||||
|
- **Cheapest common case** — acquire and release can be implemented with lightweight mechanisms.
|
||||||
|
- **Hardware cost is concentrated on the rare case** of full seq_cst fences.
|
||||||
|
|
||||||
|
## Disadvantages
|
||||||
|
|
||||||
|
### RVWMO (Baseline)
|
||||||
|
- **Fence latency is high** at 128 cores — global acknowledgment is expensive.
|
||||||
|
- **Verification complexity is high** — many legal reorderings; the litmus-test failure surface is large.
|
||||||
|
- **Software burden** — kernel and runtime code must insert fences correctly.
|
||||||
|
|
||||||
|
### RVWMO + Ztso
|
||||||
|
- **Two consistency models in one chip** — adds documentation, validation, and software education cost.
|
||||||
|
- **In-flight mixing** of ordering regimes in the same software is hard to reason about.
|
||||||
|
|
||||||
|
### Release Consistency Variants
|
||||||
|
- **Non-standard for RISC-V** — would require ISA extension. Cannot be recommended without a strong software-side driver.
|
||||||
|
- **Verification cost is the highest** of the candidates because of subtle fence semantics.
|
||||||
|
|
||||||
|
### Sequential Consistency
|
||||||
|
- **Performance cost is severe at 128 cores** — every load may stall on store-buffer drain.
|
||||||
|
- **No selective escape valve** — software cannot opt out of the strongest model.
|
||||||
|
- **Does not match RISC-V ecosystem expectations** — surprising to RISC-V software developers.
|
||||||
|
|
||||||
|
## XH-1 Considerations
|
||||||
|
|
||||||
|
OPEN QUESTION — The following XH-1-internal questions are unresolved by the repository:
|
||||||
|
|
||||||
|
1. Is XH-1's pipeline in-order or out-of-order?
|
||||||
|
2. What is the coherence protocol (MOESI, MESI, directory-based, broadcast)?
|
||||||
|
3. Is the interconnect a NoC, ring, or crossbar? Does it preserve in-order delivery of coherence responses?
|
||||||
|
4. What is the expected working set of a typical hart — does streaming-store optimization matter?
|
||||||
|
5. Does XH-1 target HPC, server, embedded, or mixed workloads? (Workload affects fence frequency.)
|
||||||
|
6. Is XH-1 a research vehicle (where SC is acceptable for simplicity) or a product (where RVWMO compliance is mandatory)?
|
||||||
|
7. Does the memory map use non-coherent regions (e.g., DMA, I/O) that interact with the ordering model?
|
||||||
|
|
||||||
|
ASSUMPTION — If XH-1 targets RISC-V ecosystem compatibility (Linux, RISC-V GCC, RISC-V LLVM), then RVWMO is mandatory. Ztso may optionally be added for x86-software porting convenience.
|
||||||
|
|
||||||
|
## 128-Core Scalability
|
||||||
|
|
||||||
|
PROPOSAL — Scalability analysis of each candidate at 128 cores:
|
||||||
|
|
||||||
|
### Fence Latency
|
||||||
|
- **Fence global acknowledgment** scales roughly with the worst-case coherence round-trip time. In a 128-core system, this can be many tens of cycles in a NoC, hundreds in a ring.
|
||||||
|
- **Store-buffer drain** latency is approximately proportional to the store buffer depth and the time to invalidate all sharers.
|
||||||
|
- **Token / epoch fences** scale better because they avoid draining the entire coherence fabric — they synchronize at a logical epoch boundary.
|
||||||
|
|
||||||
|
### Coherence Traffic
|
||||||
|
- **RVWMO fences** generate invalidation or acknowledgment traffic that competes with normal coherence traffic.
|
||||||
|
- **SC at 128 cores** can saturate a directory-based interconnect because every load's coherence transaction may stall pending store ordering.
|
||||||
|
- **RC variants** concentrate traffic on the fence path only.
|
||||||
|
|
||||||
|
### Verification Scalability
|
||||||
|
- **State-space explosion**: at 128 cores, the number of concurrently observable memory operations grows combinatorially. Memory-ordering verification is a known scaling bottleneck.
|
||||||
|
- **RC variants** are hardest to verify because of subtle fence interactions.
|
||||||
|
- **SC is easiest to verify** in principle, but worst in performance.
|
||||||
|
|
||||||
|
INSUFFICIENT EVIDENCE — No quantitative fence-latency or coherence-traffic data is available for XH-1. Estimates cannot be made without knowing the interconnect.
|
||||||
|
|
||||||
|
## Performance Considerations
|
||||||
|
|
||||||
|
PROPOSAL — Performance-impact dimensions to evaluate:
|
||||||
|
|
||||||
|
1. **Steady-state IPC** for memory-bound workloads — RVWMO best, SC worst.
|
||||||
|
2. **Fence latency** — measured as cycles from fence issue to fence completion. Dominant for synchronization-heavy code.
|
||||||
|
3. **Scalable synchronization** — how the model interacts with lock implementations, RCU, and atomic primitives.
|
||||||
|
4. **Load-use latency** — whether loads can bypass earlier stores (allowed in RVWMO and TSO, not in SC).
|
||||||
|
5. **Coherence miss rate** interaction — does the ordering model inflate or deflate cache-line ping-pong?
|
||||||
|
|
||||||
|
ASSUMPTION — For HPC workloads with high-fence frequency (e.g., MPI, OpenMP, lock-heavy code), even small differences in fence latency will dominate. For server workloads with mixed locking, the cost of fences is amortized over longer critical sections.
|
||||||
|
|
||||||
|
## Area Considerations
|
||||||
|
|
||||||
|
PROPOSAL — Area cost of ordering mechanisms:
|
||||||
|
|
||||||
|
| Mechanism | Approximate area cost per core |
|
||||||
|
|-----------|-------------------------------|
|
||||||
|
| Store buffer with forwarding | Small (few entries × cache-line width) |
|
||||||
|
| Load-load reordering queue | Small to medium (grows with IPC) |
|
||||||
|
| Fence-acknowledgment counter | Negligible (per-core counter) |
|
||||||
|
| Global epoch / token state | Medium (one per chip, not per core) |
|
||||||
|
| Speculative load reorder + rollback | Large (load-load queue, replay logic) |
|
||||||
|
|
||||||
|
ASSUMPTION — All numbers in the table are order-of-magnitude estimates. Precise XH-1 figures: INSUFFICIENT EVIDENCE.
|
||||||
|
|
||||||
|
PROPOSAL — At 128 cores, even small per-core area costs multiply by 128. Global state (epoch counters, fence acknowledgers) is O(1) and becomes attractive.
|
||||||
|
|
||||||
|
## Power and Energy Considerations
|
||||||
|
|
||||||
|
PROPOSAL — Energy implications:
|
||||||
|
|
||||||
|
- **Fences** generate global coherence traffic — energy cost scales with the number of cores that must acknowledge.
|
||||||
|
- **Speculative load reordering** wastes energy on rollback when a violation is detected.
|
||||||
|
- **Strong ordering models** keep more coherence state in flight (more invalidations, more retries).
|
||||||
|
- **RVWMO** with lightweight fences (when software permits) is the lowest-energy regime.
|
||||||
|
|
||||||
|
OPEN QUESTION — The 128-core replication means the per-fence energy cost is multiplied. A fence that costs 100 pJ per core at 16 cores costs 800× at 128 cores in the worst case (if all cores must acknowledge). No published figure for XH-1 is available — INSUFFICIENT EVIDENCE.
|
||||||
|
|
||||||
|
## Implementation Considerations
|
||||||
|
|
||||||
|
PROPOSAL — Implementation order of dependencies (the repository does not establish these, but the research area suggests the following):
|
||||||
|
|
||||||
|
1. Define coherency protocol → `cache-coherency.md`
|
||||||
|
2. Define interconnect topology → not yet in repository
|
||||||
|
3. Define memory hierarchy → `memory-hierarchy.md`
|
||||||
|
4. Define memory ordering model → this document
|
||||||
|
5. Define atomic primitives → `atomics.md`
|
||||||
|
|
||||||
|
ASSUMPTION — The Ztso extension is a small RTL delta over RVWMO. The dominant implementation cost is verification, not area.
|
||||||
|
|
||||||
|
OPEN QUESTION — Does XH-1 have a coherent accelerator fabric or non-coherent IO (e.g., CXL, DMA engines)? If so, the memory ordering model must define how these agents interact.
|
||||||
|
|
||||||
|
## Verification Considerations
|
||||||
|
|
||||||
|
PROPOSAL — Verification challenges:
|
||||||
|
|
||||||
|
- **Litmus tests**: Standard RISC-V litmus test suite must be run on RTL and on model-based environments.
|
||||||
|
- **State-space explosion**: 128 cores × N outstanding operations = combinatorial explosion. Likely requires state-space reduction or bisimulation techniques.
|
||||||
|
- **Fence semantics**: Each fence variant must be exhaustively tested at corner cases (interrupts, exceptions, MMIO).
|
||||||
|
- **Coherence-ordering interaction**: The ordering model must be checked against the chosen coherence protocol for consistency.
|
||||||
|
|
||||||
|
INSUFFICIENT EVIDENCE — No XH-1 verification infrastructure, formal spec, or litmus-test set is documented in the repository.
|
||||||
|
|
||||||
|
ASSUMPTION — RVWMO is the most documented and tested model for RISC-V. Ztso is a small addition. SC and RC variants would require custom formal infrastructure.
|
||||||
|
|
||||||
|
## Software Considerations
|
||||||
|
|
||||||
|
PROPOSAL — Software-side impact:
|
||||||
|
|
||||||
|
- **Linux kernel**: Designed for weak memory models with appropriate fences. Compatible with RVWMO.
|
||||||
|
- **C11/C++ atomics**: Map cleanly onto release/acquire semantics. Compatible with RVWMO and with RC variants.
|
||||||
|
- **x86 software ports**: Benefit from Ztso for fewer fences.
|
||||||
|
- **OpenSHMEM, MPI shmem**: Often assume TSO-like ordering; Ztso helps.
|
||||||
|
- **HPC codes with hand-rolled atomics**: Highly sensitive to fence latency.
|
||||||
|
|
||||||
|
ASSUMPTION — XH-1 will run RISC-V Linux. The kernel's memory model expectations must be satisfied by XH-1's hardware ordering rules.
|
||||||
|
|
||||||
|
## Recommendation
|
||||||
|
|
||||||
|
INSUFFICIENT EVIDENCE to make a final recommendation.
|
||||||
|
|
||||||
|
The repository does not yet establish:
|
||||||
|
- Pipeline microarchitecture
|
||||||
|
- Coherence protocol
|
||||||
|
- Interconnect topology
|
||||||
|
- Workload target
|
||||||
|
|
||||||
|
A recommendation can be made only after the following documents in the memory research area are filled in: `memory-architecture.md`, `cache-coherency.md`, `memory-hierarchy.md`. Without these, any model selection is premature.
|
||||||
|
|
||||||
|
PROPOSAL — Tentative default pending more repository context:
|
||||||
|
|
||||||
|
1. **Baseline**: RVWMO (mandatory for RISC-V ecosystem).
|
||||||
|
2. **Optional add-on**: Ztso extension for x86-software porting convenience and modest hardware cost.
|
||||||
|
3. **Avoid**: SC, custom RC variants, unless a strong workload or research driver emerges.
|
||||||
|
4. **Defer to**: `cache-coherency.md` for the coherence-side ordering mechanism.
|
||||||
|
|
||||||
|
## Confidence
|
||||||
|
|
||||||
|
| Topic | Confidence | Reason |
|
||||||
|
|-------|------------|--------|
|
||||||
|
| RVWMO is the RISC-V baseline | High | Ratified ISA spec |
|
||||||
|
| Ztso is a small optional extension | High | Ratified extension |
|
||||||
|
| SC is impractical at 128 cores | Medium | General consensus in literature; XH-1-specific number not available |
|
||||||
|
| Per-fence energy scales with core count | Medium | Logical consequence; no measurement |
|
||||||
|
| Verification is hardest for RC variants | Medium | General consensus |
|
||||||
|
| XH-1 should adopt RVWMO | Low | Repository context not yet sufficient |
|
||||||
|
| XH-1 should adopt RVWMO + Ztso | Low | Same |
|
||||||
|
|
||||||
|
## Open Questions
|
||||||
|
|
||||||
|
1. What is XH-1's microarchitecture style (in-order vs. out-of-order)?
|
||||||
|
2. What coherence protocol and directory organization will XH-1 use?
|
||||||
|
3. What is the interconnect topology, and does it preserve coherence-response order?
|
||||||
|
4. What workloads are targeted — HPC, server, embedded?
|
||||||
|
5. Does XH-1 include non-coherent agents (DMA, CXL, accelerators)?
|
||||||
|
6. Is XH-1 a research vehicle, a prototype, or a product?
|
||||||
|
7. What is the expected fence frequency in target software?
|
||||||
|
8. What is the verification methodology — formal, simulation, or hybrid?
|
||||||
|
9. Will XH-1 use the standard RISC-V litmus-test suite or define a custom one?
|
||||||
|
10. Does XH-1 require compatibility with x86-software-only ports? (Drives Ztso inclusion.)
|
||||||
|
|
||||||
|
## Sources
|
||||||
|
|
||||||
|
INSUFFICIENT EVIDENCE — The XH-1 Research repository does not yet contain source material specific to memory ordering. No citations, papers, or measurements can be honestly attributed to XH-1.
|
||||||
|
|
||||||
|
External ISA-level facts cited above (RVWMO, Ztso, FENCE instruction) are drawn from the publicly ratified RISC-V ISA specification, which is the source of truth for the ISA-level memory model. No specific paper, version, or page reference is provided because the XH-1 repository does not document which revision is being targeted.
|
||||||
|
|
||||||
|
No benchmarks, no measurements, no RTL data, no simulation results are available in the repository at this time.
|
||||||
+235
File diff suppressed because one or more lines are too long
@@ -0,0 +1,323 @@
|
|||||||
|
# Memory Ordering
|
||||||
|
|
||||||
|
## Status
|
||||||
|
|
||||||
|
Stub document. XH-1 memory ordering model has not yet been defined. The repository contains no source artifacts, RTL, ISA extension proposals, or reference implementations that establish a memory consistency model for XH-1.
|
||||||
|
|
||||||
|
## Abstract
|
||||||
|
|
||||||
|
Memory ordering defines the rules by which loads and stores issued by one or more cores become visible to other cores in the system. For a 128-core RISC-V processor, the choice of memory model — and the hardware mechanisms that enforce it — is one of the most consequential architectural decisions, affecting correctness, performance, area, power, scalability, and software portability. This document surveys the design space, identifies the candidate models available within the RISC-V ISA framework, examines alternative hardware enforcement strategies, and frames the open questions for XH-1. No recommendation is made because the repository context does not yet establish pipeline style, coherency protocol, or interconnect topology — all of which constrain the memory ordering decision.
|
||||||
|
|
||||||
|
## Research Question
|
||||||
|
|
||||||
|
What memory consistency model and hardware ordering mechanism should XH-1 adopt, given:
|
||||||
|
|
||||||
|
1. A 128-core replication target
|
||||||
|
2. The RISC-V ISA baseline
|
||||||
|
3. The need to balance software portability against hardware implementation cost
|
||||||
|
4. Unknown microarchitecture style (in-order vs. out-of-order is not established in the repository)
|
||||||
|
5. The relationship to the cache coherency protocol (defined in `cache-coherency.md`, currently SOON)
|
||||||
|
|
||||||
|
## Background
|
||||||
|
|
||||||
|
### Memory Consistency Models — Terminology
|
||||||
|
|
||||||
|
PROPOSAL — Use the following working taxonomy in this document:
|
||||||
|
|
||||||
|
- **Sequential Consistency (SC)**: All cores observe a single total order of memory operations that respects each core's program order.
|
||||||
|
- **Total Store Order (TSO)**: Loads may bypass stores; stores from one core are seen in order by all cores. Used by x86.
|
||||||
|
- **Relaxed / Weak Memory Models**: Many orderings are allowed; fences (barriers) are required to enforce constraints. Used by ARM, RISC-V, and PowerPC.
|
||||||
|
- **Release Consistency (RC) / RCO / RCpc / RCsc**: A family of models with different fence semantics. Used by academic designs and the basis for much of the C11/C++ memory model formalization.
|
||||||
|
|
||||||
|
### The RISC-V "RVWMO" Model
|
||||||
|
|
||||||
|
FACT — The RISC-V ISA specification defines the **RISC-V Weak Memory Ordering (RVWMO)** model as the baseline. Under RVWMO:
|
||||||
|
|
||||||
|
- Loads and stores to the same address from the same hart are ordered.
|
||||||
|
- A store followed by a load to a *different* address may be reordered.
|
||||||
|
- The `FENCE` instruction is used to enforce orderings. The base FENCE provides `pred` (predecessor) and `succ` (successor) bits covering device I/O (I), memory reads (R), and memory writes (W). Device ordering follows a stronger "Preserved over Higher Privilege and PMA" model (PostModel).
|
||||||
|
- The optional `Ztso` extension provides **RVTSO**, which gives x86-like TSO semantics with a relaxed form of the FENCE.
|
||||||
|
|
||||||
|
ASSUMPTION — RVWMO + the Ztso optional extension is the available ISA-level model space that XH-1 must select from. (The repository does not document any custom ISA extensions, so XH-1 is assumed to stay within ratified RISC-V.)
|
||||||
|
|
||||||
|
### Why Memory Ordering is Hard at 128 Cores
|
||||||
|
|
||||||
|
OPEN QUESTION — The repository does not establish:
|
||||||
|
- Whether XH-1 is a single die, multi-die package, or chiplet-based system
|
||||||
|
- The interconnect topology (mesh, ring, crossbar, NoC, etc.)
|
||||||
|
- Whether coherence is directory-based, broadcast-based, or hybrid
|
||||||
|
- The private-vs-shared L2/L3 split
|
||||||
|
- The memory-side directory or snoop filter organization
|
||||||
|
|
||||||
|
These decisions have direct bearing on which ordering mechanisms are feasible. Any hardware ordering scheme (invalidation queues, store buffers, fence acknowledgments, coherence-based ordering) must be analyzed in the context of these unknowns.
|
||||||
|
|
||||||
|
## Existing Approaches
|
||||||
|
|
||||||
|
### Approach A — RISC-V RVWMO (Baseline)
|
||||||
|
|
||||||
|
The standard weak memory model provided by the RISC-V ISA. Requires implementation of `FENCE`, `FENCE.I`, and the I/O ordering semantics.
|
||||||
|
|
||||||
|
**Mechanism**: Architecture permits aggressive reordering. Hardware typically uses:
|
||||||
|
- Store buffers with forwarding
|
||||||
|
- Load-load reordering queues
|
||||||
|
- Coherence-based ordering: stores become visible at coherence invalidation/upgrade points
|
||||||
|
- Fence implementation via global acknowledgment counters or by draining reorder queues
|
||||||
|
|
||||||
|
### Approach B — RVWMO + Ztso Extension
|
||||||
|
|
||||||
|
Adds RVTSO. Software that requires x86-like ordering can avoid fences in many cases. Still RISC-V compliant.
|
||||||
|
|
||||||
|
**Mechanism**: Stores from one hart are serialized in the interconnect; loads may still pass earlier stores to different addresses.
|
||||||
|
|
||||||
|
### Approach C — Sequential Consistency (SC)
|
||||||
|
|
||||||
|
All operations observed in a single global order respecting program order. Simplest software model; most expensive hardware.
|
||||||
|
|
||||||
|
**Mechanism**: Typically requires store-buffer draining on loads, broadcast-based coherence, and conservative replay mechanisms. Implementation cost is high in general; whether it is "very costly" specifically at 128 cores is workload- and topology-dependent and is treated here as a hypothesis, not a known result.
|
||||||
|
|
||||||
|
### Approach D — Release Consistency variants (RCsc, RCpc, RCO)
|
||||||
|
|
||||||
|
Academic models where release/acquire ordering is cheap but stronger fences are required for full memory ordering. Basis of C11/C++ model.
|
||||||
|
|
||||||
|
**Mechanism**: Only release/acquire fences require acknowledgment; acquire-only loads are cheap. Alignment with C11 is a strong software benefit. These variants are not part of the ratified RISC-V ISA, so adoption would require a custom ISA extension.
|
||||||
|
|
||||||
|
## Alternative Designs
|
||||||
|
|
||||||
|
### Hardware Ordering Mechanisms (to be evaluated once coherency and interconnect are known)
|
||||||
|
|
||||||
|
PROPOSAL — The following mechanisms should be considered for fence/ordering enforcement:
|
||||||
|
|
||||||
|
1. **Invalidation-acknowledgment fences** — A fence blocks until all outstanding coherence requests have been acknowledged. Simple and common; scaling impact depends on directory hop count, aggregate invalidation traffic, and invalidation filter structure.
|
||||||
|
|
||||||
|
2. **Per-core store-buffer drain fences** — A fence stalls the issuing core until its store buffer is empty and all stores have been globally observed. Predictable latency; serializes through the coherence fabric.
|
||||||
|
|
||||||
|
3. **Tournament / Token-based fences** — Fences acquire a global token or wait for a "fence epoch" counter. Avoids worst-case drain, but adds global state.
|
||||||
|
|
||||||
|
4. **FIFO coherence ordering** — Leverages the in-order completion of coherence transactions to provide ordering without explicit fences. Requires in-order interconnect, which conflicts with typical latency-optimized NoC designs.
|
||||||
|
|
||||||
|
5. **Time-to-live / epoch schemes** — Each coherence request carries an epoch tag; ordering is enforced at the L2/L3 directory by queueing requests by epoch.
|
||||||
|
|
||||||
|
6. **Speculative load reordering with rollback** — Loads can execute speculatively past stores; on conflict, the load is replayed. Common in high-performance out-of-order cores (e.g., IBM POWER).
|
||||||
|
|
||||||
|
ASSUMPTION — XH-1's microarchitecture style is not yet established. If out-of-order, the "speculative load reordering with rollback" category becomes a candidate. If in-order, mechanism choices narrow considerably.
|
||||||
|
|
||||||
|
## Comparison
|
||||||
|
|
||||||
|
| Property | SC | RVWMO | RVWMO + Ztso | RC variants |
|
||||||
|
|----------|----|----|----|----|
|
||||||
|
| Software ease | Highest | Medium | Medium-High | Low-Medium |
|
||||||
|
| Hardware cost | Highest | Low | Low-Medium | Low |
|
||||||
|
| Performance headroom | Lowest | Highest | High | High |
|
||||||
|
| Fence latency | N/A | Variable | Variable | Variable |
|
||||||
|
| Verification cost | Lowest in principle | High | High | High (likely highest) |
|
||||||
|
| C11/C++ alignment | Approximate | Approximate | Weak | Direct |
|
||||||
|
| 128-core scalability | Poor (hypothesis) | Good (claim) | Good (claim) | Good (claim) |
|
||||||
|
|
||||||
|
INSUFFICIENT EVIDENCE — Quantitative fence-latency, coherence-traffic, and scalability numbers for XH-1 are unavailable. Several rows above contain qualitative judgments rather than measured results, and the "128-core scalability" and "verification cost" rows reflect general literature consensus, not XH-1 measurements.
|
||||||
|
|
||||||
|
## Advantages
|
||||||
|
|
||||||
|
### RVWMO (Baseline)
|
||||||
|
- **Lowest hardware cost** of all candidates for the base case — most orderings are simply not enforced.
|
||||||
|
- **Maximum performance** potential — reordering is unconstrained.
|
||||||
|
- **Familiar** to RISC-V software ecosystem (Linux, RISC-V GCC, LLVM).
|
||||||
|
- **Standardized** — concrete compliance test suite exists.
|
||||||
|
|
||||||
|
### RVWMO + Ztso
|
||||||
|
- **x86-like TSO** for software that benefits from stronger ordering.
|
||||||
|
- **Backwards compatible** — software not using Ztso still runs correctly under RVWMO.
|
||||||
|
- **Modest hardware cost** — primarily requires store-store ordering at the coherence layer.
|
||||||
|
|
||||||
|
### Release Consistency Variants
|
||||||
|
- **Excellent fit for C11/C++** atomics — acquire/release are the most-used fence types in real code.
|
||||||
|
- **Cheapest common case** — acquire and release can be implemented with lightweight mechanisms.
|
||||||
|
- **Hardware cost is concentrated on the rare case** of full seq_cst fences.
|
||||||
|
|
||||||
|
## Disadvantages
|
||||||
|
|
||||||
|
### RVWMO (Baseline)
|
||||||
|
- **Fence latency depends on coherence round-trip** — at 128 cores, fence latency is expected to be high because global acknowledgment must traverse the coherence fabric. Specific cycle counts: INSUFFICIENT EVIDENCE.
|
||||||
|
- **Verification complexity is high** — many legal reorderings; the litmus-test failure surface is large.
|
||||||
|
- **Software burden** — kernel and runtime code must insert fences correctly.
|
||||||
|
|
||||||
|
### RVWMO + Ztso
|
||||||
|
- **Two consistency models in one chip** — adds documentation, validation, and software education cost.
|
||||||
|
- **In-flight mixing** of ordering regimes in the same software is hard to reason about.
|
||||||
|
|
||||||
|
### Release Consistency Variants
|
||||||
|
- **Non-standard for RISC-V** — would require a custom ISA extension. Cannot be recommended without a strong software-side driver.
|
||||||
|
- **Verification cost is high** in general because of subtle fence semantics; whether it is the "highest" among candidates is a claim about the literature, not a measured XH-1 result.
|
||||||
|
|
||||||
|
### Sequential Consistency
|
||||||
|
- **Performance cost is severe in the general case** — every load may stall on store-buffer drain; whether this is "severe at 128 cores" specifically is workload- and topology-dependent and is not established by the repository.
|
||||||
|
- **No selective escape valve** — software cannot opt out of the strongest model.
|
||||||
|
- **Does not match RISC-V ecosystem expectations** — surprising to RISC-V software developers.
|
||||||
|
|
||||||
|
## XH-1 Considerations
|
||||||
|
|
||||||
|
OPEN QUESTION — The following XH-1-internal questions are unresolved by the repository:
|
||||||
|
|
||||||
|
1. Is XH-1's pipeline in-order or out-of-order?
|
||||||
|
2. What is the coherence protocol (MOESI, MESI, directory-based, broadcast)?
|
||||||
|
3. Is the interconnect a NoC, ring, or crossbar? Does it preserve in-order delivery of coherence responses?
|
||||||
|
4. What is the expected working set of a typical hart — does streaming-store optimization matter?
|
||||||
|
5. Does XH-1 target HPC, server, embedded, or mixed workloads? (Workload affects fence frequency.)
|
||||||
|
6. Is XH-1 a research vehicle (where SC is acceptable for simplicity) or a product (where RVWMO compliance is mandatory)?
|
||||||
|
7. Does the memory map use non-coherent regions (e.g., DMA, I/O) that interact with the ordering model?
|
||||||
|
|
||||||
|
ASSUMPTION — If XH-1 targets RISC-V ecosystem compatibility (Linux, RISC-V GCC, RISC-V LLVM), then RVWMO is mandatory. Ztso may optionally be added for x86-software porting convenience.
|
||||||
|
|
||||||
|
## 128-Core Scalability
|
||||||
|
|
||||||
|
PROPOSAL — Scalability analysis of each candidate at 128 cores. Quantitative figures: INSUFFICIENT EVIDENCE. The qualitative observations below are hypotheses grounded in general literature, not XH-1 measurements.
|
||||||
|
|
||||||
|
### Fence Latency
|
||||||
|
- **Fence global acknowledgment** latency is bounded below by the worst-case coherence round-trip time. In a 128-core system, this can plausibly reach many tens of cycles in a NoC and higher in a ring or multi-hop topology; specific XH-1 figures: INSUFFICIENT EVIDENCE.
|
||||||
|
- **Store-buffer drain** latency is approximately proportional to the store buffer depth and the time to invalidate all sharers.
|
||||||
|
- **Token / epoch fences** may scale better because they avoid draining the entire coherence fabric — they synchronize at a logical epoch boundary.
|
||||||
|
|
||||||
|
### Coherence Traffic
|
||||||
|
- **RVWMO fences** generate invalidation or acknowledgment traffic that competes with normal coherence traffic. The magnitude depends on directory organization, snoop-filter effectiveness, and share-set size.
|
||||||
|
- **SC at 128 cores** can plausibly saturate a directory-based interconnect because every load's coherence transaction may stall pending store ordering; this is a hypothesis, not a measured XH-1 result.
|
||||||
|
- **RC variants** concentrate traffic on the fence path only.
|
||||||
|
|
||||||
|
### Verification Scalability
|
||||||
|
- **State-space explosion**: at 128 cores, the number of concurrently observable memory operations grows combinatorially in the unmitigated model. Memory-ordering verification is a known scaling bottleneck in the literature, and the repository does not establish what reduction or abstraction techniques XH-1 will employ.
|
||||||
|
- **RC variants** are commonly cited as among the hardest to verify because of subtle fence interactions; whether they are hardest in the absolute sense is a literature claim, not an XH-1 measurement.
|
||||||
|
- **SC is easiest to verify in principle** because it admits fewer legal behaviors, but worst in performance.
|
||||||
|
|
||||||
|
INSUFFICIENT EVIDENCE — No quantitative fence-latency or coherence-traffic data is available for XH-1. Estimates cannot be made without knowing the interconnect.
|
||||||
|
|
||||||
|
## Performance Considerations
|
||||||
|
|
||||||
|
PROPOSAL — Performance-impact dimensions to evaluate:
|
||||||
|
|
||||||
|
1. **Steady-state IPC** for memory-bound workloads — qualitatively, RVWMO permits the most reordering and SC the least; XH-1-specific IPC: INSUFFICIENT EVIDENCE.
|
||||||
|
2. **Fence latency** — measured as cycles from fence issue to fence completion. Dominant for synchronization-heavy code. XH-1-specific: INSUFFICIENT EVIDENCE.
|
||||||
|
3. **Scalable synchronization** — how the model interacts with lock implementations, RCU, and atomic primitives.
|
||||||
|
4. **Load-use latency** — whether loads can bypass earlier stores (allowed in RVWMO and TSO, not in SC).
|
||||||
|
5. **Coherence miss rate** interaction — does the ordering model inflate or deflate cache-line ping-pong? Model-dependent; XH-1-specific: INSUFFICIENT EVIDENCE.
|
||||||
|
|
||||||
|
ASSUMPTION — For HPC workloads with high-fence frequency (e.g., MPI, OpenMP, lock-heavy code), even small differences in fence latency will dominate. For server workloads with mixed locking, the cost of fences is amortized over longer critical sections. No XH-1 workload profile is documented in the repository.
|
||||||
|
|
||||||
|
## Area Considerations
|
||||||
|
|
||||||
|
PROPOSAL — Qualitative area cost of ordering mechanisms:
|
||||||
|
|
||||||
|
| Mechanism | Approximate area cost per core |
|
||||||
|
|-----------|-------------------------------|
|
||||||
|
| Store buffer with forwarding | Small (few entries × cache-line width) |
|
||||||
|
| Load-load reordering queue | Small to medium (grows with IPC and reorder depth) |
|
||||||
|
| Fence-acknowledgment counter | Negligible (per-core counter) |
|
||||||
|
| Global epoch / token state | Medium (one per chip, not per core) |
|
||||||
|
| Speculative load reorder + rollback | Large (load-load queue, replay logic) |
|
||||||
|
|
||||||
|
INSUFFICIENT EVIDENCE — All descriptions above are qualitative. Precise XH-1 area figures (in gates, µm², or mm²) are unavailable.
|
||||||
|
|
||||||
|
PROPOSAL — At 128 cores, even small per-core area costs multiply by 128. Global state (epoch counters, fence acknowledgers) is O(1) per chip and therefore becomes attractive on a per-core basis, but requires careful analysis of critical-path and clock-distribution impact, which is not established by the repository.
|
||||||
|
|
||||||
|
## Power and Energy Considerations
|
||||||
|
|
||||||
|
PROPOSAL — Qualitative energy implications:
|
||||||
|
|
||||||
|
- **Fences** generate global coherence traffic — energy cost scales with the number of cores that must acknowledge. The aggregate chip-level energy per fence is therefore expected to grow with core count.
|
||||||
|
- **Speculative load reordering** wastes energy on rollback when a violation is detected.
|
||||||
|
- **Strong ordering models** keep more coherence state in flight (more invalidations, more retries).
|
||||||
|
- **RVWMO** with lightweight fences (when software permits) is the lowest-energy regime in the general case.
|
||||||
|
|
||||||
|
INSUFFICIENT EVIDENCE — The earlier text claimed that a 100 pJ-per-core fence at 16 cores costs 800× at 128 cores in the worst case (if all cores must acknowledge). This multiplication assumed 100 pJ is a measured or even a defensible per-core figure, which is not established by the repository or by a cited source. The 100 pJ figure is removed. A defensible claim is only that per-fence chip-level energy is expected to grow with the number of acknowledging cores; quantitative XH-1 values: INSUFFICIENT EVIDENCE.
|
||||||
|
|
||||||
|
## Implementation Considerations
|
||||||
|
|
||||||
|
PROPOSAL — Implementation order of dependencies (the repository does not establish these, but the research area suggests the following):
|
||||||
|
|
||||||
|
1. Define coherency protocol → `cache-coherency.md`
|
||||||
|
2. Define interconnect topology → not yet in repository
|
||||||
|
3. Define memory hierarchy → `memory-hierarchy.md`
|
||||||
|
4. Define memory ordering model → this document
|
||||||
|
5. Define atomic primitives → `atomics.md`
|
||||||
|
|
||||||
|
ASSUMPTION — The Ztso extension is a small RTL delta over RVWMO. The dominant implementation cost is verification, not area. No RTL exists in the repository to substantiate even a relative area claim, so this is treated as a plausible qualitative statement rather than a measured result.
|
||||||
|
|
||||||
|
OPEN QUESTION — Does XH-1 have a coherent accelerator fabric or non-coherent IO (e.g., CXL, DMA engines)? If so, the memory ordering model must define how these agents interact, and the repository does not currently do so.
|
||||||
|
|
||||||
|
## Verification Considerations
|
||||||
|
|
||||||
|
PROPOSAL — Verification challenges:
|
||||||
|
|
||||||
|
- **Litmus tests**: A standard RISC-V litmus test suite is available externally; whether XH-1 will adopt it as-is or define a custom suite is OPEN.
|
||||||
|
- **State-space explosion**: 128 cores × N outstanding operations produces a combinatorial unmitigated state space. The repository does not establish whether XH-1 will use state-space reduction, abstraction, bisimulation, or formal methods. Any specific XH-1 claim: INSUFFICIENT EVIDENCE.
|
||||||
|
- **Fence semantics**: Each fence variant must be exhaustively tested at corner cases (interrupts, exceptions, MMIO).
|
||||||
|
- **Coherence-ordering interaction**: The ordering model must be checked against the chosen coherence protocol for consistency. This check cannot be performed until the coherence protocol is defined.
|
||||||
|
|
||||||
|
INSUFFICIENT EVIDENCE — No XH-1 verification infrastructure, formal spec, or litmus-test set is documented in the repository.
|
||||||
|
|
||||||
|
ASSUMPTION — RVWMO is the most documented and tested model for RISC-V. Ztso is a small addition. SC and RC variants would require custom formal infrastructure. These are qualitative claims about ecosystem maturity, not XH-1-specific measurements.
|
||||||
|
|
||||||
|
## Software Considerations
|
||||||
|
|
||||||
|
PROPOSAL — Software-side impact:
|
||||||
|
|
||||||
|
- **Linux kernel**: Designed for weak memory models with appropriate fences. Compatible with RVWMO.
|
||||||
|
- **C11/C++ atomics**: Map cleanly onto release/acquire semantics. Compatible with RVWMO and with RC variants.
|
||||||
|
- **x86 software ports**: Benefit from Ztso for fewer fences.
|
||||||
|
- **OpenSHMEM, MPI shmem**: Often assume TSO-like ordering; Ztso helps.
|
||||||
|
- **HPC codes with hand-rolled atomics**: Highly sensitive to fence latency.
|
||||||
|
|
||||||
|
ASSUMPTION — XH-1 will run RISC-V Linux. The kernel's memory model expectations must be satisfied by XH-1's hardware ordering rules. The repository does not document whether Linux is in fact the target.
|
||||||
|
|
||||||
|
## Recommendation
|
||||||
|
|
||||||
|
INSUFFICIENT EVIDENCE to make a final recommendation.
|
||||||
|
|
||||||
|
The repository does not yet establish:
|
||||||
|
- Pipeline microarchitecture
|
||||||
|
- Coherence protocol
|
||||||
|
- Interconnect topology
|
||||||
|
- Workload target
|
||||||
|
|
||||||
|
A recommendation can be made only after the following documents in the memory research area are filled in: `memory-architecture.md`, `cache-coherency.md`, `memory-hierarchy.md`. Without these, any model selection is premature.
|
||||||
|
|
||||||
|
PROPOSAL — Tentative default pending more repository context:
|
||||||
|
|
||||||
|
1. **Baseline**: RVWMO (mandatory for RISC-V ecosystem compatibility, if Linux is the target OS).
|
||||||
|
2. **Optional add-on**: Ztso extension for x86-software porting convenience and modest hardware cost, contingent on verification cost being acceptable.
|
||||||
|
3. **Avoid**: SC, custom RC variants, unless a strong workload or research driver emerges.
|
||||||
|
4. **Defer to**: `cache-coherency.md` for the coherence-side ordering mechanism.
|
||||||
|
|
||||||
|
## Confidence
|
||||||
|
|
||||||
|
| Topic | Confidence | Reason |
|
||||||
|
|-------|------------|--------|
|
||||||
|
| RVWMO is the RISC-V baseline | High | Ratified ISA spec |
|
||||||
|
| Ztso is a small optional extension | High | Ratified extension |
|
||||||
|
| SC has high implementation cost | High | General literature consensus |
|
||||||
|
| SC specifically impractical at 128 cores | Low | General claim, not XH-1-measured; workload- and topology-dependent |
|
||||||
|
| Per-fence energy scales with acknowledging-core count | Medium | Logical consequence; no measurement |
|
||||||
|
| Quantitative per-fence energy at 128 cores | None | INSUFFICIENT EVIDENCE |
|
||||||
|
| Verification is hardest for RC variants | Low–Medium | Literature claim, not XH-1-measured |
|
||||||
|
| XH-1 should adopt RVWMO | Low | Repository context not yet sufficient |
|
||||||
|
| XH-1 should adopt RVWMO + Ztso | Low | Same |
|
||||||
|
|
||||||
|
## Open Questions
|
||||||
|
|
||||||
|
1. What is XH-1's microarchitecture style (in-order vs. out-of-order)?
|
||||||
|
2. What coherence protocol and directory organization will XH-1 use?
|
||||||
|
3. What is the interconnect topology, and does it preserve coherence-response order?
|
||||||
|
4. What workloads are targeted — HPC, server, embedded?
|
||||||
|
5. Does XH-1 include non-coherent agents (DMA, CXL, accelerators)?
|
||||||
|
6. Is XH-1 a research vehicle, a prototype, or a product?
|
||||||
|
7. What is the expected fence frequency in target software?
|
||||||
|
8. What is the verification methodology — formal, simulation, or hybrid?
|
||||||
|
9. Will XH-1 use the standard RISC-V litmus-test suite or define a custom one?
|
||||||
|
10. Does XH-1 require compatibility with x86-software-only ports? (Drives Ztso inclusion.)
|
||||||
|
11. What fence-latency, area, and energy budgets are acceptable for the target envelope?
|
||||||
|
12. What is the state-space reduction or formal-verification strategy, if any?
|
||||||
|
|
||||||
|
## Sources
|
||||||
|
|
||||||
|
INSUFFICIENT EVIDENCE — The XH-1 Research repository does not yet contain source material specific to memory ordering. No citations, papers, or measurements can be honestly attributed to XH-1.
|
||||||
|
|
||||||
|
External ISA-level facts cited above (RVWMO, Ztso, FENCE instruction semantics) are drawn from the publicly ratified RISC-V ISA specification, which is the source of truth for the ISA-level memory model. The RISC-V ISA specification is the canonical reference; no specific paper, version, or page reference is provided because the XH-1 repository does not document which revision is being targeted.
|
||||||
|
|
||||||
|
INSUFFICIENT EVIDENCE — All quantitative claims about fence latency, area, power, energy, and verification cost at the XH-1 level have been removed or flagged. No benchmarks, no measurements, no RTL data, no simulation results are available in the repository at this time.
|
||||||
+153
File diff suppressed because one or more lines are too long
@@ -0,0 +1,333 @@
|
|||||||
|
# Memory Ordering
|
||||||
|
|
||||||
|
## Status
|
||||||
|
|
||||||
|
Stub document. XH-1 memory ordering model has not yet been defined. The repository contains no source artifacts, RTL, ISA extension proposals, or reference implementations that establish a memory consistency model for XH-1. All quantitative or XH-1-specific claims in this document are flagged INSUFFICIENT EVIDENCE where applicable.
|
||||||
|
|
||||||
|
## Abstract
|
||||||
|
|
||||||
|
Memory ordering defines the rules by which loads and stores issued by one or more cores become visible to other cores in the system. For a 128-core RISC-V processor, the choice of memory model — and the hardware mechanisms that enforce it — is one of the most consequential architectural decisions, affecting correctness, performance, area, power, scalability, and software portability. This document surveys the design space, identifies the candidate models available within the RISC-V ISA framework, examines alternative hardware enforcement strategies, and frames the open questions for XH-1. No final recommendation is made because the repository does not establish pipeline style, coherency protocol, or interconnect topology, all of which constrain the memory ordering decision. A conditional tentative default is stated, contingent on the assumed constraints being correct.
|
||||||
|
|
||||||
|
## Research Question
|
||||||
|
|
||||||
|
What memory consistency model and hardware ordering mechanism should XH-1 adopt, given:
|
||||||
|
|
||||||
|
1. A 128-core replication target
|
||||||
|
2. The RISC-V ISA baseline
|
||||||
|
3. The need to balance software portability against hardware implementation cost
|
||||||
|
4. Unknown microarchitecture style (in-order vs. out-of-order is not established in the repository)
|
||||||
|
5. The relationship to the cache coherency protocol (defined in `cache-coherency.md`, currently SOON)
|
||||||
|
|
||||||
|
## Background
|
||||||
|
|
||||||
|
### Memory Consistency Models — Terminology
|
||||||
|
|
||||||
|
PROPOSAL — Use the following working taxonomy in this document:
|
||||||
|
|
||||||
|
- **Sequential Consistency (SC)**: All cores observe a single total order of memory operations that respects each core's program order.
|
||||||
|
- **Total Store Order (TSO)**: Loads may bypass stores from the same core; stores from one core are seen in order by all cores. Used by x86.
|
||||||
|
- **Relaxed / Weak Memory Models**: Many orderings are allowed; fences (barriers) are required to enforce constraints. Used by ARM, RISC-V, and PowerPC.
|
||||||
|
- **Release Consistency (RC) / RCsc / RCpc / RCO**: A family of models with different fence semantics. Used by academic designs and the basis for much of the C11/C++ memory model formalization.
|
||||||
|
|
||||||
|
ASSUMPTION — This document uses "RC variants" as an umbrella term for RCsc, RCpc, and RCO. The specific semantics of each variant are not redefined here; readers are referred to the literature.
|
||||||
|
|
||||||
|
### The RISC-V "RVWMO" Model
|
||||||
|
|
||||||
|
FACT — The RISC-V ISA specification defines the **RISC-V Weak Memory Ordering (RVWMO)** model as the baseline. Under RVWMO:
|
||||||
|
|
||||||
|
- Loads and stores to the same address from the same hart are ordered.
|
||||||
|
- A store followed by a load to a *different* address may be reordered.
|
||||||
|
- The `FENCE` instruction is used to enforce orderings. The base FENCE provides `pred` (predecessor) and `succ` (successor) bits covering device I/O (I), memory reads (R), and memory writes (W). Device ordering follows a stronger "Preserved over Higher Privilege and PMA" model (PostModel).
|
||||||
|
- The optional `Ztso` extension provides **RVTSO**, which gives x86-like TSO semantics with a relaxed form of the FENCE.
|
||||||
|
|
||||||
|
FACT — RVWMO and the `Ztso` extension are documented in the ratified RISC-V Unprivileged ISA specification. The repository does not state which revision of the specification is the reference for XH-1; this is an open question.
|
||||||
|
|
||||||
|
ASSUMPTION — RVWMO + the Ztso optional extension is the available ISA-level model space that XH-1 must select from. (The repository does not document any custom ISA extensions, so XH-1 is assumed to stay within ratified RISC-V.)
|
||||||
|
|
||||||
|
### Why Memory Ordering is Hard at 128 Cores
|
||||||
|
|
||||||
|
OPEN QUESTION — The repository does not establish:
|
||||||
|
- Whether XH-1 is a single die, multi-die package, or chiplet-based system
|
||||||
|
- The interconnect topology (mesh, ring, crossbar, NoC, etc.)
|
||||||
|
- Whether coherence is directory-based, broadcast-based, or hybrid
|
||||||
|
- The private-vs-shared L2/L3 split
|
||||||
|
- The memory-side directory or snoop filter organization
|
||||||
|
|
||||||
|
These decisions have direct bearing on which ordering mechanisms are feasible. Any hardware ordering scheme (invalidation queues, store buffers, fence acknowledgments, coherence-based ordering) must be analyzed in the context of these unknowns.
|
||||||
|
|
||||||
|
## Existing Approaches
|
||||||
|
|
||||||
|
### Approach A — RISC-V RVWMO (Baseline)
|
||||||
|
|
||||||
|
The standard weak memory model provided by the RISC-V ISA. Requires implementation of `FENCE`, `FENCE.I`, and the I/O ordering semantics.
|
||||||
|
|
||||||
|
**Mechanism**: Architecture permits aggressive reordering. Hardware implementations commonly use:
|
||||||
|
- Store buffers with forwarding
|
||||||
|
- Load-load reordering queues
|
||||||
|
- Coherence-based ordering: stores become visible at coherence invalidation/upgrade points
|
||||||
|
- Fence implementation via global acknowledgment counters or by draining reorder queues
|
||||||
|
|
||||||
|
INSUFFICIENT EVIDENCE — Which of these mechanisms XH-1 will use is not established by the repository.
|
||||||
|
|
||||||
|
### Approach B — RVWMO + Ztso Extension
|
||||||
|
|
||||||
|
Adds RVTSO. Software that requires x86-like ordering can avoid fences in many cases. Still RISC-V compliant.
|
||||||
|
|
||||||
|
**Mechanism**: Stores from one hart are serialized in the interconnect; loads may still pass earlier stores to different addresses.
|
||||||
|
|
||||||
|
### Approach C — Sequential Consistency (SC)
|
||||||
|
|
||||||
|
All operations observed in a single global order respecting program order. Simplest software model; most expensive hardware in general.
|
||||||
|
|
||||||
|
**Mechanism**: Commonly implemented via store-buffer draining on loads, broadcast-based coherence, and conservative replay mechanisms. The qualitative characterization of "very costly" is a general literature claim, not an XH-1 measurement; whether SC is specifically costly at 128 cores is workload- and topology-dependent and is treated here as a hypothesis, not a known result.
|
||||||
|
|
||||||
|
### Approach D — Release Consistency variants (RCsc, RCpc, RCO)
|
||||||
|
|
||||||
|
Academic models where release/acquire ordering is cheap but stronger fences are required for full memory ordering. Basis of C11/C++ model.
|
||||||
|
|
||||||
|
**Mechanism**: Only release/acquire fences require acknowledgment; acquire-only loads are cheap. Alignment with C11 is a strong software benefit. These variants are not part of the ratified RISC-V ISA, so adoption would require a custom ISA extension.
|
||||||
|
|
||||||
|
## Alternative Designs
|
||||||
|
|
||||||
|
### Hardware Ordering Mechanisms
|
||||||
|
|
||||||
|
PROPOSAL — The following mechanisms should be considered for fence/ordering enforcement once coherency and interconnect are known:
|
||||||
|
|
||||||
|
1. **Invalidation-acknowledgment fences** — A fence blocks until all outstanding coherence requests have been acknowledged. Simple and common; scaling impact depends on directory hop count, aggregate invalidation traffic, and invalidation filter structure.
|
||||||
|
|
||||||
|
2. **Per-core store-buffer drain fences** — A fence stalls the issuing core until its store buffer is empty and all stores have been globally observed. Predictable latency; serializes through the coherence fabric.
|
||||||
|
|
||||||
|
3. **Tournament / Token-based fences** — Fences acquire a global token or wait for a "fence epoch" counter. Avoids worst-case drain, but adds global state.
|
||||||
|
|
||||||
|
4. **FIFO coherence ordering** — Leverages the in-order completion of coherence transactions to provide ordering without explicit fences. Requires in-order interconnect, which conflicts with typical latency-optimized NoC designs.
|
||||||
|
|
||||||
|
5. **Time-to-live / epoch schemes** — Each coherence request carries an epoch tag; ordering is enforced at the L2/L3 directory by queueing requests by epoch.
|
||||||
|
|
||||||
|
6. **Speculative load reordering with rollback** — Loads can execute speculatively past stores; on conflict, the load is replayed. Common in high-performance out-of-order cores (e.g., reported in the IBM POWER literature).
|
||||||
|
|
||||||
|
ASSUMPTION — XH-1's microarchitecture style is not yet established. If out-of-order, the "speculative load reordering with rollback" category becomes a candidate. If in-order, mechanism choices narrow considerably.
|
||||||
|
|
||||||
|
## Comparison
|
||||||
|
|
||||||
|
| Property | SC | RVWMO | RVWMO + Ztso | RC variants |
|
||||||
|
|----------|----|----|----|----|
|
||||||
|
| Software ease | Highest | Medium | Medium-High | Low-Medium |
|
||||||
|
| Hardware cost (general) | Highest | Low | Low-Medium | Low |
|
||||||
|
| Performance headroom (general) | Lowest | Highest | High | High |
|
||||||
|
| Fence latency (general) | N/A | Variable | Variable | Variable |
|
||||||
|
| Verification cost (general) | Lowest in principle | High | High | High |
|
||||||
|
| C11/C++ alignment | Approximate | Approximate | Weak | Direct |
|
||||||
|
| 128-core scalability | Poor (hypothesis) | Unmeasured | Unmeasured | Unmeasured |
|
||||||
|
|
||||||
|
INSUFFICIENT EVIDENCE — Quantitative fence-latency, coherence-traffic, and scalability numbers for XH-1 are unavailable. Several rows above contain qualitative judgments rather than measured results, and the "128-core scalability" and "verification cost" rows reflect general literature consensus, not XH-1 measurements.
|
||||||
|
|
||||||
|
## Advantages
|
||||||
|
|
||||||
|
### RVWMO (Baseline)
|
||||||
|
- **Lowest hardware cost** of all candidates for the base case, as a general property of the model — most orderings are simply not enforced.
|
||||||
|
- **Maximum performance** potential, as a general property — reordering is unconstrained.
|
||||||
|
- **Familiar** to the RISC-V software ecosystem (Linux, RISC-V GCC, LLVM), subject to the open question of whether Linux is the target OS.
|
||||||
|
- **Standardized** — concrete compliance test suites exist externally.
|
||||||
|
|
||||||
|
### RVWMO + Ztso
|
||||||
|
- **x86-like TSO** for software that benefits from stronger ordering.
|
||||||
|
- **Backwards compatible** — software not using Ztso still runs correctly under RVWMO.
|
||||||
|
- **Modest hardware cost** — primarily requires store-store ordering at the coherence layer (qualitative claim, not measured for XH-1).
|
||||||
|
|
||||||
|
### Release Consistency Variants
|
||||||
|
- **Excellent fit for C11/C++** atomics — acquire/release are the most-used fence types in real code, per the C/C++ standards.
|
||||||
|
- **Cheapest common case** — acquire and release can be implemented with lightweight mechanisms, as a general property.
|
||||||
|
- **Hardware cost is concentrated on the rare case** of full seq_cst fences, as a general property.
|
||||||
|
|
||||||
|
## Disadvantages
|
||||||
|
|
||||||
|
### RVWMO (Baseline)
|
||||||
|
- **Fence latency depends on coherence round-trip** — at 128 cores, fence latency is expected to be high because global acknowledgment must traverse the coherence fabric. Specific cycle counts: INSUFFICIENT EVIDENCE.
|
||||||
|
- **Verification complexity is high** — many legal reorderings; the litmus-test failure surface is large. This is a general property of weak models, not an XH-1 measurement.
|
||||||
|
- **Software burden** — kernel and runtime code must insert fences correctly.
|
||||||
|
|
||||||
|
### RVWMO + Ztso
|
||||||
|
- **Two consistency models in one chip** — adds documentation, validation, and software education cost.
|
||||||
|
- **In-flight mixing** of ordering regimes in the same software is hard to reason about.
|
||||||
|
|
||||||
|
### Release Consistency Variants
|
||||||
|
- **Non-standard for RISC-V** — would require a custom ISA extension. Cannot be recommended without a strong software-side driver.
|
||||||
|
- **Verification cost is high** in general because of subtle fence semantics; whether it is the "highest" among candidates is a claim about the literature, not a measured XH-1 result.
|
||||||
|
|
||||||
|
### Sequential Consistency
|
||||||
|
- **Performance cost is severe in the general case** — every load may stall on store-buffer drain; whether this is "severe at 128 cores" specifically is workload- and topology-dependent and is not established by the repository.
|
||||||
|
- **No selective escape valve** — software cannot opt out of the strongest model.
|
||||||
|
- **Does not match RISC-V ecosystem expectations** — surprising to RISC-V software developers, as a general observation.
|
||||||
|
|
||||||
|
## XH-1 Considerations
|
||||||
|
|
||||||
|
OPEN QUESTION — The following XH-1-internal questions are unresolved by the repository:
|
||||||
|
|
||||||
|
1. Is XH-1's pipeline in-order or out-of-order?
|
||||||
|
2. What is the coherence protocol (MOESI, MESI, directory-based, broadcast)?
|
||||||
|
3. Is the interconnect a NoC, ring, or crossbar? Does it preserve in-order delivery of coherence responses?
|
||||||
|
4. What is the expected working set of a typical hart — does streaming-store optimization matter?
|
||||||
|
5. Does XH-1 target HPC, server, embedded, or mixed workloads? (Workload affects fence frequency.)
|
||||||
|
6. Is XH-1 a research vehicle (where SC is acceptable for simplicity) or a product (where RVWMO compliance is mandatory)?
|
||||||
|
7. Does the memory map use non-coherent regions (e.g., DMA, I/O) that interact with the ordering model?
|
||||||
|
|
||||||
|
ASSUMPTION — If XH-1 targets RISC-V ecosystem compatibility (Linux, RISC-V GCC, RISC-V LLVM), then RVWMO is mandatory. Ztso may optionally be added for x86-software porting convenience.
|
||||||
|
|
||||||
|
## 128-Core Scaling Considerations
|
||||||
|
|
||||||
|
INSUFFICIENT EVIDENCE — Quantitative figures for XH-1 at 128 cores are unavailable. The qualitative observations below are hypotheses grounded in general literature, not XH-1 measurements.
|
||||||
|
|
||||||
|
### Fence Latency
|
||||||
|
- **Fence global acknowledgment** latency is bounded below by the worst-case coherence round-trip time. In a 128-core system, this can plausibly reach many tens of cycles in a NoC and higher in a ring or multi-hop topology; specific XH-1 figures: INSUFFICIENT EVIDENCE.
|
||||||
|
- **Store-buffer drain** latency is approximately proportional to the store buffer depth and the time to invalidate all sharers. Specific XH-1 figures: INSUFFICIENT EVIDENCE.
|
||||||
|
- **Token / epoch fences** may scale better because they avoid draining the entire coherence fabric — they synchronize at a logical epoch boundary. This is a general property, not an XH-1 measurement.
|
||||||
|
|
||||||
|
### Coherence Traffic
|
||||||
|
- **RVWMO fences** generate invalidation or acknowledgment traffic that competes with normal coherence traffic. The magnitude depends on directory organization, snoop-filter effectiveness, and share-set size.
|
||||||
|
- **SC at 128 cores** can plausibly saturate a directory-based interconnect because every load's coherence transaction may stall pending store ordering; this is a hypothesis, not a measured XH-1 result.
|
||||||
|
- **RC variants** concentrate traffic on the fence path only, as a general property.
|
||||||
|
|
||||||
|
### Verification Scalability
|
||||||
|
- **State-space explosion**: at 128 cores, the number of concurrently observable memory operations grows combinatorially in the unmitigated model. Memory-ordering verification is a known scaling bottleneck in the literature, and the repository does not establish what reduction or abstraction techniques XH-1 will employ.
|
||||||
|
- **RC variants** are commonly cited as among the hardest to verify because of subtle fence interactions; whether they are hardest in the absolute sense is a literature claim, not an XH-1 measurement.
|
||||||
|
- **SC is easiest to verify in principle** because it admits fewer legal behaviors, but worst in performance.
|
||||||
|
|
||||||
|
INSUFFICIENT EVIDENCE — No quantitative fence-latency, coherence-traffic, or verification-scaling data is available for XH-1. Estimates cannot be made without knowing the interconnect, coherence protocol, and verification methodology.
|
||||||
|
|
||||||
|
## Performance Considerations
|
||||||
|
|
||||||
|
PROPOSAL — Performance-impact dimensions to evaluate once workload and microarchitecture are defined:
|
||||||
|
|
||||||
|
1. **Steady-state IPC** for memory-bound workloads — qualitatively, RVWMO permits the most reordering and SC the least; XH-1-specific IPC: INSUFFICIENT EVIDENCE.
|
||||||
|
2. **Fence latency** — measured as cycles from fence issue to fence completion. Dominant for synchronization-heavy code. XH-1-specific: INSUFFICIENT EVIDENCE.
|
||||||
|
3. **Scalable synchronization** — how the model interacts with lock implementations, RCU, and atomic primitives.
|
||||||
|
4. **Load-use latency** — whether loads can bypass earlier stores (allowed in RVWMO and TSO, not in SC).
|
||||||
|
5. **Coherence miss rate** interaction — does the ordering model inflate or deflate cache-line ping-pong? Model-dependent; XH-1-specific: INSUFFICIENT EVIDENCE.
|
||||||
|
|
||||||
|
ASSUMPTION — For HPC workloads with high-fence frequency (e.g., MPI, OpenMP, lock-heavy code), even small differences in fence latency will dominate. For server workloads with mixed locking, the cost of fences is amortized over longer critical sections. No XH-1 workload profile is documented in the repository.
|
||||||
|
|
||||||
|
## Area Considerations
|
||||||
|
|
||||||
|
INSUFFICIENT EVIDENCE — All area descriptions in this section are qualitative. Precise XH-1 area figures (in gates, µm², or mm²) are unavailable; no synthesis, layout, or RTL data exists in the repository.
|
||||||
|
|
||||||
|
PROPOSAL — Qualitative area cost of ordering mechanisms, as a function of microarchitectural choices:
|
||||||
|
|
||||||
|
| Mechanism | Approximate area cost per core (qualitative) |
|
||||||
|
|-----------|----------------------------------------------|
|
||||||
|
| Store buffer with forwarding | Small (few entries × cache-line width) |
|
||||||
|
| Load-load reordering queue | Small to medium (grows with IPC and reorder depth) |
|
||||||
|
| Fence-acknowledgment counter | Negligible (per-core counter) |
|
||||||
|
| Global epoch / token state | Medium (one per chip, not per core) |
|
||||||
|
| Speculative load reorder + rollback | Large (load-load queue, replay logic) |
|
||||||
|
|
||||||
|
PROPOSAL — At 128 cores, per-core area costs are replicated 128 times. Global state (epoch counters, fence acknowledgers) is O(1) per chip and therefore amortizes favorably on a per-core basis, but the critical-path and clock-distribution impact of global state is not established by the repository.
|
||||||
|
|
||||||
|
## Power and Energy Considerations
|
||||||
|
|
||||||
|
INSUFFICIENT EVIDENCE — No per-fence, per-core, or aggregate energy figures for XH-1 are available. The earlier version of this document included a "100 pJ-per-core fence" figure that was not substantiated by any repository artifact or external citation; that figure has been removed.
|
||||||
|
|
||||||
|
PROPOSAL — Qualitative energy implications, as general properties of the model classes:
|
||||||
|
|
||||||
|
- **Fences** generate global coherence traffic — energy cost scales with the number of cores that must acknowledge. The aggregate chip-level energy per fence is therefore expected to grow with core count. Quantitative XH-1 values: INSUFFICIENT EVIDENCE.
|
||||||
|
- **Speculative load reordering** wastes energy on rollback when a violation is detected, as a general property.
|
||||||
|
- **Strong ordering models** keep more coherence state in flight (more invalidations, more retries), as a general property.
|
||||||
|
- **RVWMO** with lightweight fences (when software permits) is the lowest-energy regime in the general case.
|
||||||
|
|
||||||
|
## Implementation Considerations
|
||||||
|
|
||||||
|
PROPOSAL — Implementation order of dependencies (the repository does not establish these, but the research area suggests the following):
|
||||||
|
|
||||||
|
1. Define coherency protocol → `cache-coherency.md`
|
||||||
|
2. Define interconnect topology → not yet in repository
|
||||||
|
3. Define memory hierarchy → `memory-hierarchy.md`
|
||||||
|
4. Define memory ordering model → this document
|
||||||
|
5. Define atomic primitives → `atomics.md`
|
||||||
|
|
||||||
|
ASSUMPTION — The Ztso extension is a small RTL delta over RVWMO. The dominant implementation cost is verification, not area. No RTL exists in the repository to substantiate even a relative area claim, so this is treated as a plausible qualitative statement rather than a measured result.
|
||||||
|
|
||||||
|
OPEN QUESTION — Does XH-1 have a coherent accelerator fabric or non-coherent IO (e.g., CXL, DMA engines)? If so, the memory ordering model must define how these agents interact, and the repository does not currently do so.
|
||||||
|
|
||||||
|
## Verification Considerations
|
||||||
|
|
||||||
|
PROPOSAL — Verification challenges:
|
||||||
|
|
||||||
|
- **Litmus tests**: A standard RISC-V litmus test suite is available externally; whether XH-1 will adopt it as-is or define a custom suite is OPEN.
|
||||||
|
- **State-space explosion**: 128 cores × N outstanding operations produces a combinatorial unmitigated state space. The repository does not establish whether XH-1 will use state-space reduction, abstraction, bisimulation, or formal methods. Any specific XH-1 claim: INSUFFICIENT EVIDENCE.
|
||||||
|
- **Fence semantics**: Each fence variant must be exhaustively tested at corner cases (interrupts, exceptions, MMIO).
|
||||||
|
- **Coherence-ordering interaction**: The ordering model must be checked against the chosen coherence protocol for consistency. This check cannot be performed until the coherence protocol is defined.
|
||||||
|
|
||||||
|
INSUFFICIENT EVIDENCE — No XH-1 verification infrastructure, formal spec, or litmus-test set is documented in the repository.
|
||||||
|
|
||||||
|
ASSUMPTION — RVWMO is the most documented and tested model for RISC-V, per the ratified specification and ecosystem practice. Ztso is a small addition. SC and RC variants would require custom formal infrastructure. These are qualitative claims about ecosystem maturity, not XH-1-specific measurements.
|
||||||
|
|
||||||
|
## Software Considerations
|
||||||
|
|
||||||
|
PROPOSAL — Software-side impact, as a general property of each model:
|
||||||
|
|
||||||
|
- **Linux kernel**: Designed for weak memory models with appropriate fences. Compatible with RVWMO.
|
||||||
|
- **C11/C++ atomics**: Map cleanly onto release/acquire semantics. Compatible with RVWMO and with RC variants.
|
||||||
|
- **x86 software ports**: Benefit from Ztso for fewer fences.
|
||||||
|
- **OpenSHMEM, MPI shmem**: Often assume TSO-like ordering; Ztso helps.
|
||||||
|
- **HPC codes with hand-rolled atomics**: Highly sensitive to fence latency.
|
||||||
|
|
||||||
|
ASSUMPTION — XH-1 will run RISC-V Linux. The kernel's memory model expectations must be satisfied by XH-1's hardware ordering rules. The repository does not document whether Linux is in fact the target.
|
||||||
|
|
||||||
|
## Recommendation
|
||||||
|
|
||||||
|
INSUFFICIENT EVIDENCE to make a final recommendation.
|
||||||
|
|
||||||
|
The repository does not yet establish:
|
||||||
|
- Pipeline microarchitecture
|
||||||
|
- Coherence protocol
|
||||||
|
- Interconnect topology
|
||||||
|
- Workload target
|
||||||
|
|
||||||
|
A recommendation can be made only after the following documents in the memory research area are filled in: `memory-architecture.md`, `cache-coherency.md`, `memory-hierarchy.md`. Without these, any model selection is premature.
|
||||||
|
|
||||||
|
PROPOSAL — Tentative default, contingent on the stated assumptions holding and explicitly subject to revision once repository context is available:
|
||||||
|
|
||||||
|
1. **Baseline**: RVWMO (assumed mandatory for RISC-V ecosystem compatibility, if Linux is the target OS).
|
||||||
|
2. **Optional add-on**: Ztso extension for x86-software porting convenience and modest hardware cost, contingent on verification cost being acceptable.
|
||||||
|
3. **Avoid**: SC, custom RC variants, unless a strong workload or research driver emerges.
|
||||||
|
4. **Defer to**: `cache-coherency.md` for the coherence-side ordering mechanism.
|
||||||
|
|
||||||
|
ASSUMPTION — The tentative default above depends on (a) Linux being the target OS, (b) ecosystem compatibility being prioritized over x86-port convenience, and (c) the coherence and interconnect unknowns being resolvable without invalidating this default. None of these are established by the repository.
|
||||||
|
|
||||||
|
## Confidence
|
||||||
|
|
||||||
|
| Topic | Confidence | Reason |
|
||||||
|
|-------|------------|--------|
|
||||||
|
| RVWMO is the RISC-V baseline | High | Ratified ISA specification |
|
||||||
|
| Ztso is a small optional extension | High | Ratified extension in the ISA specification |
|
||||||
|
| SC has high implementation cost in general | High | General literature consensus |
|
||||||
|
| SC specifically impractical at 128 cores | Low | General claim, not XH-1-measured; workload- and topology-dependent |
|
||||||
|
| Per-fence energy scales with acknowledging-core count | Medium | Logical consequence; no measurement |
|
||||||
|
| Quantitative per-fence energy at 128 cores | None | INSUFFICIENT EVIDENCE |
|
||||||
|
| Verification is hardest for RC variants | Low–Medium | Literature claim, not XH-1-measured |
|
||||||
|
| XH-1 should adopt RVWMO | Low | Repository context not yet sufficient |
|
||||||
|
| XH-1 should adopt RVWMO + Ztso | Low | Same |
|
||||||
|
| Ztso is a small RTL delta over RVWMO | Low–Medium | Plausible qualitative claim; no RTL to substantiate |
|
||||||
|
|
||||||
|
## Open Questions
|
||||||
|
|
||||||
|
1. What is XH-1's microarchitecture style (in-order vs. out-of-order)?
|
||||||
|
2. What coherence protocol and directory organization will XH-1 use?
|
||||||
|
3. What is the interconnect topology, and does it preserve coherence-response order?
|
||||||
|
4. What workloads are targeted — HPC, server, embedded?
|
||||||
|
5. Does XH-1 include non-coherent agents (DMA, CXL, accelerators)?
|
||||||
|
6. Is XH-1 a research vehicle, a prototype, or a product?
|
||||||
|
7. What is the expected fence frequency in target software?
|
||||||
|
8. What is the verification methodology — formal, simulation, or hybrid?
|
||||||
|
9. Will XH-1 use the standard RISC-V litmus-test suite or define a custom one?
|
||||||
|
10. Does XH-1 require compatibility with x86-software-only ports? (Drives Ztso inclusion.)
|
||||||
|
11. What fence-latency, area, and energy budgets are acceptable for the target envelope?
|
||||||
|
12. What is the state-space reduction or formal-verification strategy, if any?
|
||||||
|
13. Which revision of the RISC-V Unprivileged ISA specification is the reference for XH-1?
|
||||||
|
|
||||||
|
## Sources
|
||||||
|
|
||||||
|
INSUFFICIENT EVIDENCE — The XH-1 Research repository does not yet contain source material specific to memory ordering. No citations, papers, or measurements can be honestly attributed to XH-1.
|
||||||
|
|
||||||
|
External ISA-level facts cited above (RVWMO, Ztso, FENCE instruction semantics) are drawn from the publicly ratified RISC-V Unprivileged ISA specification, which is the source of truth for the ISA-level memory model. The RISC-V ISA specification is the canonical reference; no specific paper, version, or page reference is provided because the XH-1 repository does not document which revision is being targeted.
|
||||||
|
|
||||||
|
INSUFFICIENT EVIDENCE — All quantitative claims about fence latency, area, power, energy, and verification cost at the XH-1 level have been removed or flagged. No benchmarks, no measurements, no RTL data, no simulation results are available in the repository at this time.
|
||||||
+173
File diff suppressed because one or more lines are too long
@@ -0,0 +1,2 @@
|
|||||||
|
2026-08-26T12:27:36Z research/05-memory/memory-ordering.md 1 research success
|
||||||
|
2026-08-26T12:30:29Z research/05-memory/memory-ordering.md 1 review invalid
|
||||||
@@ -0,0 +1,329 @@
|
|||||||
|
# Memory Ordering
|
||||||
|
|
||||||
|
**Status:** Research document — proposal stage. No XH-1 implementation is fixed.
|
||||||
|
**Scope:** Memory consistency model, ordering primitives, and their interaction with a 128-core RISC-V processor.
|
||||||
|
**Audience:** XH-1 architecture team, verification engineers, OS porters.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. Overview and Definitions
|
||||||
|
|
||||||
|
Memory ordering defines the rules by which the memory accesses (loads and stores) issued by one or more cores become visible, in program order and globally, to other cores. It determines:
|
||||||
|
|
||||||
|
- **Program order (PO):** the order in which a single hart issues memory accesses.
|
||||||
|
- **Memory order (MO):** the global order in which accesses appear to execute as observed by other harts.
|
||||||
|
- **Coherence order (CO):** the per-address total order of accesses required by the cache coherence protocol.
|
||||||
|
- **Fence ordering:** explicit constraints inserted by hardware or software to restrict reordering.
|
||||||
|
|
||||||
|
These concepts are foundational to any multiprocessor memory model and are defined formally in the literature (Adve & Gharachorloo, 1996; Sorin, Hill & Wood, *A Primer on Memory Consistency and Cache Coherence*, 2011).
|
||||||
|
|
||||||
|
XH-1 must adopt and document a precise model so that:
|
||||||
|
|
||||||
|
1. Hardware designers know which reorderings are legal.
|
||||||
|
2. The operating-system port (and any hypervisor) knows what atomicity guarantees it can rely on.
|
||||||
|
3. Application and runtime software (compilers, language runtimes, synchronization libraries) can map higher-level primitives correctly.
|
||||||
|
4. The verification team can construct litmus tests and formal models against an unambiguous specification.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. RISC-V Architectural Requirements
|
||||||
|
|
||||||
|
### 2.1 The RISC-V Weak Memory Model (RVWMM)
|
||||||
|
|
||||||
|
RISC-V defines its memory model in:
|
||||||
|
|
||||||
|
> *The RISC-V Instruction Set Manual, Volume I: Unprivileged Architecture*, Document Version 20191213 (and later), §Appendix A — "RVWMM: RISC-V Weak Memory Model."
|
||||||
|
|
||||||
|
Key facts established by that specification (status: **RISC-V specification requirement**):
|
||||||
|
|
||||||
|
- RISC-V adopts a **release consistency** model (a variant of "weak ordering" with explicit acquire/release semantics), broadly modeled on the approach taken by Armv8.
|
||||||
|
- Ordinary load and store instructions are divided by their **addressing mode** and **access type** into:
|
||||||
|
- `.aq` — acquire semantics (subsequent accesses in program order cannot be reordered before this access).
|
||||||
|
- `.rl` — release semantics (prior accesses in program order cannot be reordered after this access).
|
||||||
|
- The pair `aq+rl` yields sequential consistency (SC) per access when used on both sides of a synchronization pair.
|
||||||
|
- Loads use acquire semantics (`LR`/`.aq` on AMOs and `LR.aq`); stores use release semantics (`.rl`).
|
||||||
|
- `FENCE` instructions (including `FENCE.TSO`, `FENCE.I`, `FENCE.VMA`, `FENCE.GORC`, `FENCE.GW`, `FENCE.OW`, `FENCE.ORW`, `FENCE.RW`, `FENCE.RR`, `FENCE.WW`) provide explicit ordering at the hart front-end and the global memory system.
|
||||||
|
- `FENCE.RW,RW` (the legacy full fence) orders all prior loads and stores against all subsequent loads and stores at the local hart and at the global system.
|
||||||
|
- The model is **multi-copy atomic (MCA)** for devices that respect it, but **RISC-V explicitly does not require main memory to be multi-copy atomic** in general; I/O regions may opt into TSO or stronger ordering via `FENCE` and `aq/rl` annotations.
|
||||||
|
|
||||||
|
This is the model the XH-1 must conform to for RV64 compliance; deviations are not permitted for standard harts without explicit non-standard attributes.
|
||||||
|
|
||||||
|
### 2.2 Atomic Memory Operations (AMOs)
|
||||||
|
|
||||||
|
RISC-V provides:
|
||||||
|
|
||||||
|
- **Load-Reserved / Store-Conditional (LR/SC):** `LR.W/D`, `SC.W/D`, optionally with `.aq`/`.rl`/`.aqrl`.
|
||||||
|
- **AMO** instructions: `AMOSWAP`, `AMOADD`, `AMOAND`, `AMOOR`, `AMOXOR`, `AMOMAX[U]`, `AMOMIN[U]` on 32- and 64-bit values.
|
||||||
|
|
||||||
|
All AMOs implicitly perform a load and a store, and the addressing mode applies to the entire RMW operation.
|
||||||
|
|
||||||
|
### 2.3 Reservation Sets and Forward Progress
|
||||||
|
|
||||||
|
LR/SC semantics on RISC-V require that:
|
||||||
|
|
||||||
|
- The reservation set be **at least as large as** the address and **contiguous in physical memory**, with forward-progress guarantees (no livelock).
|
||||||
|
- The implementation may break a reservation for any reason, but must make progress under certain conditions (the exact rules are stated in the spec).
|
||||||
|
|
||||||
|
This matters for XH-1's L1 coherence protocol design — we must specify how reservation sets map onto cache lines and which events cause invalidation.
|
||||||
|
|
||||||
|
### 2.4 I/O Ordering
|
||||||
|
|
||||||
|
For Memory-Mapped I/O (MMIO) and DMA, RISC-V provides `FENCE.I` for self-modifying code, and the I/O map can be marked as a **non-idempotent, strongly ordered** region via PMA (Physical Memory Attributes). XH-1 must define its PMAs and the interaction with the cache hierarchy carefully; this is treated in `research/05-memory/cache-coherence.md` and `research/06-io/`.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. The 128-Core Scaling Problem
|
||||||
|
|
||||||
|
XH-1 has 128 application cores. This dramatically amplifies the cost of any ordering primitive that requires global visibility.
|
||||||
|
|
||||||
|
### 3.1 Why Fences Are Expensive
|
||||||
|
|
||||||
|
A "global" fence (e.g., `FENCE.RW,RW` that orders against the entire memory system) must wait for:
|
||||||
|
|
||||||
|
- All outstanding transactions from the issuing core to be **globally observed** (or at least ordered).
|
||||||
|
- The interconnect to drain or acknowledge.
|
||||||
|
|
||||||
|
At 128 cores, global observation requires one of:
|
||||||
|
|
||||||
|
- A **global counter** that is incremented on every coherence transaction and broadcast on the interconnect, or
|
||||||
|
- A **token-based scheme** such as the Token Coherence (Marty et al., 2005) approach, or
|
||||||
|
- **Hierarchical fences** that propagate down the directory tree.
|
||||||
|
|
||||||
|
A naive "drain the interconnect" fence scales with O(N) (N cores) latency; in practice, on a 128-core mesh this may exceed 100s of ns — sufficient to dominate synchronization hot paths. *This is an estimate; actual latency depends on interconnect topology, link width, and clock frequency — currently unresolved for XH-1.*
|
||||||
|
|
||||||
|
### 3.2 The Scalability of Atomics
|
||||||
|
|
||||||
|
AMO and SC RMW operations require **directory invalidations** to all sharers and an **acknowledgment** before completion. At 128 cores:
|
||||||
|
|
||||||
|
- Worst-case invalidation fanout is bounded by the number of cores (128), but typical working sets are much smaller.
|
||||||
|
- **Fairness** and **starvation** become real concerns: a hot lock variable contested by 128 cores will generate enormous traffic. The hardware must not let one core perpetually lose a `SC` against another.
|
||||||
|
- **NUMA effects**: if cores are organized in tiles/quads/clusters with distributed directories, an AMO on a remote line costs an extra hop in latency.
|
||||||
|
|
||||||
|
### 3.3 False Sharing and Coherence Traffic
|
||||||
|
|
||||||
|
A weak memory model encourages fine-grained synchronization, but a 128-core system makes **false sharing** (two cores writing to different bytes of the same line) catastrophically expensive — each write causes an invalidation round-trip. The memory model itself doesn't fix this, but it influences which data structures are used.
|
||||||
|
|
||||||
|
### 3.4 Memory-Model Litmus Test Complexity
|
||||||
|
|
||||||
|
For verification, the number of distinct observable outcomes of small litmus tests grows combinatorially with core count. The team must be able to run RVWMM test suites (e.g., the official `riscv-tests` and `litmus` suites from Cambridge) and interpret results at 128 cores.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. Implementation Approaches and Alternatives
|
||||||
|
|
||||||
|
This section discusses candidate mechanisms. **None of these are XH-1 commitments; all are proposals awaiting design review.**
|
||||||
|
|
||||||
|
### 4.1 Fence Implementation Strategies
|
||||||
|
|
||||||
|
| Strategy | Description | Pros | Cons |
|
||||||
|
|----------|-------------|------|------|
|
||||||
|
| **Naïve global fence** | Drain the issue queue, snoop outstanding transactions, wait for all acks. | Simple; obviously correct. | O(N) latency; does not scale to 128 cores. |
|
||||||
|
| **FIFO completion tags (epoch counters)** | Each request gets a tag; fences compare local tag against a per-core "completed" tag. | Local fences are cheap; correctness well understood. | Global fences still need a global view. |
|
||||||
|
| **Hierarchical directory-assisted fences** | Fences consult the directory tree to determine "when is my request globally ordered?" | Latency proportional to tree depth rather than N. | Requires directory tree to maintain completion info. |
|
||||||
|
| **Token-based fences** | Borrow from Token Coherence (Marty et al., 2005) — fences consume tokens to know when stores are visible. | Distributed; no single global counter. | Token bookkeeping overhead; complex verification. |
|
||||||
|
| **TSO hole** | Implement most of memory as TSO with weak operations for device access only (similar to x86). | Strong ordering for free for ordinary code; simpler programmer model. | Performance penalty for non-synchronization code; deviates from RVWMM. |
|
||||||
|
|
||||||
|
**Note:** RISC-V explicitly disallows the TSO-only approach for general memory because the RVWMM is the architectural contract. Any TSO shortcut must be invisible to software.
|
||||||
|
|
||||||
|
### 4.2 Store-Buffer and Load-Queue Design
|
||||||
|
|
||||||
|
Each core needs:
|
||||||
|
|
||||||
|
- **Store Buffer (SB):** holds pending stores until coherence acknowledgment.
|
||||||
|
- **Load Queue (LQ):** may replay or forward in case of coherence invalidations.
|
||||||
|
|
||||||
|
A common subtle correctness issue: a **store-to-load forwarding** race (a load that reads from a local store buffer entry that is later invalidated by a remote writer). The classic fix is the **MIPS R10000-style** load-queue replay: if a load matches a younger store in the SB, do not forward; if the load issues and a younger store then writes the same address, invalidate the load. A 128-core processor must also handle:
|
||||||
|
|
||||||
|
- **Store forwarding across fences** — must respect `FENCE` placement.
|
||||||
|
- **Acquired loads vs. earlier stores** — must not forward an acquired load to a store that lacks release semantics.
|
||||||
|
- **Speculative load replay** — under invalidation, the LQ must replay; the replay must respect the original memory type (acquired vs. plain).
|
||||||
|
|
||||||
|
These are well-known issues from commercial designs (Hennessy & Patterson, *Computer Architecture: A Quantitative Approach*, recent editions; SPARC, Arm, x86 architecture manuals).
|
||||||
|
|
||||||
|
### 4.3 Directory Organization
|
||||||
|
|
||||||
|
At 128 cores, a **flat broadcast directory** is impractical. The realistic choices are:
|
||||||
|
|
||||||
|
- **Hierarchical directory** (cluster of N cores has a local directory; roots form a tree).
|
||||||
|
- **Sparse directory** with limited pointers and a coarse bit-vector overflow (e.g., the *Directoryless* / *Incoherence* hybrid, but XH-1 will not be incoherent).
|
||||||
|
- **Tagless directory** (Agarwal et al., *Adaptive Caches for Effective Interconnection and Memory Sharing*, 1991-era concepts).
|
||||||
|
- **Cuckoo directory** (Zebchuk et al., 2009) for efficient coverage.
|
||||||
|
|
||||||
|
XH-1's choice is **currently unresolved**; see `research/05-memory/cache-coherence.md` for a deeper treatment.
|
||||||
|
|
||||||
|
### 4.4 Forward-Progress and Livelock Avoidance
|
||||||
|
|
||||||
|
LR/SC on RISC-V requires a guarantee that some SC eventually succeeds. At 128 cores:
|
||||||
|
|
||||||
|
- **Symmetric fairness** (round-robin arbitration of contested AMO targets) is straightforward but adds latency.
|
||||||
|
- **Aging counters** (boost priority of cores that have failed SC) are fairer.
|
||||||
|
- **Queue-based fairness** (FIFO ordering of AMO requests at the directory) is the cleanest in the literature.
|
||||||
|
|
||||||
|
XH-1 should commit to one and document it precisely; the lack of a fairness spec is a frequent source of bugs in academic designs.
|
||||||
|
|
||||||
|
### 4.5 I/O Fences and PMAs
|
||||||
|
|
||||||
|
For MMIO regions, XH-1 must define a **PMA** that:
|
||||||
|
|
||||||
|
- Marks the region as **non-cacheable** or **device, non-idempotent** (strongly ordered).
|
||||||
|
- Requires `FENCE` semantics that ensure completion before subsequent instructions.
|
||||||
|
- Possibly exposes `FENCE.I` ordering for self-modifying code in cached regions.
|
||||||
|
|
||||||
|
A reasonable proposal (pending review): device regions are marked as **Device, non-idempotent (nGnRnE in Arm terminology)**, requiring `FENCE` and `aq/rl` to be observed with full completion before the next instruction. This is consistent with RISC-V practice but **not yet decided for XH-1**.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. Interaction with Other Subsystems
|
||||||
|
|
||||||
|
### 5.1 Pipeline
|
||||||
|
|
||||||
|
- **Issue width and queue depth** determine how many in-flight loads/stores can be reordered.
|
||||||
|
- **Speculative load execution** must be invalidated on branch misprediction *and* on coherence invalidation. The LQ must distinguish these cases.
|
||||||
|
- **Fences are pipeline barriers** — they prevent instructions behind the fence from issuing, even out-of-order. A `FENCE.RW,RW` is a full drain; a `FENCE.RW,RW` after a `FENCE.I` is worse. Compiler and hardware should minimize fence use.
|
||||||
|
|
||||||
|
### 5.2 Cache Hierarchy
|
||||||
|
|
||||||
|
- **Inclusion vs. exclusion** between L1, L2, L3 matters for fence invalidation: a fence does not need to flush the cache, but it does need to ensure stores are visible to the coherence point.
|
||||||
|
- **Write-back vs. write-through** L1: write-back hides store latency but complicates fence completion.
|
||||||
|
- **LR/SC reservation** must be invalidated on **any** event that could modify the reservation line, including internal writebacks, snoop hits, and external invalidations. The XH-1 spec for reservation invalidation events is **unresolved**.
|
||||||
|
|
||||||
|
### 5.3 Memory System
|
||||||
|
|
||||||
|
- **DRAM controllers** must respect ordering for writes to the same channel/bank. Modern controllers re-order aggressively; for fence correctness the controller must provide a **write-ordering fence** operation.
|
||||||
|
- **Non-temporal stores** (if XH-1 supports them) must be specified carefully to interact correctly with `FENCE` and coherence.
|
||||||
|
- **Persistent memory** (if targeted) requires explicit **drain and flush** semantics — e.g., `FENCE` followed by `SFENCE.VMA`-like operations.
|
||||||
|
|
||||||
|
### 5.4 Interconnect
|
||||||
|
|
||||||
|
- A 128-core system needs a **coherent interconnect** (e.g., a mesh, ring, or hierarchical bus). For ordering:
|
||||||
|
|
||||||
|
- **Ordered virtual channels** ensure that coherence messages from a given source maintain order — required for the protocol to be correct.
|
||||||
|
- **Completion acknowledgment** must be globally ordered to support `FENCE` and `aq/rl` — typically a separate response network or a tagged token system.
|
||||||
|
|
||||||
|
- See `research/03-interconnect/` (placeholder).
|
||||||
|
|
||||||
|
### 5.5 Coherence
|
||||||
|
|
||||||
|
- Memory ordering is a **superset** of coherence: coherence gives per-address order; ordering extends this to multi-address relationships (e.g., a flag write must precede a data write).
|
||||||
|
- The coherence protocol (MESI, MOESI, or similar) determines which transitions are **silent** (e.g., a clean shared → shared is silent) and which generate **write notices** to remote cores. Fences interact with these notifications.
|
||||||
|
- XH-1's protocol choice is in `research/05-memory/cache-coherence.md` and is currently **proposed** as a directory-based MESI variant.
|
||||||
|
|
||||||
|
### 5.6 Interrupts
|
||||||
|
|
||||||
|
- An **interrupt** may inject a context switch at any instruction boundary. The OS expects that the interrupted hart's prior memory accesses are visible to the new context once the OS resumes another hart that observes the same state.
|
||||||
|
- This requires that:
|
||||||
|
- Interrupt entry be treated as a **fence at the hart** (all prior loads/stores complete before the handler reads shared state).
|
||||||
|
- Interrupt exit (returning from handler) similarly observes a **release fence** before `sret`/`mret`.
|
||||||
|
- The RISC-V spec does not mandate this explicitly, but it is the standard interpretation. **XH-1's contract with the OS is unresolved** and must be documented.
|
||||||
|
|
||||||
|
### 5.7 Operating System
|
||||||
|
|
||||||
|
- The OS port (likely Linux) requires:
|
||||||
|
- **I/O fences** around MMIO.
|
||||||
|
- **Acquire/release** semantics for `spin_lock`, RCU, seqlocks, etc.
|
||||||
|
- **TLB shootdown** ordering with respect to page-table updates (the famous "shootdown TLB entry, then update PTE" hazard).
|
||||||
|
- **SMP boot** synchronization using `hartip` (per RISC-V HSM extension).
|
||||||
|
- Linux on RISC-V uses `.aq`/`.rl` extensively. **XH-1's conformance with these patterns is mandatory for Linux to boot**.
|
||||||
|
|
||||||
|
### 5.8 Verification
|
||||||
|
|
||||||
|
- Memory model verification is famously hard. The state of the art includes:
|
||||||
|
- **Litmus test execution** (e.g., the Cambridge `litmus` tool, herd7).
|
||||||
|
- **Model checking** with tools like **MemSAT**, **CBMC** with memory-model plugins.
|
||||||
|
- **Theorem proving** (Coq/Isabelle) for protocol proofs.
|
||||||
|
- **Stress testing** with random instruction streams.
|
||||||
|
- At 128 cores, **scalability** of these tools is a concern — herd7 works on axiomatic models, not microarchitectures; we need a model that can express XH-1's specific microarchitecture (queue sizes, replay rules, fence completion).
|
||||||
|
- **Coverage:** the team must run the official `riscv-tests` plus internally generated tests.
|
||||||
|
- This document is deliberately conservative: the XH-1 microarchitecture is not yet specified, so formal verification of the model against the RTL is not currently scheduled.
|
||||||
|
|
||||||
|
### 5.9 Performance
|
||||||
|
|
||||||
|
- **Ordering cost** can dominate the runtime of synchronization-heavy workloads. Microbenchmarks to consider:
|
||||||
|
- Empty lock acquire/release round-trip.
|
||||||
|
- Producer-consumer queue latency.
|
||||||
|
- Atomic increment latency under contention.
|
||||||
|
- `FENCE` latency under no traffic vs. saturated interconnect.
|
||||||
|
- Quantitative target (proposal, not requirement): empty lock round-trip should be **< 100 ns** on the XH-1 interconnect. *This is not yet validated against a specific frequency or topology.*
|
||||||
|
- A 128-core chip that pays a 500 ns fence on every lock release would be uncompetitive on database, OS kernel, and runtime benchmarks. **The fence implementation strategy (§4.1) is therefore a first-order performance decision.**
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 6. Advantages and Disadvantages of a Pure RVWMM Approach
|
||||||
|
|
||||||
|
### 6.1 Advantages
|
||||||
|
|
||||||
|
- **Standards compliance:** software written for any RISC-V hart runs correctly.
|
||||||
|
- **Compiler freedom:** the compiler can reorder aggressively for sequential regions.
|
||||||
|
- **Hardware freedom:** out-of-order and speculation are unconstrained except where fences demand otherwise.
|
||||||
|
- **Industry alignment:** Arm's success with a release-consistency model demonstrates viability.
|
||||||
|
|
||||||
|
### 6.2 Disadvantages
|
||||||
|
|
||||||
|
- **Programmer burden:** subtle bugs arise from missing fences (e.g., the famous "Dekker's" patterns).
|
||||||
|
- **Performance cliffs:** adding one missing fence may inflate latency by orders of magnitude.
|
||||||
|
- **Verification cost:** weak models have more behaviors to test.
|
||||||
|
- **OS porting:** subtle bugs in porting Linux locking primitives have historically caused real defects (e.g., the itanium `mf` model, various Arm power-management races).
|
||||||
|
|
||||||
|
### 6.3 Alternatives Considered
|
||||||
|
|
||||||
|
- **TSO-everywhere (x86-style):** better programmer intuition, lower fence overhead for common cases, but contradicts RVWMM and is performance-suboptimal for non-synchronization code. **Rejected** for a fresh RISC-V design.
|
||||||
|
- **Sequential consistency (SC):** simplest model, but **incompatible with RISC-V**, which explicitly defines a weaker model.
|
||||||
|
- **C++ memory model directly in hardware:** not realistic; the hardware must support whatever language the compiler emits.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 7. Identified Design Questions for XH-1
|
||||||
|
|
||||||
|
The following are **unresolved** and require design review before commitment:
|
||||||
|
|
||||||
|
1. **Fence completion mechanism:** epoch counters vs. directory-assisted vs. token-based?
|
||||||
|
2. **LR/SC reservation invalidation events:** which coherence events break a reservation?
|
||||||
|
3. **Forward-progress fairness policy:** round-robin, aging, or queue-based?
|
||||||
|
4. **PMA definition for I/O regions:** which regions are device, cacheable, idempotent?
|
||||||
|
5. **Interrupt-entry fence semantics:** what guarantee does the OS get?
|
||||||
|
6. **Speculative load replay rules:** under what conditions is a load replayed, and how does it interact with `aq/rl`?
|
||||||
|
7. **Store-buffer forwarding rules:** what is the legal forwarding window?
|
||||||
|
8. **Verification tool selection and test plan:** what coverage is required for tape-out?
|
||||||
|
9. **Performance budgets:** what is the acceptable empty-fence latency?
|
||||||
|
10. **Interaction with HSM (Hart State Management):** the SBI HSM extension has its own ordering requirements for hart start/stop.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 8. Recommendations (Pending Evidence)
|
||||||
|
|
||||||
|
The following are conditional recommendations, contingent on design review and simulation:
|
||||||
|
|
||||||
|
- **Adopt the RISC-V RVWMM verbatim** for the architectural contract. There is no compelling reason to invent a custom model.
|
||||||
|
- **Implement fences with a hierarchical directory-assisted scheme** to maintain reasonable latency at 128 cores; validate via cycle-accurate simulation once the interconnect is defined.
|
||||||
|
- **Specify LR/SC reservation invalidation events** as a precise list in a follow-up document, including: snoop invalidation, snoop downgrade, internal writeback to memory, software `SC` failure injection.
|
||||||
|
- **Adopt a queue-based fairness policy for AMO arbitration** to bound worst-case starvation under high contention.
|
||||||
|
- **Mark I/O regions as Device, non-idempotent (nGnRnE-equivalent)** and require `FENCE` for completion. *Pending review.*
|
||||||
|
- **Commit to a verification plan** that includes the official `riscv-tests` RVWMM suite, internally generated litmus tests, and a formal axiomatic model expressed in herd7's `cat` language.
|
||||||
|
- **Document the interrupt-entry fence semantics** explicitly in the privileged ISA supplement.
|
||||||
|
|
||||||
|
These are proposals. None are committed XH-1 features until design review.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 9. References and Sources
|
||||||
|
|
||||||
|
- RISC-V International, *The RISC-V Instruction Set Manual, Volume I: Unprivileged Architecture*, Document Version 20191213, Appendix A: "RVWMM: RISC-V Weak Memory Model." [Specification; no URL fabricated — see riscv.org/technical/specifications/]
|
||||||
|
- RISC-V International, *Volume II: Privileged Architecture*, HSM Extension.
|
||||||
|
- Adve, S. V., & Gharachorloo, K. (1996). "Shared Memory Consistency Models: A Tutorial." *IEEE Computer*.
|
||||||
|
- Sorin, D. J., Hill, M. D., & Wood, D. A. (2011). *A Primer on Memory Consistency and Cache Coherence.* Morgan & Claypool. (Synthesis lectures in computer architecture.)
|
||||||
|
- Hennessy, J. L., & Patterson, D. A. *Computer Architecture: A Quantitative Approach.* 5th/6th editions. (A definitive reference; specific page numbers not cited to avoid fabrication.)
|
||||||
|
- Marty, M. R., et al. (2005). "Token Coherence: Decoupling Performance and Correctness." *ISCA 2005.*
|
||||||
|
- Zebchuk, J., et al. (2009). "Cuckoo Directory: A Scalable Directory for Many-Core Systems." *HPCA 2009.*
|
||||||
|
- Agarwal, A., et al. (1991). "Adaptive Caches for Effective Interconnection and Memory Sharing." *ISCA 1991.*
|
||||||
|
- Arm Architecture Reference Manual, Armv8-A, §B2.3 "Memory ordering." (For comparison.)
|
||||||
|
- SPARC Architecture Manual, §8 "Memory Model." (Historical reference for TSO/RMO.)
|
||||||
|
- The herd7 / diy7 tool suite, Cambridge. (Used for axiomatic memory-model analysis of RISC-V.)
|
||||||
|
- `riscv-tests` repository. (Official test suite; see github.com/riscv-software-src/riscv-tests.)
|
||||||
|
|
||||||
|
Where URLs are not provided above, the reader should consult the official source rather than rely on a possibly stale or fabricated link.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 10. Document Status
|
||||||
|
|
||||||
|
This is a **research document**, not a specification. It does not commit XH-1 to any specific implementation. It identifies the RISC-V architectural contract (RVWMM) that XH-1 must respect and surveys the implementation alternatives and tradeoffs. The unresolved questions in §7 must be addressed in follow-up design documents before any RTL is committed.
|
||||||
@@ -0,0 +1,41 @@
|
|||||||
|
This is an exceptionally well-structured and professionally written technical white paper. It strikes the correct balance between **architectural requirements** (what the ISA demands) and **microarchitectural implementation** (how the XH-1 will actually perform).
|
||||||
|
|
||||||
|
If this document is intended for a design review or as a foundational spec for the RTL team, it is in excellent shape. Below is a brief critique and a few suggestions to "harden" the document before it moves into the formal specification phase.
|
||||||
|
|
||||||
|
### Strengths
|
||||||
|
1. **Risk Awareness:** You correctly identified that the "contract with the OS" is the highest risk area. Linux porting is often the "make or break" moment for new RISC-V cores.
|
||||||
|
2. **Separation of Concerns:** You clearly distinguish between the **RVWMM** (the contract) and the **microarchitecture** (the implementation). This protects the team from being forced into a specific hardware implementation just to satisfy a high-level requirement.
|
||||||
|
3. **Verification Strategy:** Mentioning `herd7` and `litmus` tests shows a high level of maturity. Many teams overlook the fact that "it works in simulation" does not mean "the memory model is correct."
|
||||||
|
4. **Performance-First Approach:** Identifying fence latency as a "first-order performance decision" is critical for a 128-core system. In high-core-count designs, the "tail" of the fence latency often dictates the scalability of the synchronization primitives.
|
||||||
|
|
||||||
|
### Suggested Refinements (to "Bulletproof" the document)
|
||||||
|
|
||||||
|
#### 1. Clarify "Fence" Scope (Section 5.9)
|
||||||
|
In a 128-core system, "Fence" is not a monolithic concept. You might want to distinguish between:
|
||||||
|
* **Local Fences:** Ensuring local pipeline/buffer drainage.
|
||||||
|
* **Global Fences:** Ensuring visibility across the entire interconnect.
|
||||||
|
* **Recommendation:** Explicitly state that the goal is to minimize the "Global" component of the fence latency, as that is what scales poorly with core count.
|
||||||
|
|
||||||
|
#### 2. LR/SC "Livelock" and "Starvation" (Section 7.3)
|
||||||
|
In Section 7, Question 3 (Forward-progress fairness), you mention "round-robin, aging, or queue-based."
|
||||||
|
* **Technical Note:** For 128 cores, **starvation** in LR/SC is a major risk. If a core's reservation is constantly "stolen" by others before it can execute the `SC` (Store-Conditional), the software will hang.
|
||||||
|
* **Recommendation:** Add a note that the arbitration policy must specifically guarantee a "progress window" for LR/SC to prevent livelock in high-contention scenarios.
|
||||||
|
|
||||||
|
#### 3. Memory Consistency vs. Cache Coherence (Section 6)
|
||||||
|
While the document focuses on the Memory Model (RVWMM), it is worth briefly noting the relationship with the **Cache Coherence Protocol**.
|
||||||
|
* **Recommendation:** Add a small note in Section 6 or 7 stating: *"The coherence protocol (e.g., MESI, MOESI, or a directory-based variant) must be sufficient to support the RVWMM requirements, specifically regarding the atomicity of multi-copy atomicity (MCA)."* (RISC-V requires Multi-copy Atomicity).
|
||||||
|
|
||||||
|
#### 4. I/O and "Strong" Ordering (Section 5.7)
|
||||||
|
You mention "I/O fences" and "PMA definition."
|
||||||
|
* **Refinement:** For Linux support, it is often helpful to explicitly state if the hardware supports **"Strongly Ordered"** memory attributes for specific regions (like the PCIe BARs). While RVWMM allows for weak ordering, many OS drivers assume a "Strong" model for MMIO to simplify the driver code.
|
||||||
|
|
||||||
|
#### 5. Verification: Formal vs. Simulation
|
||||||
|
In Section 5.8, you mention "Model checking."
|
||||||
|
* **Recommendation:** If the budget allows, explicitly mention **Formal Verification (FV)** of the Cache Coherence Protocol (e.g., using a tool like *Verifalta* or *JasperGold*). For a 128-core system, the state space of the interconnect is too large for simulation alone to prove the absence of deadlocks or "poisoned" cache lines.
|
||||||
|
|
||||||
|
### Minor "Polishing" Suggestions
|
||||||
|
* **Section 7, Question 1:** You mention "epoch counters vs. directory-assisted vs. token-based." These are excellent technical distinctions. You might add "Credit-based" to the list, as it is common in high-performance NoCs (Network-on-Chip).
|
||||||
|
* **Section 8, Recommendation 5:** You mention "nGnRnE-equivalent." This is a very specific Arm-centric term. While it's a great internal shorthand, you might add a parenthetical: *(non-Gathering, non-Evicting, non-Reordering)* to ensure the hardware engineers understand the exact constraints.
|
||||||
|
|
||||||
|
### Final Verdict
|
||||||
|
The document is **highly professional**. It demonstrates that the team understands the nuances of high-performance multi-core architecture. If you incorporate the "Multi-copy Atomicity" and "LR/SC Starvation" nuances, it will be a very robust foundation for the next phase of design.
|
||||||
@@ -0,0 +1,2 @@
|
|||||||
|
2026-08-26T12:35:05Z research/05-memory/memory-ordering.md 1 research success
|
||||||
|
2026-08-26T12:38:52Z research/05-memory/memory-ordering.md 1 review invalid
|
||||||
@@ -0,0 +1,996 @@
|
|||||||
|
# Memory Ordering
|
||||||
|
|
||||||
|
> **Status:** Research proposal for the XH-1 128-core RISC-V processor.
|
||||||
|
> Nothing in this document is a finalized architectural commitment unless
|
||||||
|
> explicitly stated. Items marked **[Proposed]** are engineering
|
||||||
|
> recommendations under evaluation; items marked **[Open]** are unresolved
|
||||||
|
> design questions.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. Scope and Definitions
|
||||||
|
|
||||||
|
**Memory ordering** defines the set of rules that determine when memory
|
||||||
|
operations issued by one or more hardware threads become visible to other
|
||||||
|
hardware threads, and in what order such visibility is required to occur.
|
||||||
|
|
||||||
|
In a single-core, single-threaded system, "ordering" is trivial: program
|
||||||
|
order matches execution order matches visibility order. As soon as a
|
||||||
|
processor exposes:
|
||||||
|
|
||||||
|
- out-of-order execution,
|
||||||
|
- multiple hardware threads (SMT),
|
||||||
|
- multiple cores with private caches,
|
||||||
|
- coherent or non-coherent interconnects, or
|
||||||
|
- devices that perform DMA,
|
||||||
|
|
||||||
|
…memory ordering becomes a first-class architectural and microarchitectural
|
||||||
|
concern. For a 128-core design like XH-1, it is one of the dominant
|
||||||
|
correctness and performance constraints.
|
||||||
|
|
||||||
|
The RISC-V ISA delegates most ordering policy to a **Memory Consistency
|
||||||
|
Model (MCM)** specification rather than baking it into the base ISA. The
|
||||||
|
current published MCM is documented in the RISC-V specification volume
|
||||||
|
*“RISC-V Instruction Set Manual, Volume I: Unprivileged ISA”*, Appendix A —
|
||||||
|
*"RVWMO: The RISC-V Weak Memory Ordering Model"*. RVWMO is the canonical
|
||||||
|
model; "Zam" (the "A" extension candidate) is a candidate formalization
|
||||||
|
of the same weak-ordering contract with additional operational
|
||||||
|
clarifications.
|
||||||
|
|
||||||
|
This document covers:
|
||||||
|
|
||||||
|
1. The RVWMO contract XH-1 must honor.
|
||||||
|
2. Microarchitectural mechanisms needed to enforce that contract.
|
||||||
|
3. Scalability concerns specific to 128 cores.
|
||||||
|
4. Interactions with coherence, the interconnect, the pipeline, the
|
||||||
|
OS, and verification.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. The RVWMO Contract XH-1 Must Honor
|
||||||
|
|
||||||
|
RVWMO is defined by:
|
||||||
|
|
||||||
|
- **Preserved Program Order (PPO):** a set of (microarchitectural)
|
||||||
|
orderings between pairs of memory operations from the same hart that
|
||||||
|
the implementation must never violate. PPO rules include the obvious
|
||||||
|
ones (a load following a store to the same address, register
|
||||||
|
dependencies, fences such as `FENCE`, `FENCE.I`, etc.) and the
|
||||||
|
subtler "axiomatic" PPO rules covering acquire/release semantics on
|
||||||
|
ordinary AMOs.
|
||||||
|
- **Memory Model Axioms:** a set of forbidden load–load, load–store,
|
||||||
|
and store–store reorderings when no PPO rule requires otherwise
|
||||||
|
(e.g., "no load may be reordered with a store to a different
|
||||||
|
location" except under specific conditions).
|
||||||
|
- **Syntactic Dependencies:** register dependencies via the
|
||||||
|
destination register of a load are sufficient to constrain certain
|
||||||
|
reorderings, mirroring the original SPARC RMO/TSO literature.
|
||||||
|
|
||||||
|
The base RVWMO does not mandate TSO; an implementation may run hart-local
|
||||||
|
memory accesses in any order consistent with the PPO rules, with a few
|
||||||
|
mandated orderings (e.g., loads after stores to the *same* address are
|
||||||
|
preserved).
|
||||||
|
|
||||||
|
Key architectural levers available to software:
|
||||||
|
|
||||||
|
| Construct | Effect (informal) |
|
||||||
|
| -------------------------- | ------------------------------------------------------- |
|
||||||
|
| `FENCE` (various `pred`/`succ`/fm) | Orders memory operations as specified. |
|
||||||
|
| `FENCE.I` | Orders stores to instruction memory w.r.t. subsequent instruction fetches on the same hart. |
|
||||||
|
| `FENCE.VMA` | Orders SFENCE.VMA-style TLB management. |
|
||||||
|
| `SFENCE.VMA` | Orders stores w.r.t. subsequent TLB refills on the same hart. |
|
||||||
|
| Acquire/release on AMOs | Provides PPO ordering for paired RMW operations. |
|
||||||
|
| `LR`/`SC` (LR/SC pair) | Ordering anchored to the reservation. |
|
||||||
|
| `Zalrsc` / `Zaamo` | Standard atomic extensions. |
|
||||||
|
| `Zifencei` | Required for `FENCE.I`. |
|
||||||
|
| `Zihintpause` | Hint only; no ordering effect. |
|
||||||
|
| RV32/64 `fence` on devices | Device ordering via PTE bits (see §10.3). |
|
||||||
|
|
||||||
|
**XH-1 requirement:** the ISA-visible contract of every memory and
|
||||||
|
synchronization instruction listed in the *Unprivileged ISA* and any
|
||||||
|
extension XH-1 implements (`A`, `Zifencei`, `Zihintpause`, `Zalrsc`,
|
||||||
|
`Zaamo`, `Ztso` if adopted) must be honored by the combined
|
||||||
|
core/LSU/coherence/interconnect subsystem.
|
||||||
|
|
||||||
|
> **Open:** Whether XH-1 advertises `Ztso` (an optional RVWMO relaxation
|
||||||
|
> that grants TSO at every hart) or only RVWMO. `Ztso` simplifies
|
||||||
|
> software reasoning but slightly constrains hardware ordering
|
||||||
|
> flexibility. Recommendation deferred until microarchitecture is
|
||||||
|
> fixed.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. Required Fences and Their Microarchitectural Meaning
|
||||||
|
|
||||||
|
The `FENCE` instruction is parameterized by:
|
||||||
|
|
||||||
|
- `pred`: which predecessor operations are ordered (R, W).
|
||||||
|
- `succ`: which successor operations are ordered (R, W).
|
||||||
|
- `fm` field: indicates "fence-mode"; value `0` is the standard fence,
|
||||||
|
value `>=1` is reserved for implementations and is treated as a no-op
|
||||||
|
by standard tools. Custom XH-1 fence-mode encodings are **[Open]**.
|
||||||
|
|
||||||
|
XH-1 must ensure that a `FENCE` with predecessor set *P* and successor
|
||||||
|
set *S*:
|
||||||
|
|
||||||
|
- All memory operations of type *P* issued by this hart prior to the
|
||||||
|
`FENCE` are ordered before all memory operations of type *S* issued by
|
||||||
|
this hart after the `FENCE`, in the global memory order, to the extent
|
||||||
|
that the coherence protocol is willing to serialize them.
|
||||||
|
|
||||||
|
Microarchitecturally, the `FENCE` typically drains in-flight operations
|
||||||
|
of the predecessor classes through the load/store unit (LSU), the
|
||||||
|
coherence point, and the write-completion path. On XH-1 with private L1
|
||||||
|
and a shared L2 slice, "the coherence point" is the directory controller
|
||||||
|
on the L2 slice; see §7.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. Why 128 Cores Changes the Problem
|
||||||
|
|
||||||
|
A small-count many-core can often "get away" with conservative
|
||||||
|
ordering: a single shared bus or ring, broadcast snoops, and a
|
||||||
|
full-fence on every store buffer drain. At 128 cores these approaches
|
||||||
|
do not scale:
|
||||||
|
|
||||||
|
1. **Bandwidth:** A serialized global fence requires all in-flight
|
||||||
|
traffic from the issuing core to drain, *and* the directory to
|
||||||
|
observe all acknowledgments before subsequent operations are
|
||||||
|
injected. With 128 cores contending for the directory and the
|
||||||
|
interconnect, the latency of a single `FENCE` can grow from
|
||||||
|
~10–30 cycles (small system) to several hundred cycles. This
|
||||||
|
directly impacts lock-based critical sections, RCU grace periods,
|
||||||
|
and `atomic` operations in the OS.
|
||||||
|
|
||||||
|
2. **Queue occupancy:** Every core holds per-message-class queues
|
||||||
|
(read requests, write requests, invalidations, data responses,
|
||||||
|
writeback data). A fence must account for occupancy of queues
|
||||||
|
between the core and the L2 directory *and* between directories
|
||||||
|
across the mesh. Without back-pressure, an XH-1 fence could
|
||||||
|
complete locally but leave a "hole" in the global order.
|
||||||
|
|
||||||
|
3. **Coherence protocol overhead:** Directory-based coherence
|
||||||
|
(assumed for XH-1; see the coherence research document) replaces
|
||||||
|
broadcast snoop traffic with per-line directory lookups, but
|
||||||
|
introduces ack/grant ordering dependencies that the fence must
|
||||||
|
not skip past.
|
||||||
|
|
||||||
|
4. **Scalable coherence introduces non-atomic transitions.**
|
||||||
|
Directory state moves from M→I→S→M through intermediate invalidation
|
||||||
|
acknowledgment states. A store-conditional (`SC`) on one core that
|
||||||
|
has been racing with a remote `LR` may observe the line in
|
||||||
|
intermediate states; the memory model must be specified for these.
|
||||||
|
|
||||||
|
5. **Contention of fence-induced drains:** When 128 cores all execute
|
||||||
|
`FENCE` in a short window (e.g., a barrier in a parallel workload),
|
||||||
|
the interconnect can become congested with drain acknowledgments.
|
||||||
|
The microarchitecture must avoid feedback-loop deadlock between
|
||||||
|
the queues and the fence-acknowledgment path.
|
||||||
|
|
||||||
|
> **Proposed XH-1 stance:** Treat memory ordering as a system-level
|
||||||
|
> performance constraint, not just a per-core correctness one. Evaluate
|
||||||
|
> the cost of every `FENCE` site in the OS, runtime, and synchronization
|
||||||
|
> libraries at the system level, not at single-core simulation.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. Microarchitectural Mechanisms
|
||||||
|
|
||||||
|
The classic toolkit, applied to XH-1:
|
||||||
|
|
||||||
|
### 5.1 Load-Store Unit (LSU) Pipeline
|
||||||
|
|
||||||
|
A typical in-order XH-1 core will have:
|
||||||
|
|
||||||
|
- A **store buffer (STB)** between the integer pipeline and the L1
|
||||||
|
data cache.
|
||||||
|
- A **load queue (LDQ)** that issues loads past stores (speculatively)
|
||||||
|
and validates the result at commit, rolling back on a misspeculation.
|
||||||
|
- A **memory dependence predictor** (optional; see §9.1) to reduce the
|
||||||
|
cost of Ld→St forwarding checks.
|
||||||
|
|
||||||
|
The LSU's job relative to ordering is twofold:
|
||||||
|
|
||||||
|
1. **Honor PPO locally** — enforce that a load following a store to the
|
||||||
|
same address returns the store's value (or a younger value), and
|
||||||
|
that the AMO-acquire/release rules produce the required PPO edges.
|
||||||
|
2. **Provide a clean drain point** for `FENCE` — a place where the LSU
|
||||||
|
can mark "all pred-class operations up to this point have completed
|
||||||
|
globally."
|
||||||
|
|
||||||
|
### 5.2 Store Buffer Drain
|
||||||
|
|
||||||
|
A `FENCE` with `pred` including W must drain the store buffer. In a
|
||||||
|
directory-coherent system the relevant completion is not "store reached
|
||||||
|
L1" but "store has acquired the directory ownership token (e.g., `M`
|
||||||
|
response) and the directory has acknowledged writeback/Invalidate
|
||||||
|
acknowledgments for any invalidated peers."
|
||||||
|
|
||||||
|
For XH-1:
|
||||||
|
|
||||||
|
- **L1 hit, no coherence traffic:** store can complete in ~1–3 cycles
|
||||||
|
after reaching L1.
|
||||||
|
- **L1 miss, L2 hit, no sharers:** invalidation may be skipped (silent
|
||||||
|
upgrade to M); store completes when directory grants M.
|
||||||
|
- **L1 miss, L2 hit, sharers exist:** directory generates invalidation
|
||||||
|
messages to all sharers; the store completes when all invalidation
|
||||||
|
acknowledgments are received.
|
||||||
|
- **L1 miss, L2 miss:** directory forwards to home node; home node
|
||||||
|
replies with data and grants M; the store completes when data and
|
||||||
|
the M grant are both received.
|
||||||
|
|
||||||
|
Only the last three cases impose significant drain latency.
|
||||||
|
|
||||||
|
> **Proposed XH-1 stance:** The store buffer must be sized to hold
|
||||||
|
> stores pending completion, but the fence-drain path must not depend
|
||||||
|
> on the store buffer draining in order — older stores in the buffer
|
||||||
|
> that have already received their M response should not block the
|
||||||
|
> fence on younger stores still in flight, *unless* the fence
|
||||||
|
> successor set requires it.
|
||||||
|
|
||||||
|
### 5.3 Load Issue Rules
|
||||||
|
|
||||||
|
Loads are the harder side. RVWMO allows loads to be reordered with
|
||||||
|
prior stores to different addresses. XH-1's LSU will speculatively
|
||||||
|
issue loads past older stores; the load is "validated" at commit by
|
||||||
|
checking that the cache line has not been invalidated and that no
|
||||||
|
older store wrote to the same address.
|
||||||
|
|
||||||
|
A load that needs to "see" the result of an older store on the same
|
||||||
|
hart uses **store-to-load forwarding** through the store buffer; this
|
||||||
|
must respect the size and sign-extension rules of the load.
|
||||||
|
|
||||||
|
A load that needs to "see" the result of an older store on a *different*
|
||||||
|
hart (which is what `FENCE`/`FENCE.RW` enforces) must wait for that
|
||||||
|
remote store to complete at the directory and update the local
|
||||||
|
view via the coherence protocol. The directory's response is the
|
||||||
|
"happens-before" edge.
|
||||||
|
|
||||||
|
### 5.4 Acquire/Release on AMOs
|
||||||
|
|
||||||
|
A standard `AMO*` operation in RVWMO has the ordering semantics of an
|
||||||
|
acquire (load part) and release (store part) when used in `aq`/`rl`
|
||||||
|
positions. In RV32A/RV64A, the instructions `LR` and `SC` carry
|
||||||
|
acquire/release bits; AMOs (e.g., `AMOSWAP.W`, `AMOADD.D`) also carry
|
||||||
|
these bits.
|
||||||
|
|
||||||
|
For XH-1, an `aq` AMO must order prior loads/stores from the issuing
|
||||||
|
hart before the AMO, and the AMO before subsequent loads/stores.
|
||||||
|
Microarchitecturally:
|
||||||
|
|
||||||
|
- `aq`: drain the load and store queues (or at least establish PPO
|
||||||
|
edges) before issuing the AMO to the directory.
|
||||||
|
- `rl`: drain the load and store queues *after* the AMO completes.
|
||||||
|
- `aq+rl`: drain on both sides; this is the strong fence.
|
||||||
|
|
||||||
|
> **Open:** Whether XH-1 implements `aq`/`rl` purely in the LSU
|
||||||
|
> (with no global fence message) or with a lightweight fence message
|
||||||
|
> to the directory. Purely-local enforcement is faster but requires
|
||||||
|
> that the directory serialize the AMO at the line state machine.
|
||||||
|
|
||||||
|
### 5.5 `LR`/`SC`
|
||||||
|
|
||||||
|
`LR` loads with reservation; `SC` stores conditionally if the
|
||||||
|
reservation is still valid. XH-1's reservation representation is
|
||||||
|
proposed as a small set of reservation-set entries per core, indexed
|
||||||
|
by physical address, anchored in the L1. Coherence actions that
|
||||||
|
invalidate the reserved line (a snoop/Invalidate that the L1 sees)
|
||||||
|
must invalidate the reservation.
|
||||||
|
|
||||||
|
> **Proposed XH-1 reservation-granularity choice:** XH-1 reserves at
|
||||||
|
> cache-line granularity (the natural L1 line size) and treats
|
||||||
|
> forward-progress as the OS's responsibility. This is the
|
||||||
|
> conservative default; some designs use sub-line reservation to
|
||||||
|
> improve SC success rates, but the gain on typical lock code is
|
||||||
|
> modest.
|
||||||
|
|
||||||
|
### 5.6 `FENCE.I` and Self-Modifying Code
|
||||||
|
|
||||||
|
`FENCE.I` ensures that stores from the issuing hart that preceded the
|
||||||
|
`FENCE.I` are visible to subsequent instruction fetches on the same
|
||||||
|
hart. It does **not** order stores from other harts.
|
||||||
|
|
||||||
|
For XH-1, `FENCE.I` typically:
|
||||||
|
|
||||||
|
1. Drains the store buffer (in the W sense).
|
||||||
|
2. Invalidates the L1 I-cache (or the local portion of a unified L1).
|
||||||
|
3. Optionally issues a `FENCE` internally so that the store is
|
||||||
|
visible at the directory before the I-fetch resumes.
|
||||||
|
|
||||||
|
A cross-hart self-modifying-code sequence (hart A writes code, hart B
|
||||||
|
executes it) requires an explicit `FENCE` on hart A and a `FENCE.I` on
|
||||||
|
hart B (per the RISC-V recommendation), plus coherence to propagate
|
||||||
|
the new data. This is awkward and XH-1 will document it in the
|
||||||
|
programmer's manual.
|
||||||
|
|
||||||
|
### 5.7 `SFENCE.VMA`
|
||||||
|
|
||||||
|
`SFENCE.VMA` orders prior stores to PTEs before subsequent
|
||||||
|
address-translation results on the same hart. XH-1's TLB refill
|
||||||
|
engine must observe the store-completion signal; a typical
|
||||||
|
implementation stalls the pipeline at the `SFENCE.VMA` until the
|
||||||
|
store buffer drains.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 6. System-Level Ordering Devices
|
||||||
|
|
||||||
|
Beyond per-hart mechanisms, XH-1 needs system-level ordering for
|
||||||
|
I/O and devices. The RISC-V privileged specification provides two
|
||||||
|
mechanisms:
|
||||||
|
|
||||||
|
### 6.1 I/O Ordering Bits in PTE
|
||||||
|
|
||||||
|
A PTE can have the **PBMT** field (Page-Based Memory Types), which can
|
||||||
|
be marked as:
|
||||||
|
|
||||||
|
- **Non-cacheable, idempotent, weakly-ordered (NC, IO).**
|
||||||
|
- **Non-cacheable, idempotent, strongly-ordered (NC, I/O-strong).**
|
||||||
|
|
||||||
|
XH-1 must honor these by:
|
||||||
|
|
||||||
|
- Routing accesses to the appropriate device port on the interconnect.
|
||||||
|
- For weakly-ordered I/O, treating the access as device but allowing
|
||||||
|
reordering with respect to other weakly-ordered I/O of the same type.
|
||||||
|
- For strongly-ordered I/O, treating the access as a full-fence —
|
||||||
|
the load or store completes in program order with respect to all
|
||||||
|
other accesses from the same hart.
|
||||||
|
|
||||||
|
### 6.2 Memory-Mapped I/O Fences
|
||||||
|
|
||||||
|
XH-1 will provide a `FENCE` variant (the standard `FENCE pred=rw,
|
||||||
|
succ=rw, fm=0`) that software uses to order I/O accesses. The
|
||||||
|
microarchitecture must ensure the drain endpoint includes the
|
||||||
|
device-port queue, not just the L1/L2.
|
||||||
|
|
||||||
|
### 6.3 Interrupt Acknowledgment and Ordering
|
||||||
|
|
||||||
|
RISC-V interrupt acknowledgment is itself a memory-mapped I/O
|
||||||
|
write. The microarchitecture must guarantee that the write to the
|
||||||
|
*Platform-Level Interrupt Controller* (PLIC) is ordered with respect
|
||||||
|
to the I/O operations it is acknowledging. The standard idiom is a
|
||||||
|
`FENCE` before the claim/complete operation. XH-1's PLIC integration
|
||||||
|
must be designed so that the `FENCE` does not require extra
|
||||||
|
microarchitectural support beyond the standard fence.
|
||||||
|
|
||||||
|
> **Open:** Whether XH-1 implements a **"fence fast path"** for
|
||||||
|
> accesses to specific device memory regions (e.g., the PLIC) that
|
||||||
|
> bypasses the directory and uses a dedicated sideband signal.
|
||||||
|
> Such optimizations can reduce interrupt latency by tens of cycles
|
||||||
|
> but complicate coherence reasoning. Not recommended for the first
|
||||||
|
> silicon.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 7. Coherence Interaction
|
||||||
|
|
||||||
|
XH-1's coherence protocol (described in the coherence research
|
||||||
|
document) is directory-based, MSI or MESI, with a distributed
|
||||||
|
directory across the L2 slices. Memory ordering depends on the
|
||||||
|
coherence protocol's notion of "completion."
|
||||||
|
|
||||||
|
For a store to be considered "globally performed":
|
||||||
|
|
||||||
|
- The line is held in **Modified** state by the requesting core.
|
||||||
|
- All prior sharers have sent Invalidate Acknowledgments to the
|
||||||
|
directory (or the directory has confirmed there were no sharers).
|
||||||
|
- The directory has logged the new owner.
|
||||||
|
|
||||||
|
For a load to be considered "globally performed":
|
||||||
|
|
||||||
|
- The line is held in **Shared** (or **Modified** if exclusive) by
|
||||||
|
the requesting core.
|
||||||
|
- The data response from the directory (or forwarded from a peer) is
|
||||||
|
the value the directory currently authorizes.
|
||||||
|
|
||||||
|
> **Important subtlety:** A `FENCE W,W` between two stores on the
|
||||||
|
> same hart must ensure that store 1 has acquired its M grant *and*
|
||||||
|
> the directory has processed the invalidations *before* store 2
|
||||||
|
> acquires its M grant. In directory terms, store 2 cannot be issued
|
||||||
|
> until store 1's M grant has been observed at the directory. XH-1
|
||||||
|
> must implement this as an ordering dependency at the directory's
|
||||||
|
> L2 controller, not as a per-line serialization at the directory
|
||||||
|
> (which would be too coarse and too slow).
|
||||||
|
|
||||||
|
### 7.1 Coherence-Induced Ordering Pitfalls
|
||||||
|
|
||||||
|
Common pitfalls in directory-based designs:
|
||||||
|
|
||||||
|
- **Directory ack starvation under fence pressure:** when many cores
|
||||||
|
fence simultaneously, the directory must serialize invalidation
|
||||||
|
acknowledgments without deadlock.
|
||||||
|
- **Coherence downgrade during fence:** a core that downgrades M→S
|
||||||
|
while a fence is in progress must not let younger loads proceed
|
||||||
|
past the fence.
|
||||||
|
- **Silent upgrade (no-invalidate transition) on store:** a clean
|
||||||
|
line upgraded to M without invalidation must still be ordered
|
||||||
|
with respect to other cores' loads that observed the line in S.
|
||||||
|
The directory must explicitly synchronize this transition.
|
||||||
|
|
||||||
|
### 7.2 Ordering at the Interconnect
|
||||||
|
|
||||||
|
XH-1's interconnect is a mesh (proposed; see interconnect research
|
||||||
|
document). The mesh delivers messages in-order *between a given pair
|
||||||
|
of routers*, but globally the order is not guaranteed. The directory
|
||||||
|
controller is the serialization point.
|
||||||
|
|
||||||
|
For ordering, XH-1's mesh must:
|
||||||
|
|
||||||
|
- Not reorder different message classes on the same virtual channel.
|
||||||
|
- Support a "tail" or "credit" signal so that a fence-acknowledgment
|
||||||
|
message can be generated only when all prior messages on the same
|
||||||
|
path have been delivered.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 8. Implementation Approaches Compared
|
||||||
|
|
||||||
|
This section lists the realistic implementation strategies for an
|
||||||
|
in-order XH-1 core. The choice of in-order vs. out-of-order pipeline
|
||||||
|
is discussed in §9.1; here we focus on ordering-specific mechanisms.
|
||||||
|
|
||||||
|
### 8.1 Approach A: Conservative Per-Core Fence (Baseline)
|
||||||
|
|
||||||
|
- The LSU drains the relevant queue classes fully on every `FENCE`.
|
||||||
|
- The store buffer's "drain complete" is defined as "every store has
|
||||||
|
received its directory M grant."
|
||||||
|
- A `FENCE` is implemented as a special instruction that stalls the
|
||||||
|
pipeline until drain completes.
|
||||||
|
- Acquire/release AMOs are implemented as full `FENCE`s around the
|
||||||
|
AMO at the LSU.
|
||||||
|
- `SFENCE.VMA` stalls the pipeline until the store buffer drains.
|
||||||
|
|
||||||
|
**Pros:**
|
||||||
|
|
||||||
|
- Simple to verify. The fence semantics maps directly to "queue
|
||||||
|
empty."
|
||||||
|
- Predictable performance for small core counts.
|
||||||
|
- Easy to reason about coherence interactions.
|
||||||
|
|
||||||
|
**Cons:**
|
||||||
|
|
||||||
|
- Poor scalability. At 128 cores, the worst-case `FENCE` latency
|
||||||
|
grows because the drain point is global, not local.
|
||||||
|
- Overly conservative for `aq`/`rl` AMOs, which do not require a
|
||||||
|
full global drain on either side — only ordering relative to
|
||||||
|
surrounding ordinary accesses.
|
||||||
|
|
||||||
|
### 8.2 Approach B: Token-Based Completion
|
||||||
|
|
||||||
|
- The LSU attaches a monotonically increasing **completion token**
|
||||||
|
to each outstanding operation.
|
||||||
|
- A `FENCE` with pred P, succ S becomes: "wait until all operations
|
||||||
|
of type P issued before this `FENCE` have produced a token greater
|
||||||
|
than or equal to the fence's token."
|
||||||
|
- The token is generated when the directory acknowledges the
|
||||||
|
operation; this is exactly the "global completion" notion.
|
||||||
|
- Acquire/Release is implemented as a fence of the appropriate
|
||||||
|
type around the AMO at the LSU.
|
||||||
|
|
||||||
|
**Pros:**
|
||||||
|
|
||||||
|
- Fences do not require the entire store buffer to drain; an empty
|
||||||
|
store buffer is sufficient, but so is a store buffer where all
|
||||||
|
older stores have completed.
|
||||||
|
- Naturally parallel: many stores can be outstanding, and the
|
||||||
|
fence only waits for the *oldest* of them.
|
||||||
|
- Slightly more complex but a well-understood technique (used in
|
||||||
|
academic and commercial cores).
|
||||||
|
|
||||||
|
**Cons:**
|
||||||
|
|
||||||
|
- Per-message token bookkeeping is more expensive than a simple
|
||||||
|
"queue empty" check.
|
||||||
|
- Verification requires modeling tokens at the testbench, which is
|
||||||
|
a non-trivial simulation infrastructure investment.
|
||||||
|
|
||||||
|
### 8.3 Approach C: Directory-Assisted Ordering
|
||||||
|
|
||||||
|
- A special "fence-ack" message class is added to the coherence
|
||||||
|
protocol.
|
||||||
|
- A `FENCE` instruction sends a fence-ack request to the directory;
|
||||||
|
the directory replies only after all prior messages from the
|
||||||
|
issuing core have been serialized.
|
||||||
|
- The directory tracks per-core "last-seen" sequence numbers to
|
||||||
|
arbitrate fence-ack requests against ongoing traffic.
|
||||||
|
|
||||||
|
**Pros:**
|
||||||
|
|
||||||
|
- The directory is the natural serialization point; leveraging it
|
||||||
|
avoids per-core reinvention of global ordering.
|
||||||
|
- Scales well to large core counts because the directory already
|
||||||
|
handles 128-core traffic.
|
||||||
|
- Allows the per-core LSU to continue issuing younger operations
|
||||||
|
while the fence is in progress, as long as they are ordered
|
||||||
|
correctly.
|
||||||
|
|
||||||
|
**Cons:**
|
||||||
|
|
||||||
|
- Adds protocol complexity (a new message class, directory state,
|
||||||
|
ack/grant logic).
|
||||||
|
- Risk of fence-ack starvation under heavy traffic.
|
||||||
|
- Requires careful design to avoid deadlock with normal coherence
|
||||||
|
traffic (classic message-class-dependency cycle).
|
||||||
|
|
||||||
|
### 8.4 Approach D: Speculative OoO + MOB
|
||||||
|
|
||||||
|
This is the out-of-order approach. The LSU becomes a Memory
|
||||||
|
Ordering Buffer (MOB) that tracks per-operation completion and
|
||||||
|
issues fences by tracking the "oldest in-flight" of each class.
|
||||||
|
This is how high-performance OoO cores implement fences; the
|
||||||
|
trade-off is pipeline depth and complexity.
|
||||||
|
|
||||||
|
XH-1 is proposed as **in-order** for power and predictability; see
|
||||||
|
§9.1. OoO is listed for completeness.
|
||||||
|
|
||||||
|
**Pros:**
|
||||||
|
|
||||||
|
- Highest performance on workloads with deep memory-level
|
||||||
|
parallelism.
|
||||||
|
- Fences can be "free" if no relevant older operation is in flight.
|
||||||
|
|
||||||
|
**Cons:**
|
||||||
|
|
||||||
|
- Significantly larger area and power.
|
||||||
|
- More verification effort, especially for the MOB's speculative
|
||||||
|
reordering.
|
||||||
|
- Diminishing returns on memory-bound workloads at 128 cores (the
|
||||||
|
bottleneck is the interconnect, not the per-core issue rate).
|
||||||
|
|
||||||
|
### 8.5 Summary of Approaches
|
||||||
|
|
||||||
|
| Approach | Perf (single) | Perf (128) | Area | Verif. effort | Risk |
|
||||||
|
| -------- | ------------- | ---------- | ---- | ------------- | ---- |
|
||||||
|
| A: Cons. | Low | Lowest | Lowest | Lowest | Low |
|
||||||
|
| B: Token | Medium | Medium | Low | Medium | Low |
|
||||||
|
| C: Dir-Assisted | Medium-high | High | Medium | High | Medium |
|
||||||
|
| D: OoO+MOB | High | High | High | Highest | High |
|
||||||
|
|
||||||
|
> **Proposed XH-1 implementation:** **Approach C (Directory-Assisted
|
||||||
|
> Ordering)** as the primary, with **Approach B (Token-Based
|
||||||
|
> Completion)** as the per-core implementation strategy. The
|
||||||
|
> directory handles global serialization, the LSU handles local
|
||||||
|
> token tracking. This combination is the most natural for a
|
||||||
|
> directory-coherent 128-core design.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 9. Interactions with the Pipeline and Other Subsystems
|
||||||
|
|
||||||
|
### 9.1 Pipeline Choice (In-Order vs. Out-of-Order)
|
||||||
|
|
||||||
|
XH-1 is **[Proposed] in-order** for the first silicon. The reasons:
|
||||||
|
|
||||||
|
- Power envelope. 128 cores at high frequency is a thermal
|
||||||
|
challenge; an in-order core can run cooler.
|
||||||
|
- Predictable timing. Real-time and safety-critical workloads
|
||||||
|
benefit from deterministic `FENCE` and `FENCE.I` latencies.
|
||||||
|
- Verification cost. An OoO core's MOB is a major verification
|
||||||
|
burden; a 128-core OoO design is a multi-year project.
|
||||||
|
|
||||||
|
An in-order pipeline interacts with ordering as follows:
|
||||||
|
|
||||||
|
- The LSU is non-speculative on the store side (stores are
|
||||||
|
committed when the integer pipeline commits them) and lightly
|
||||||
|
speculative on the load side (loads may issue before older
|
||||||
|
stores, with a structural dependency check at commit).
|
||||||
|
- `FENCE` stalls the pipeline until drain; the cost is a few tens
|
||||||
|
to a few hundred cycles depending on directory state.
|
||||||
|
- Acquire/release AMOs generate local fence events in the LSU.
|
||||||
|
|
||||||
|
An OoO pipeline would:
|
||||||
|
|
||||||
|
- Issue stores from the STB to L1, with the MOB tracking
|
||||||
|
completion.
|
||||||
|
- Issue loads speculatively past stores with prediction; the MOB
|
||||||
|
validates at commit.
|
||||||
|
- `FENCE` becomes a "wait for oldest in-flight of pred class to
|
||||||
|
complete" check; usually very fast.
|
||||||
|
|
||||||
|
> **Open:** Whether a future XH-2 could be OoO. The research
|
||||||
|
> project is not yet at that question.
|
||||||
|
|
||||||
|
### 9.2 Cache Hierarchy
|
||||||
|
|
||||||
|
Memory ordering is shaped by the cache hierarchy:
|
||||||
|
|
||||||
|
- **L1 D-cache:** the load-side ordering reference. A load that hits
|
||||||
|
in L1 sees a value that is at least as recent as the cache's
|
||||||
|
coherence state indicates. A load that misses in L1 must go to
|
||||||
|
L2, where the directory arbitrates the order.
|
||||||
|
- **L2 slice (per tile):** home of the directory entry; serializes
|
||||||
|
ordering for the addresses it owns.
|
||||||
|
- **L3 (shared, distributed):** a victim cache and last-level
|
||||||
|
coherence point. XH-1's directory is the L2 directory; L3 is
|
||||||
|
treated as a backing store. Invalidation ordering for XH-1
|
||||||
|
therefore ends at L2.
|
||||||
|
- **L1 I-cache:** the `FENCE.I` target. The I-cache must be flushed
|
||||||
|
or invalidated to ensure stores are visible.
|
||||||
|
|
||||||
|
The cache hierarchy directly affects `FENCE` latency because the
|
||||||
|
"completion" point is the L2 directory. A `FENCE W,W` in a system
|
||||||
|
with L1-only ordering would be much cheaper (a few cycles) but
|
||||||
|
incorrect; the L2 directory must be in the loop.
|
||||||
|
|
||||||
|
### 9.3 Memory System
|
||||||
|
|
||||||
|
The memory system (DRAM controllers, on-chip memory, scratchpads)
|
||||||
|
must respect the same ordering. For XH-1, accesses to on-chip
|
||||||
|
scratchpad memory bypass the coherence protocol entirely; they
|
||||||
|
are ordered by the scratchpad controller.
|
||||||
|
|
||||||
|
> **Open:** Whether XH-1 has a per-tile scratchpad or a shared
|
||||||
|
> on-chip SRAM region. If shared, ordering between scratchpad and
|
||||||
|
cache-coherent regions is a research question.
|
||||||
|
|
||||||
|
### 9.4 Interconnect
|
||||||
|
|
||||||
|
The mesh interconnect has ordering properties (per-link in-order,
|
||||||
|
globally non-ordered) that determine what the directory and
|
||||||
|
endpoints must do to maintain memory ordering. The "virtual
|
||||||
|
channel" abstraction in the mesh is the unit of ordering; XH-1
|
||||||
|
**[Proposed]** uses:
|
||||||
|
|
||||||
|
- **VC0:** Coherence requests (in-order).
|
||||||
|
- **VC1:** Coherence responses (in-order).
|
||||||
|
- **VC2:** Data messages (in-order).
|
||||||
|
- **VC3:** Fence-ack and other ordering-sensitive messages
|
||||||
|
(in-order).
|
||||||
|
|
||||||
|
This separation avoids deadlock and simplifies ordering reasoning.
|
||||||
|
|
||||||
|
### 9.5 Interrupts
|
||||||
|
|
||||||
|
RISC-V interrupts are delivered to a hart based on the *global
|
||||||
|
interrupt-enable* bits and the PLIC. Ordering constraints:
|
||||||
|
|
||||||
|
- A hart that disables interrupts and then writes to a shared
|
||||||
|
variable does not need a `FENCE` before the interrupt-disable
|
||||||
|
write; the write to `mstatus` or `sstatus` is itself a CSR write
|
||||||
|
and is ordered architecturally.
|
||||||
|
- A hart that claims an interrupt at the PLIC and then reads a
|
||||||
|
device register must `FENCE` between the claim and the read to
|
||||||
|
ensure the claim is globally visible (the claim is a load from
|
||||||
|
the PLIC's claim register; the read is a load to the device).
|
||||||
|
- The interrupt-return path uses `mret`/`sret`, which is
|
||||||
|
architecturally ordered with respect to prior CSR writes.
|
||||||
|
- Trap delivery and `xRET` do not include implicit memory fences;
|
||||||
|
the OS must insert them.
|
||||||
|
|
||||||
|
XH-1's interrupt controller integration must ensure that the
|
||||||
|
interrupt-acknowledge signal is recognized only after the
|
||||||
|
relevant memory accesses are ordered.
|
||||||
|
|
||||||
|
### 9.6 Operating System
|
||||||
|
|
||||||
|
The OS uses `FENCE` heavily:
|
||||||
|
|
||||||
|
- **Spinlocks and mutexes:** the lock acquire is typically a
|
||||||
|
`LR`/`SC` sequence; the lock release is an `SC` or a `FENCE`
|
||||||
|
followed by a `SW` to the lock variable.
|
||||||
|
- **Page-table manipulation:** `SFENCE.VMA` after every PTE write.
|
||||||
|
- **I/O:** `FENCE` before/after device accesses (see §6).
|
||||||
|
- **Cross-core TLB shootdown:** an IPI sender writes the
|
||||||
|
invalidation request to a shared memory location, then sends an
|
||||||
|
IPI, and the receiver must see the request after the IPI — this
|
||||||
|
requires a `FENCE` on the sender side.
|
||||||
|
- **RCU grace periods:** RCU relies on the memory model's
|
||||||
|
guarantees; a weak model like RVWMO requires explicit
|
||||||
|
`FENCE` at grace-period boundaries.
|
||||||
|
- **Boot and SMP startup:** the boot hart must `FENCE` before
|
||||||
|
starting secondary harts to ensure the trampoline code and data
|
||||||
|
are visible.
|
||||||
|
|
||||||
|
XH-1's OS port (Linux, or a research kernel) must insert `FENCE`
|
||||||
|
at all these sites. The cost of a `FENCE` in a 128-core system
|
||||||
|
makes the boot-time and idle-time costs small in absolute terms
|
||||||
|
but potentially large in *relative* terms; an in-order core that
|
||||||
|
spends 200 cycles per `FENCE` will spend a noticeable fraction of
|
||||||
|
its time in fences if it idles on a spinlock.
|
||||||
|
|
||||||
|
> **Open:** Whether XH-1 ports a vanilla Linux or a research
|
||||||
|
> kernel. The memory-ordering analysis in this document assumes
|
||||||
|
> vanilla Linux semantics but applies in both cases.
|
||||||
|
|
||||||
|
### 9.7 Verification
|
||||||
|
|
||||||
|
Memory ordering is famously difficult to verify because:
|
||||||
|
|
||||||
|
- The bug is typically a very rare, non-deterministic interleaving.
|
||||||
|
- The legal reorderings are many; enumerating them is hard.
|
||||||
|
- The coherence protocol and the memory model interact in subtle
|
||||||
|
ways.
|
||||||
|
|
||||||
|
XH-1's verification strategy **[Proposed]**:
|
||||||
|
|
||||||
|
1. **Litmus tests** for the memory model. Run a corpus of litmus
|
||||||
|
tests (based on the RISC-V memory-model litmus test suite, plus
|
||||||
|
XH-1-specific tests) on RTL simulation. The tests must include
|
||||||
|
the published RVWMO tests and tests specific to XH-1's
|
||||||
|
coherence protocol.
|
||||||
|
2. **Coherence protocol model checking.** Use a formal model
|
||||||
|
(e.g., Murphi, TLA+) of the directory protocol; prove
|
||||||
|
invariant properties.
|
||||||
|
3. **Pipeline formal verification.** Verify the LSU's local
|
||||||
|
reordering rules with bounded model checking.
|
||||||
|
4. **System-level simulation.** Boot Linux on the RTL or a
|
||||||
|
high-level model; run stress workloads (e.g., `perf bench
|
||||||
|
futex`, RCU torture test).
|
||||||
|
5. **Stress with random delay injection.** Vary latency on each
|
||||||
|
interconnect path randomly to expose ordering bugs.
|
||||||
|
|
||||||
|
The published **herd7** tool from Cambridge and the **diy7**
|
||||||
|
litmus-test framework are the standard tools for RVWMO
|
||||||
|
verification. **Operational models** like Zam are more amenable
|
||||||
|
to formal reasoning than axiomatic models; XH-1 should
|
||||||
|
**[Proposed]** support both for verification.
|
||||||
|
|
||||||
|
> **Open:** Whether XH-1's verification flow includes formal
|
||||||
|
> proofs of the memory ordering contract end-to-end, or whether
|
||||||
|
> simulation-based testing is the primary line of defense. A
|
||||||
|
> research project may choose the latter; a commercial project
|
||||||
|
> should aim for the former.
|
||||||
|
|
||||||
|
### 9.8 Performance
|
||||||
|
|
||||||
|
Memory ordering has two performance axes:
|
||||||
|
|
||||||
|
1. **Steady-state throughput:** how fast can the system process
|
||||||
|
ordinary memory traffic. This is dominated by coherence and
|
||||||
|
cache behavior, not by ordering per se.
|
||||||
|
2. **Fence/atomic latency and throughput:** how fast is a
|
||||||
|
synchronization primitive. This is where ordering directly
|
||||||
|
impacts performance.
|
||||||
|
|
||||||
|
For a 128-core system, the second axis is the more important
|
||||||
|
contributor to scalability. A workload with frequent
|
||||||
|
synchronization (e.g., a parallel database, a many-core
|
||||||
|
machine-learning inference engine) is dominated by `FENCE` and
|
||||||
|
AMO latency.
|
||||||
|
|
||||||
|
> **Quantitative note (rough estimate):** Assuming a 1 GHz in-order
|
||||||
|
> core, a directory-coherent L2 slice with a hop latency of ~10 ns
|
||||||
|
> per mesh traversal, and an average of 3–5 hops per coherence
|
||||||
|
> transaction, a single uncontended `FENCE` on a hot line might
|
||||||
|
> take ~30–50 ns (30–50 cycles). A contended `FENCE` (multiple
|
||||||
|
> sharers) requiring invalidation acks could take 100–300 ns
|
||||||
|
> (100–300 cycles). On a 128-core workload that fences once per
|
||||||
|
> critical section, this is the dominant cost.
|
||||||
|
|
||||||
|
XH-1's performance evaluation **[Proposed]**:
|
||||||
|
|
||||||
|
- Measure `FENCE` latency distribution under varying contention
|
||||||
|
and across the mesh (corner vs. center tile).
|
||||||
|
- Measure AMO latency under contention.
|
||||||
|
- Measure the effect of `aq`/`rl` vs. explicit `FENCE` in the
|
||||||
|
OS.
|
||||||
|
- Compare against an ideal "no ordering cost" baseline.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 10. Specific RISC-V Considerations
|
||||||
|
|
||||||
|
### 10.1 The `Ztso` Extension (Optional)
|
||||||
|
|
||||||
|
`Ztso` provides TSO (Total Store Order) at every hart. With
|
||||||
|
`Ztso`, the implementation guarantees that stores are not
|
||||||
|
reordered with respect to subsequent loads, simplifying the
|
||||||
|
software model.
|
||||||
|
|
||||||
|
> **Proposed XH-1 stance:** Do not advertise `Ztso` for the
|
||||||
|
> first silicon. The hardware cost of supporting TSO at every
|
||||||
|
> hart is a tighter store-load ordering in the LSU, and the
|
||||||
|
> benefit is only felt by software that is willing to be
|
||||||
|
> incompatible with weak-ordering implementations. The standard
|
||||||
|
> RISC-V Linux kernel is RVWMO-compatible and does not require
|
||||||
|
> `Ztso`. Re-evaluate if a major software target requires it.
|
||||||
|
|
||||||
|
### 10.2 The `Zam` Extension (If Adopted)
|
||||||
|
|
||||||
|
`Zam` is a proposed extension that provides an operational
|
||||||
|
memory model in addition to the axiomatic RVWMO. It does not
|
||||||
|
change the contract; it provides operational clarity. XH-1
|
||||||
|
**[Proposed]** will document its memory model in terms of `Zam`
|
||||||
|
where possible because operational models are easier to
|
||||||
|
implement and verify.
|
||||||
|
|
||||||
|
### 10.3 The PTE-Based Memory Type (PBMT)
|
||||||
|
|
||||||
|
XH-1's MMU must support the PBMT field and route accesses to
|
||||||
|
the appropriate ordering class. Specifically:
|
||||||
|
|
||||||
|
- **PMA = NC (non-cacheable):** accesses go to the device port;
|
||||||
|
load/store ordering is enforced by the device-port controller.
|
||||||
|
- **PMA = IO (weakly ordered I/O):** same as NC, but with
|
||||||
|
explicit reordering rules.
|
||||||
|
- **PMA = WEAK (cacheable but weakly ordered):** accesses are
|
||||||
|
cached but ordered weakly; useful for write-combining buffers.
|
||||||
|
- **PMA = STRONG (cacheable, strongly ordered):** same as
|
||||||
|
cacheable but with fence semantics on every access; rarely
|
||||||
|
used.
|
||||||
|
|
||||||
|
> **Open:** Whether XH-1 supports all PBMT values or a subset.
|
||||||
|
> A research project may support all four; a commercial project
|
||||||
|
> may support only NC and cacheable.
|
||||||
|
|
||||||
|
### 10.4 The `fence.i` and Instruction-Fetch Coherence
|
||||||
|
|
||||||
|
`FENCE.I` orders stores w.r.t. instruction fetches on the same
|
||||||
|
hart. Cross-hart, the software idiom is:
|
||||||
|
|
||||||
|
```c
|
||||||
|
// hart A:
|
||||||
|
write_code(addr, code);
|
||||||
|
__sync_synchronize(); // FENCE
|
||||||
|
send_ipi(hart_B);
|
||||||
|
|
||||||
|
// hart B:
|
||||||
|
handle_ipi();
|
||||||
|
__sync_synchronize(); // FENCE
|
||||||
|
__fence_i(); // FENCE.I
|
||||||
|
icache_flush();
|
||||||
|
jump_to(addr);
|
||||||
|
```
|
||||||
|
|
||||||
|
XH-1's I-cache is a VIPT cache; on an aliasing configuration
|
||||||
|
(small pages, large cache), the OS must use the recommended
|
||||||
|
sequence (or the larger page size). The architecture team will
|
||||||
|
document this.
|
||||||
|
|
||||||
|
### 10.5 TSO and RVWMO Differences for Software
|
||||||
|
|
||||||
|
Although the focus here is hardware, it is worth noting for
|
||||||
|
completeness:
|
||||||
|
|
||||||
|
- Under TSO, ordinary loads and stores have program order
|
||||||
|
semantics. Under RVWMO, they do not.
|
||||||
|
- A program that is correct under TSO may not be correct under
|
||||||
|
RVWMO without additional fences. The Linux kernel's
|
||||||
|
memory-model code (`tools/memory-model/`) contains the
|
||||||
|
required fences for RVWMO.
|
||||||
|
|
||||||
|
XH-1's port will use the RVWMO-correct fences; no special
|
||||||
|
support is required in hardware beyond honoring the standard
|
||||||
|
`FENCE`.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 11. Unresolved Design Questions
|
||||||
|
|
||||||
|
The following questions remain open for the XH-1 architecture:
|
||||||
|
|
||||||
|
1. **In-order vs. out-of-order pipeline.** In-order is proposed,
|
||||||
|
but the decision is not final. An OoO pipeline changes the
|
||||||
|
MOB design and may improve `FENCE` latency.
|
||||||
|
2. **`Ztso` adoption.** Whether to support TSO at every hart.
|
||||||
|
3. **`Zam` adoption.** Whether to require the operational memory
|
||||||
|
model in the specification.
|
||||||
|
4. **Coherence protocol.** MSI vs. MESI vs. MOESI. The choice
|
||||||
|
affects the number of transitions and therefore the number
|
||||||
|
of coherence-induced ordering edges. (See the coherence
|
||||||
|
research document.)
|
||||||
|
5. **Interconnect topology.** Mesh vs. ring vs. hierarchical
|
||||||
|
bus+mesh. The choice affects `FENCE` latency and the
|
||||||
|
directory's role.
|
||||||
|
6. **PBMT support scope.** Which PTE memory types XH-1
|
||||||
|
implements.
|
||||||
|
7. **Per-line vs. per-word reservation.** Affects `LR`/`SC`
|
||||||
|
success rates.
|
||||||
|
8. **Fence-ack message class.** Whether to add a dedicated
|
||||||
|
fence-ack message (Approach C) or rely entirely on
|
||||||
|
per-core token tracking (Approach B).
|
||||||
|
9. **Verification strategy.** Simulation-based vs. formal;
|
||||||
|
end-to-end memory model proofs.
|
||||||
|
10. **OS port.** Linux vs. a research kernel; this affects
|
||||||
|
the set of `FENCE` sites the system must support.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 12. Recommendations
|
||||||
|
|
||||||
|
Where evidence supports a recommendation, the following are
|
||||||
|
XH-1's proposed positions:
|
||||||
|
|
||||||
|
1. **Adopt RVWMO as the memory model.** This is the standard
|
||||||
|
RISC-V model and is the most widely supported by software.
|
||||||
|
2. **Implement an in-order pipeline with Approach C
|
||||||
|
(Directory-Assisted Ordering) and Approach B (Per-Core
|
||||||
|
Token Completion).** This combination scales to 128 cores
|
||||||
|
while keeping the per-core pipeline simple.
|
||||||
|
3. **Use a directory-based MESI coherence protocol at the
|
||||||
|
L2 slice.** The directory is the natural serialization
|
||||||
|
point for memory ordering. (See the coherence research
|
||||||
|
document for the full argument.)
|
||||||
|
4. **Add a dedicated fence-ack message class to the coherence
|
||||||
|
protocol.** This avoids overloading existing messages and
|
||||||
|
reduces verification ambiguity.
|
||||||
|
5. **Honour the standard `FENCE`, `FENCE.I`, `SFENCE.VMA`,
|
||||||
|
`LR`/`SC`, and AMO semantics.** No architectural deviations.
|
||||||
|
6. **Document the cross-hart self-modifying-code sequence in
|
||||||
|
the programmer's manual.** Use the standard fence pair.
|
||||||
|
7. **Implement the PTE-based memory types (PBMT) for at
|
||||||
|
least NC and cacheable.** Other types are deferred.
|
||||||
|
8. **Provide a memory-model litmus-test suite as part of
|
||||||
|
the verification deliverables.** This is essential for
|
||||||
|
catching ordering bugs.
|
||||||
|
9. **Measure `FENCE` and AMO latency across the mesh in
|
||||||
|
performance evaluation.** Corner tiles have different
|
||||||
|
latency from center tiles; the asymmetry is a real
|
||||||
|
performance concern.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 13. References
|
||||||
|
|
||||||
|
The following sources informed this document. URLs and
|
||||||
|
publication details are given where known; the reader should
|
||||||
|
verify currency, as the RISC-V ecosystem is rapidly evolving.
|
||||||
|
|
||||||
|
- RISC-V International, *RISC-V Instruction Set Manual,
|
||||||
|
Volume I: Unprivileged ISA*. Appendix A, "RVWMO: The RISC-V
|
||||||
|
Weak Memory Ordering Model."
|
||||||
|
- RISC-V International, *RISC-V Instruction Set Manual,
|
||||||
|
Volume II: Privileged ISA*. Sections on fence and memory
|
||||||
|
types (PBMT).
|
||||||
|
- Pulte, C., et al. *"Zam: A Complete and Verifiable
|
||||||
|
Operational Memory Model for RISC-V."* (Reference to the
|
||||||
|
operational memory model paper.)
|
||||||
|
- Owens, S., Sarkar, S., Sewell, P. *"A Better x86 Memory
|
||||||
|
Model: x86-TSO."* (TSO reference; the TSO model is the
|
||||||
|
inspiration for `Ztso`.)
|
||||||
|
- Lamport, L. *"How to Make a Multiprocessor Computer That
|
||||||
|
Correctly Executes Multiprocess Programs."* IEEE TC, 1979.
|
||||||
|
(The foundational sequential consistency paper.)
|
||||||
|
- Adve, S. V., Gharachorloo, K. *"Shared Memory Consistency
|
||||||
|
Models: A Tutorial."* IEEE Computer, 1996. (Tutorial on
|
||||||
|
memory models.)
|
||||||
|
- Martin, M. M. K., et al. *"Token Coherence."* (Token-based
|
||||||
|
ordering reference, applicable to Approach B.)
|
||||||
|
- ACM SIGOPS Operating Systems Review. *Linux Kernel Memory
|
||||||
|
Model Documentation.* (`tools/memory-model/` in the Linux
|
||||||
|
kernel source tree.)
|
||||||
|
- AMD. *AMD64 Architecture Programmer's Manual*, Volume 2.
|
||||||
|
(TSO reference for comparison.)
|
||||||
|
- Intel. *Intel 64 and IA-32 Architectures Software
|
||||||
|
Developer's Manual*, Volume 3. (TSO reference for
|
||||||
|
comparison.)
|
||||||
|
- ARM. *ARM Architecture Reference Manual*. (Comparison
|
||||||
|
with ARM's weak memory model.)
|
||||||
|
|
||||||
|
Specific tool references:
|
||||||
|
|
||||||
|
- **herd7**: litmus-test framework for weak memory models,
|
||||||
|
by Cambridge. https://github.com/herd/herdtools7
|
||||||
|
- **diy7**: companion tool for herd7.
|
||||||
|
- **Murphi**, **TLA+**: protocol model checkers used for
|
||||||
|
coherence protocol verification.
|
||||||
|
|
||||||
|
The RISC-V Memory Model working group's published
|
||||||
|
documentation, including the operational model papers and
|
||||||
|
the litmus-test corpus, is the primary source for the
|
||||||
|
RVWMO contract. Specifics of XH-1's coherence protocol are
|
||||||
|
out of scope for this document and are addressed in the
|
||||||
|
coherence research document.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 14. Document Status
|
||||||
|
|
||||||
|
This document is a **research proposal**. It does not represent
|
||||||
|
a finalized XH-1 architectural commitment. Items marked
|
||||||
|
**[Proposed]** are engineering recommendations under evaluation;
|
||||||
|
items marked **[Open]** are unresolved design questions.
|
||||||
|
|
||||||
|
The next revisions of this document should:
|
||||||
|
|
||||||
|
- Reflect decisions on the open questions in §11.
|
||||||
|
- Add quantitative `FENCE` latency measurements once RTL or
|
||||||
|
high-level models are available.
|
||||||
|
- Incorporate results from the verification litmus-test
|
||||||
|
suite.
|
||||||
|
- Align with the final coherence protocol choice.
|
||||||
|
- Document the OS-level `FENCE` insertion policy.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
*End of document.*
|
||||||
@@ -0,0 +1,6 @@
|
|||||||
|
{
|
||||||
|
"verdict": "PASS",
|
||||||
|
"confidence": "HIGH",
|
||||||
|
"issues": [],
|
||||||
|
"required_fixes": []
|
||||||
|
}
|
||||||
@@ -0,0 +1,15 @@
|
|||||||
|
The document is a well-structured and technically sound research proposal for the XH-1 architecture. It correctly identifies the primary challenges associated with scaling a RISC-V implementation to 128 cores, specifically regarding memory consistency, coherence, and the latency of synchronization primitives in a mesh interconnect.
|
||||||
|
|
||||||
|
### Technical Analysis:
|
||||||
|
|
||||||
|
1. **Memory Model (RVWMO vs. TSO):** The document correctly identifies the distinction between the standard RISC-V Weak Memory Ordering (RVWMO) and the Total Store Order (TSO) model. By opting for RVWMO as the base and acknowledging the `Ztso` extension, the proposal aligns with the RISC-V ecosystem while providing a path for TSO-specific optimizations if required.
|
||||||
|
2. **Scalability and "Fence-at-Distance":** One of the most significant challenges in large-scale mesh networks (like a 128-core system) is the latency of `FENCE` instructions. The proposal's inclusion of "Approach B" (Per-Core Token Completion) and "Approach C" (Directory-Assisted Ordering) along with a dedicated "fence-ack" message class demonstrates a sophisticated understanding of how to mitigate the "fence-at-distance" problem without over-complicating the standard coherence protocol.
|
||||||
|
3. **Coherence and Interconnect:** The choice of a directory-based MESI protocol at the L2 slice is the industry standard for high-core-count systems. This localizes the coherence traffic and provides a clear point of serialization. The recommendation to use a mesh topology is also appropriate for the target scale.
|
||||||
|
4. **Verification Strategy:** The inclusion of `herd7`, `diy7`, and formal tools like `Murphi` and `TLA+` indicates a rigorous approach to verification. These are the standard tools for ensuring that a complex memory model is correctly implemented in hardware.
|
||||||
|
5. **RISC-V Compliance:** The document correctly references the relevant parts of the RISC-V ISA manuals (RVWMO, PBMT, and standard fence semantics), ensuring that the hardware remains compliant
|
||||||
|
{
|
||||||
|
"verdict": "PASS",
|
||||||
|
"confidence": "HIGH",
|
||||||
|
"issues": [],
|
||||||
|
"required_fixes": []
|
||||||
|
}
|
||||||
@@ -0,0 +1,2 @@
|
|||||||
|
2026-08-26T12:41:33Z research/05-memory/memory-ordering.md 1 research success
|
||||||
|
2026-08-26T12:45:19Z research/05-memory/memory-ordering.md 1 review PASS
|
||||||
@@ -0,0 +1,920 @@
|
|||||||
|
```
|
||||||
|
# Memory Ordering
|
||||||
|
|
||||||
|
## Scope
|
||||||
|
|
||||||
|
This document investigates the memory consistency model of the XH-1
|
||||||
|
processor, a 128-core RISC-V design, and the hardware mechanisms required
|
||||||
|
to implement that model at scale. It covers the RISC-V "RVWMO" base
|
||||||
|
ordering rules, the Ztso and Fence extensions, the I/O and acquire/release
|
||||||
|
ordering primitives, the role of the coherence protocol, the implications
|
||||||
|
of 128 cores on ordering scalability, and the interactions with pipeline,
|
||||||
|
cache hierarchy, interconnect, interrupts, operating system, verification,
|
||||||
|
and performance.
|
||||||
|
|
||||||
|
The document is research material. Where the XH-1 design has not been
|
||||||
|
finalized, statements are marked explicitly as proposals, assumptions, or
|
||||||
|
open questions.
|
||||||
|
|
||||||
|
## 1. Background: What "Memory Ordering" Means
|
||||||
|
|
||||||
|
A memory consistency model defines the legal values that a load may return
|
||||||
|
as a function of the loads and stores issued by all cores and by I/O
|
||||||
|
devices. This is necessary because:
|
||||||
|
|
||||||
|
1. Modern cores issue memory operations out of program order.
|
||||||
|
2. Caches and store buffers cause additional reorderings.
|
||||||
|
3. Interconnects may deliver messages out of order relative to program
|
||||||
|
order.
|
||||||
|
4. Multiple coherence agents may service requests concurrently.
|
||||||
|
|
||||||
|
A weaker model allows more reordering, permits more aggressive
|
||||||
|
optimization, and exposes more burden to software. A stronger model
|
||||||
|
exposes less burden to software but constrains hardware and costs
|
||||||
|
performance. The RISC-V ISA explicitly chooses a comparatively weak
|
||||||
|
model, "RVWMO", documented in the RISC-V *Unprivileged* and
|
||||||
|
*Privileged* ISA specifications and the RISC-V Memory Model
|
||||||
|
documentation.
|
||||||
|
|
||||||
|
Sources:
|
||||||
|
|
||||||
|
- RISC-V International, *The RISC-V Instruction Set Manual, Volume I:
|
||||||
|
Unprivileged ISA*, Document Version 20240411 (or current).
|
||||||
|
- RISC-V International, *The RISC-V Instruction Set Manual, Volume II:
|
||||||
|
Privileged ISA*, current version.
|
||||||
|
- RISC-V International, *RISC-V Memory Model*, in the same manual
|
||||||
|
package.
|
||||||
|
- Adve and Gharachorloo, "Shared Memory Consistency Models: A Tutorial",
|
||||||
|
IEEE Computer, 1996.
|
||||||
|
- Sorin, Hill, and Wood, *A Primer on Memory Consistency and Cache
|
||||||
|
Coherence*, Synthesis Lectures on Computer Architecture, 2011.
|
||||||
|
|
||||||
|
## 2. RISC-V Architectural Requirements
|
||||||
|
|
||||||
|
### 2.1 The RVWMO Base Model
|
||||||
|
|
||||||
|
RISC-V defines a baseline memory model called **RVWMO** ("RISC-V
|
||||||
|
Weak Memory Ordering"). The normative definition consists of:
|
||||||
|
|
||||||
|
1. A set of **preserved program order (PPO)** relations between pairs
|
||||||
|
of memory operations issued by the same hart.
|
||||||
|
2. A set of **memory consistency axioms** in the style of "global
|
||||||
|
memory order" (GMO), which is a total order over all memory
|
||||||
|
operations that respects PPO and coherence.
|
||||||
|
3. An **axiomatic specification** of which loads may read which stores
|
||||||
|
under the model.
|
||||||
|
|
||||||
|
The PPO rules include (non-exhaustive):
|
||||||
|
|
||||||
|
- A store followed by a load to an overlapping address does not
|
||||||
|
reorder (Write→Read same address).
|
||||||
|
- Acquire annotations prevent preceding memory operations from being
|
||||||
|
reordered past subsequent loads and stores.
|
||||||
|
- Release annotations prevent subsequent memory operations from being
|
||||||
|
reordered past preceding loads and stores.
|
||||||
|
- Fence instructions impose explicit ordering constraints.
|
||||||
|
- A `FENCE` after an AMO prevents the AMO from being reordered with
|
||||||
|
subsequent loads and stores.
|
||||||
|
- Dependencies (`addr`, `ctrl`, `data`) on the address of a load
|
||||||
|
enforce ordering with respect to preceding stores.
|
||||||
|
|
||||||
|
RISC-V is an **architecturally** weak model. The specification permits
|
||||||
|
the implementation to reorder, in the absence of ordering constraints,
|
||||||
|
loads with respect to other loads, stores with respect to other stores,
|
||||||
|
loads with respect to preceding stores, and stores with respect to
|
||||||
|
subsequent loads. It does **not** require TSO (Total Store Order) by
|
||||||
|
default.
|
||||||
|
|
||||||
|
### 2.2 The Ztso Extension
|
||||||
|
|
||||||
|
RISC-V defines a separate **Ztso** extension that makes the memory
|
||||||
|
model **TSO-compatible**:
|
||||||
|
|
||||||
|
- Loads may not be reordered with respect to other loads.
|
||||||
|
- Stores may not be reordered with respect to other stores.
|
||||||
|
- Loads may be reordered with respect to earlier stores to a different
|
||||||
|
address (as in SPARC V9 TSO and x86-TSO).
|
||||||
|
- Stores may not be reordered with respect to earlier loads.
|
||||||
|
|
||||||
|
If a hart implements `Ztso`, `fence rw,rw` may be implemented as a no-op
|
||||||
|
and the memory consistency rules are simplified accordingly.
|
||||||
|
|
||||||
|
The Ztso extension is **optional**. RISC-V implementations are not
|
||||||
|
required to support TSO.
|
||||||
|
|
||||||
|
### 2.3 The Fence and Fence.i Instructions
|
||||||
|
|
||||||
|
The base `FENCE` instruction takes a four-bit predecessor set and a
|
||||||
|
four-bit successor set, encoded in `fm` and `predecessor/successor`
|
||||||
|
fields. The standard encoding is:
|
||||||
|
|
||||||
|
```
|
||||||
|
FENCE pred, succ
|
||||||
|
```
|
||||||
|
|
||||||
|
where `pred` and `succ` select among:
|
||||||
|
|
||||||
|
- `r` (device input and memory reads)
|
||||||
|
- `w` (device output and memory writes)
|
||||||
|
- `i` (instruction stream)
|
||||||
|
|
||||||
|
In the absence of the `Zihintpause` extension, the FENCE instruction
|
||||||
|
also serves as a PAUSE hint when `pred=0, succ=0` (the "fence.tso"
|
||||||
|
form, distinct from the Ztso extension).
|
||||||
|
|
||||||
|
`FENCE.I` orders writes to instruction memory with respect to
|
||||||
|
subsequent instruction fetches, and is required to invalidate the
|
||||||
|
local I-cache.
|
||||||
|
|
||||||
|
### 2.4 The Zaamo and Zalrsc Extensions
|
||||||
|
|
||||||
|
- **Zaamo** defines atomic memory operations (AMOs) such as
|
||||||
|
`AMOSWAP.W`, `AMOADD.W`, `AMOAND.W`, `AMOOR.W`, `AMOXOR.W`,
|
||||||
|
`AMOMAX[U].W`, `AMOMIN[U].W`, and their `.D` and `.Q` (if `XLEN` is
|
||||||
|
large enough) variants. AMOs are defined to be `aq`-ordered before
|
||||||
|
and `rl`-ordered after the atomic access by default in RVWMO.
|
||||||
|
- **Zalrsc** defines the `LR`/`SC` pair, with `.aq`/`.rl` ordering
|
||||||
|
bits.
|
||||||
|
|
||||||
|
In RVWMO, an AMO with default ordering has acquire semantics on its
|
||||||
|
predecessors and release semantics on its successors. Setting `.aq` or
|
||||||
|
`.rl` is redundant for AMOs but is permitted. For `LR`/`SC`, the `.aq`
|
||||||
|
and `.rl` bits **are** meaningful: `LR` with `.aq` imposes acquire
|
||||||
|
ordering; `LR` with `.rl` is reserved; `SC` with `.rl` imposes release
|
||||||
|
ordering; `SC` with `.aq` is reserved.
|
||||||
|
|
||||||
|
This asymmetry between AMO and LR/SC default ordering is a documented
|
||||||
|
quirk of the current spec; it is preserved for compatibility but
|
||||||
|
considered a likely future correction.
|
||||||
|
|
||||||
|
### 2.5 The Zicbom and Zicboz Extensions
|
||||||
|
|
||||||
|
These are not strictly part of the memory model but interact with it:
|
||||||
|
|
||||||
|
- **Zicbom** defines cache-block management instructions
|
||||||
|
(`CBO.INVAL`, `CBO.CLEAN`, `CBO.FLUSH`). Their ordering behavior
|
||||||
|
relative to surrounding loads and stores is specified by the
|
||||||
|
standard.
|
||||||
|
- **Zicboz** defines `CBO.ZERO`, used to zero a cache block with
|
||||||
|
defined ordering and without polluting caches with external data.
|
||||||
|
|
||||||
|
### 2.6 I/O Ordering: PMAs and the I/O Fence
|
||||||
|
|
||||||
|
Device memory, in the RISC-V model, is a matter of **physical memory
|
||||||
|
attributes (PMAs)**: each address range carries attributes describing
|
||||||
|
whether it supports non-idempotent writes, the supported access size and
|
||||||
|
atomicity, whether the device can tolerate speculative accesses, and
|
||||||
|
whether coherence agents must snoop or forward. The RISC-V model
|
||||||
|
requires an **I/O fence** (a FENCE with `pred=ow, succ=ow`, or a
|
||||||
|
stronger form) before a non-idempotent device write and after a
|
||||||
|
non-idempotent device read. Without such a fence the architectural
|
||||||
|
guarantees do not apply.
|
||||||
|
|
||||||
|
This is a significant departure from architectures such as x86, where
|
||||||
|
I/O ordering is guaranteed for in-order MMIO. RISC-V requires
|
||||||
|
**software** to insert fences before and after MMIO accesses to
|
||||||
|
non-idempotent devices. (Idempotent and well-behaved devices are
|
||||||
|
implicitly supported; the spec describes the boundary precisely.)
|
||||||
|
|
||||||
|
### 2.7 Acquire and Release Annotations
|
||||||
|
|
||||||
|
Load-acquire (`LR.W.aq` / `LR.D.aq` / equivalent AMOs) and
|
||||||
|
store-release (`SC.W.rl` / `SC.D.rl` / equivalent AMOs and the plain
|
||||||
|
`AMOSWAP` with `rl` bit) are first-class citizens of the RISC-V model.
|
||||||
|
They are the preferred synchronization primitives, since they are
|
||||||
|
narrower in scope than a full fence and can be optimized by hardware.
|
||||||
|
|
||||||
|
RV32A and RV64A define a small set of AMO operations. The `.aq` and
|
||||||
|
`.rl` bits are 1-bit fields. The Zalasr extension, when present,
|
||||||
|
generalizes acquire/release to ordinary loads and stores (e.g.
|
||||||
|
`LD.aq` / `ST.rl`), as a proposed but not yet ratified extension at
|
||||||
|
the time of writing. **Status of Zalasr should be confirmed against
|
||||||
|
the current RISC-V International extension list before adoption.**
|
||||||
|
|
||||||
|
## 3. Implications for a 128-Core RISC-V Processor
|
||||||
|
|
||||||
|
### 3.1 Hardware Required to Implement RVWMO
|
||||||
|
|
||||||
|
Implementing the RVWMO model correctly in hardware requires that the
|
||||||
|
**observable behavior** of memory operations match the axiomatic
|
||||||
|
definition. This is normally achieved by combining:
|
||||||
|
|
||||||
|
1. A **store buffer** (often called a write-combining buffer or
|
||||||
|
write queue) that holds pending stores in program order and may
|
||||||
|
commit them to the coherent subsystem out of order only as
|
||||||
|
permitted by the PPO rules.
|
||||||
|
2. A **load queue** (often called a memory dependence speculation
|
||||||
|
table or load reorder buffer) that tracks in-flight loads so that
|
||||||
|
a violating load can be replayed or, in some designs, so that
|
||||||
|
forwarding from the store buffer can be checked.
|
||||||
|
3. A **coherence protocol** that defines the per-cache-block
|
||||||
|
ordering of writes and provides a single-writer (modified) state
|
||||||
|
for exclusive ownership.
|
||||||
|
4. An **interconnect** (e.g., a NoC) that provides point-to-point
|
||||||
|
ordering on each channel, and a system-level framework for
|
||||||
|
ensuring that the global memory order is consistent with PPO at
|
||||||
|
every observer.
|
||||||
|
5. **Fence handling**: a mechanism to drain the store buffer and to
|
||||||
|
block the issue of subsequent loads until prior stores complete,
|
||||||
|
to the extent required by the FENCE operand set.
|
||||||
|
|
||||||
|
The classical TSO implementation, as in the original SPARC and as
|
||||||
|
described in the SPARC V9 manual and in Sorin/Hill/Wood, is the
|
||||||
|
reference point. RVWMO requires additional handling for the
|
||||||
|
non-TSO reorderings, particularly load-load and store-store
|
||||||
|
reordering across different addresses.
|
||||||
|
|
||||||
|
### 3.2 Per-Hart Pipeline State
|
||||||
|
|
||||||
|
For each hart, XH-1 (proposed) would include:
|
||||||
|
|
||||||
|
- A **store buffer** with entries holding address, data, size, and
|
||||||
|
the originating hart's local program order. Entries are dispatched
|
||||||
|
to the coherent subsystem when downstream permits (e.g., when a
|
||||||
|
cache line is granted in Modified state, or when the buffer is
|
||||||
|
drained by a fence).
|
||||||
|
- A **load queue** that records in-flight loads and their
|
||||||
|
dependencies. A load may forward from the store buffer only if the
|
||||||
|
PPO rules permit (in particular, must check the "load must not
|
||||||
|
reorder with respect to a prior store to the same or overlapping
|
||||||
|
address" rule and must consider the various "load may not bypass
|
||||||
|
fence" rules).
|
||||||
|
- A **dependency-tracking mechanism** for the PPO rules that depend
|
||||||
|
on `addr`, `ctrl`, and `data` dependencies, since these are
|
||||||
|
significant for FENCE relaxation around address-dependent loads.
|
||||||
|
RVWMO allows a FENCE to be elided between a store and a subsequent
|
||||||
|
load whose address depends on the store, which is a key
|
||||||
|
optimization.
|
||||||
|
|
||||||
|
A typical implementation tracks addresses in the store buffer to
|
||||||
|
determine when a FENCE can be retired (i.e., when the buffer has
|
||||||
|
drained) and when a subsequent load can bypass a prior store. For
|
||||||
|
TSO (Ztso), the design simplifies: a FENCE `rw,rw` is a no-op, and
|
||||||
|
the store buffer enforces program order among stores automatically.
|
||||||
|
|
||||||
|
### 3.3 Coherence Protocol and Memory Ordering
|
||||||
|
|
||||||
|
The coherence protocol provides the underlying **single-writer,
|
||||||
|
multiple-reader (SWMR)** invariant for each cache line, which is a
|
||||||
|
prerequisite for any memory model. The RISC-V coherence extension
|
||||||
|
(Zicsr aside) is defined by the privileged manual:
|
||||||
|
|
||||||
|
- A directory-based or broadcast-based MESI / MOESI family is
|
||||||
|
expected. The directory is the typical choice for 128 cores to
|
||||||
|
avoid broadcast storms.
|
||||||
|
- The **directory** must record which caches hold a line in which
|
||||||
|
state. For a 128-core system, a full-map directory is 128 bits
|
||||||
|
per line plus a state field, which is acceptable for L2 or L3 tag
|
||||||
|
storage.
|
||||||
|
|
||||||
|
Memory ordering interacts with coherence at several points:
|
||||||
|
|
||||||
|
- **Write atomicity**: the system must ensure that a store to
|
||||||
|
address A becomes visible to all other harts **atomically**, in
|
||||||
|
the sense that no other hart can observe a value older than the
|
||||||
|
store and a value newer than the store at the same address
|
||||||
|
simultaneously. Directory invalidations and acknowledgments
|
||||||
|
enforce this.
|
||||||
|
- **Invalidation acknowledgment**: an acquire fence or load-acquire
|
||||||
|
must wait for invalidations of the same address to be
|
||||||
|
acknowledged before issuing a subsequent load to that address.
|
||||||
|
- **Forwarding from store buffer**: in a directory system, a
|
||||||
|
subsequent load that hits in the L1 may also need to be ordered
|
||||||
|
with respect to coherence invalidations in flight. A typical
|
||||||
|
implementation uses a **store buffer forwarding path** that
|
||||||
|
snoops incoming invalidations and invalidates matching store
|
||||||
|
buffer entries.
|
||||||
|
|
||||||
|
The protocol chosen (MESI vs. MOESI vs. MESIF; directory vs.
|
||||||
|
broadcast) is a design choice independent of the memory model but
|
||||||
|
strongly affects how the model is enforced. Directory-based MESI
|
||||||
|
is the conventional choice for a 128-core machine.
|
||||||
|
|
||||||
|
### 3.4 Interconnect and Ordering
|
||||||
|
|
||||||
|
The on-chip interconnect (NoC) is the medium through which coherence
|
||||||
|
messages travel. For the memory model to be implementable, the NoC
|
||||||
|
must provide:
|
||||||
|
|
||||||
|
1. **Per-channel ordering**: for each pair of source–destination
|
||||||
|
endpoints, the NoC must deliver messages in the order in which
|
||||||
|
they are injected. Most on-chip NoCs (e.g., a 2D mesh with
|
||||||
|
virtual-channel routers) provide this at the link layer for
|
||||||
|
within-VC traffic.
|
||||||
|
2. **In-order delivery of coherence messages for the same address**:
|
||||||
|
if two messages for the same address travel through different
|
||||||
|
paths, the protocol must include a sequence number or
|
||||||
|
acknowledgment so that they are processed in order. Directory
|
||||||
|
controllers serialize responses per address by design.
|
||||||
|
3. **FIFO ordering between specific message classes**: in
|
||||||
|
particular, a data response that delivers a value to a requester
|
||||||
|
must be ordered with respect to invalidations flowing to the
|
||||||
|
same requester. A typical implementation uses a per-request
|
||||||
|
completion queue.
|
||||||
|
|
||||||
|
For a 128-core XH-1 (proposed), a **2D mesh NoC** with XY routing,
|
||||||
|
virtual channels, and per-router flow control is a reasonable
|
||||||
|
baseline. The NoC must be designed so that an acquire fence can
|
||||||
|
wait for outstanding invalidations on all channels.
|
||||||
|
|
||||||
|
### 3.5 Realistic Implementation Approaches for XH-1
|
||||||
|
|
||||||
|
The following are realistic approaches for an XH-1 implementation.
|
||||||
|
Each is presented as a proposal; the final choice is open.
|
||||||
|
|
||||||
|
**Approach A: Strict RVWMO with per-hart store buffer + load
|
||||||
|
queue + directory MESI.**
|
||||||
|
|
||||||
|
- Each hart has a 16–32 entry store buffer and a 32–64 entry load
|
||||||
|
queue.
|
||||||
|
- Stores are dispatched to the directory when the line is in
|
||||||
|
Modified state in the L1. The store buffer coalesces
|
||||||
|
same-address stores.
|
||||||
|
- Loads can forward from the store buffer under the PPO rules
|
||||||
|
(same-address, no dependence violation).
|
||||||
|
- FENCE `w,w` waits for the store buffer to drain at the
|
||||||
|
interconnect boundary. FENCE `r,r` waits for the load queue to
|
||||||
|
drain.
|
||||||
|
- Directory is a 128-bit sharer vector per line plus state.
|
||||||
|
|
||||||
|
This is the classical RVWMO implementation, well understood from
|
||||||
|
academic and commercial practice (e.g., the original RISC-V
|
||||||
|
research chips, MIPS R10000, ARM Cortex-A series, recent
|
||||||
|
high-performance x86 cores).
|
||||||
|
|
||||||
|
**Approach B: TSO (Ztso) implementation.**
|
||||||
|
|
||||||
|
- Store buffer enforces program order among stores. Loads can
|
||||||
|
bypass the store buffer (per TSO rules).
|
||||||
|
- FENCE `rw,rw` is a no-op.
|
||||||
|
- Performance penalty relative to Approach A is small for most
|
||||||
|
workloads; benefit is greatly simplified verification and
|
||||||
|
software compatibility with code written for x86 / SPARC.
|
||||||
|
|
||||||
|
This may be attractive if the XH-1 project targets compatibility
|
||||||
|
with existing operating systems or synchronization libraries
|
||||||
|
that expect TSO. However, adopting Ztso as a **system-wide**
|
||||||
|
property is awkward on RISC-V, since the ISA does not require it
|
||||||
|
on every hart and binaries may mix RVWMO and Ztso harts. The
|
||||||
|
recommended approach is to support both via a **Zicfiss / Ztso
|
||||||
|
feature CSR** that allows each hart to opt into TSO. (Note:
|
||||||
|
Zicfiss is an extension for control-flow integrity, not
|
||||||
|
ordering; a different CSR-based opt-in mechanism would be needed;
|
||||||
|
status of CSR-based TSO opt-in is **open** and should be
|
||||||
|
confirmed against the RISC-V Privileged specification.)
|
||||||
|
|
||||||
|
**Approach C: A hybrid, with RCpc / RCsc style acquire/release and
|
||||||
|
a lightweight fence.**
|
||||||
|
|
||||||
|
- Implement the standard RVWMO.
|
||||||
|
- Use the standard AMO.acquire / AMO.release where possible.
|
||||||
|
- Implement FENCE as the lightweight `FENCE.TSO` (`fm=0, pred=0,
|
||||||
|
succ=0`) when the operand set permits, which is a hint to drain
|
||||||
|
the store buffer but allows it to be implemented as a full fence.
|
||||||
|
- Treat all device memory accesses as requiring the I/O fence
|
||||||
|
pattern (FENCE `ow,ow` before, FENCE `iorw,iorw` after).
|
||||||
|
|
||||||
|
This is a typical approach in commercial RISC-V cores. The
|
||||||
|
performance is essentially the same as Approach A but the design
|
||||||
|
intent is cleaner.
|
||||||
|
|
||||||
|
**Recommendation (tentative):** Approach A or C, with FENCE
|
||||||
|
implemented strictly, AMO.acquire/release for most
|
||||||
|
synchronization, and Ztso support as an optional feature. Final
|
||||||
|
choice depends on the operating system port, the performance
|
||||||
|
budget, and the verification strategy.
|
||||||
|
|
||||||
|
## 4. The 128-Core Scaling Problem
|
||||||
|
|
||||||
|
The memory model and its hardware implementation are sensitive to
|
||||||
|
the number of cores. The 128-core XH-1 introduces several
|
||||||
|
scalability concerns:
|
||||||
|
|
||||||
|
### 4.1 Fence Scalability
|
||||||
|
|
||||||
|
A "fence" in the global sense — a FENCE that requires all
|
||||||
|
prior memory operations to be globally visible before the fence
|
||||||
|
is retired — is a **global barrier**. As the number of cores
|
||||||
|
grows, the latency of a global fence grows with the diameter of
|
||||||
|
the coherence system.
|
||||||
|
|
||||||
|
In a directory-MESI system, a FENCE `w,w` must wait for the store
|
||||||
|
buffer to drain and for each outstanding store to be acknowledged
|
||||||
|
by the directory and, if necessary, by all remote sharers. The
|
||||||
|
worst-case latency is bounded by the longest in-flight store,
|
||||||
|
which is typically a few hundred cycles for an L2 miss.
|
||||||
|
|
||||||
|
For a 128-core machine, a **FENCE is local by default** in
|
||||||
|
RVWMO: it orders only with respect to the issuing hart. Software
|
||||||
|
that needs a global barrier must use a `FENCE` followed by an
|
||||||
|
`AMO` or a custom mechanism (e.g., an IPI). This is part of the
|
||||||
|
RVWMO design: there is no architectural "sync" or "membar"
|
||||||
|
instruction with global scope; software must compose one.
|
||||||
|
|
||||||
|
Open question: Should XH-1 provide a custom global fence
|
||||||
|
instruction via a custom opcode (non-standard) to accelerate
|
||||||
|
operating system boot and shutdown? This is a real engineering
|
||||||
|
trade-off and the answer is not yet decided.
|
||||||
|
|
||||||
|
### 4.2 Directory Storage and Bandwidth
|
||||||
|
|
||||||
|
A full-map directory of 128 sharers requires 128 bits per cache
|
||||||
|
line, plus a state field. For an L3 of 64 MiB at 64 B line size
|
||||||
|
(1M lines), this is 16 MiB of directory state — feasible but
|
||||||
|
significant. For 128-byte lines it is 2 MiB — much cheaper but
|
||||||
|
worse for spatial locality and false sharing.
|
||||||
|
|
||||||
|
Coalescing the directory (e.g., a 16-entry coarse vector) saves
|
||||||
|
space at the cost of extra invalidation traffic. The trade-off
|
||||||
|
is well known; see Cuesta et al., "Increasing the Effectiveness
|
||||||
|
of Directory Caches by Deactivating Coarse-Grained Coherence
|
||||||
|
Tracking", ICS 2011.
|
||||||
|
|
||||||
|
**Proposal for XH-1 (open):** a coarse-vector directory with 8
|
||||||
|
or 16 pointers per entry, plus a broadcast fallback when the
|
||||||
|
vector overflows. This is conventional in commercial designs of
|
||||||
|
this scale.
|
||||||
|
|
||||||
|
### 4.3 Interconnect Contention
|
||||||
|
|
||||||
|
At 128 cores, the on-chip NoC must carry invalidations, data
|
||||||
|
responses, and acknowledgments between every pair of cores. A
|
||||||
|
2D mesh with bisection bandwidth of B bytes/cycle/hop becomes the
|
||||||
|
bottleneck. The memory model's influence is mainly through
|
||||||
|
**invalidation storms**: when one core writes a shared line, all
|
||||||
|
128 sharers (or all 128 caches, for a miss) must be invalidated
|
||||||
|
or notified. The protocol must bound the worst-case
|
||||||
|
invalidation fanout.
|
||||||
|
|
||||||
|
Standard mitigations:
|
||||||
|
|
||||||
|
- **Region coherence / hierarchical directory** (L2 directories,
|
||||||
|
L3 directory).
|
||||||
|
- **Token coherence** (Martin, Hill, Sorin; HPCA 2003), which
|
||||||
|
replaces directory entries with tokens and bounds
|
||||||
|
invalidation traffic at the cost of extra protocol messages.
|
||||||
|
- **In-network coherence** (e.g., the Intel Mesh / Emesh design),
|
||||||
|
which performs directory lookups inside the NoC routers.
|
||||||
|
|
||||||
|
These are all reasonable for a research project but introduce
|
||||||
|
complexity. A baseline MESI directory is the recommended
|
||||||
|
starting point.
|
||||||
|
|
||||||
|
### 4.4 False Sharing and Ordering
|
||||||
|
|
||||||
|
A 128-core machine is particularly sensitive to **false
|
||||||
|
sharing**, where two cores write to different words in the same
|
||||||
|
cache line. Each such write triggers a coherence ping-pong.
|
||||||
|
|
||||||
|
The memory model does not by itself mitigate this — the cache
|
||||||
|
line is the unit of coherence, and any store to it forces a
|
||||||
|
write. XH-1 (proposed) may include hardware support for
|
||||||
|
detecting hot lines, but the primary mitigation is software
|
||||||
|
(cache-line padding, avoiding shared writable data).
|
||||||
|
|
||||||
|
The memory model is relevant for one specific case: a load
|
||||||
|
following a store to the same line in program order must not
|
||||||
|
reorder past the store. In a directory system, the store must
|
||||||
|
be acknowledged before the load can complete. This is a
|
||||||
|
"store-load forwarding" case and is a major source of pipeline
|
||||||
|
stalls.
|
||||||
|
|
||||||
|
### 4.5 Acquire/Release Scalability
|
||||||
|
|
||||||
|
Load-acquire and store-release are the preferred RVWMO
|
||||||
|
synchronization primitives. They are local: they order only
|
||||||
|
with respect to the issuing hart. The hardware cost is a small
|
||||||
|
amount of state in the load/store units to delay the issue of
|
||||||
|
subsequent operations until the acquire is acknowledged.
|
||||||
|
|
||||||
|
At 128 cores, the **aggregate** acquire/release traffic (from
|
||||||
|
synchronization in the OS, runtime, and applications) is
|
||||||
|
significant but not fundamentally different from smaller
|
||||||
|
machines. The scalability limit is the protocol, not the
|
||||||
|
model.
|
||||||
|
|
||||||
|
## 5. Interactions
|
||||||
|
|
||||||
|
### 5.1 Pipeline
|
||||||
|
|
||||||
|
The store buffer and load queue are pipeline structures. The
|
||||||
|
pipeline depth determines how many in-flight loads/stores can
|
||||||
|
coexist. For a high-frequency XH-1 core, the store buffer should
|
||||||
|
hold at least as many entries as the load-to-use latency times
|
||||||
|
the issue rate — typically 16–32 entries for a 4-wide issue.
|
||||||
|
|
||||||
|
The FENCE instruction is a **pipeline-flushing event** in the
|
||||||
|
broadest sense: it must drain the store buffer and/or load queue
|
||||||
|
to the extent required. A FENCE `rw,rw` is the most expensive
|
||||||
|
and may stall the issue queue for tens of cycles. The
|
||||||
|
implementation can optimize by allowing subsequent
|
||||||
|
non-dependent operations to issue (out-of-order) but must
|
||||||
|
constrain loads to a coherent subset until the fence retires.
|
||||||
|
|
||||||
|
### 5.2 Cache Hierarchy
|
||||||
|
|
||||||
|
The cache hierarchy is the agent that enforces the model's
|
||||||
|
single-writer invariant. The L1 must be **write-through** to the
|
||||||
|
L2 or **write-back** with a coherence protocol; in modern
|
||||||
|
designs, write-back with a directory is standard.
|
||||||
|
|
||||||
|
The store buffer sits **between the L1 and the L2**: stores
|
||||||
|
commit to the L1 in program order but the L1 marks the line
|
||||||
|
Modified without necessarily writing it to the L2. The
|
||||||
|
coherence protocol performs the L2 writeback when the line is
|
||||||
|
evicted or when another core requests the line.
|
||||||
|
|
||||||
|
For RVWMO, the L1 store buffer must:
|
||||||
|
|
||||||
|
- Coalesce same-line stores.
|
||||||
|
- Hold the address until the store is acknowledged by the L2
|
||||||
|
(or further).
|
||||||
|
- Service forwarding requests from the load unit.
|
||||||
|
|
||||||
|
For Ztso, the same structure is required, with the
|
||||||
|
simplification that FENCE is a no-op.
|
||||||
|
|
||||||
|
### 5.3 Memory System
|
||||||
|
|
||||||
|
Main memory (DRAM) introduces a new variable: the **DRAM
|
||||||
|
controller** may reorder requests to the same rank/bank to
|
||||||
|
improve bandwidth. This reordering is observable to the cores
|
||||||
|
unless the controller enforces ordering.
|
||||||
|
|
||||||
|
The RISC-V memory model requires that **the system behaves as
|
||||||
|
if there is a single global memory order**. A DRAM controller
|
||||||
|
that reorders must do so in a way that is invisible to the
|
||||||
|
model, which usually means serializing conflicting requests
|
||||||
|
(e.g., reads to the same bank after a write must wait).
|
||||||
|
|
||||||
|
Standard practice: a **write queue** in the memory controller
|
||||||
|
that tracks in-flight requests and serializes conflicting ones.
|
||||||
|
For XH-1 (proposed), the memory controller would include a
|
||||||
|
per-bank ordering queue.
|
||||||
|
|
||||||
|
### 5.4 Interconnect
|
||||||
|
|
||||||
|
The on-chip interconnect (NoC) is the medium through which
|
||||||
|
coherence flows. As discussed in §3.4, the NoC must provide
|
||||||
|
in-order delivery per channel and per-address serialization
|
||||||
|
across channels. A 2D mesh with virtual channels is the
|
||||||
|
conventional choice. For 128 cores, a 16x8 or 8x16 mesh
|
||||||
|
balances wire length and bisection bandwidth.
|
||||||
|
|
||||||
|
The memory model's FENCE instruction interacts with the NoC
|
||||||
|
via **end-to-end acknowledgment**: a FENCE that requires global
|
||||||
|
ordering cannot retire until all messages it has issued have
|
||||||
|
been acknowledged by their destinations. In a directory
|
||||||
|
system, this is a per-message ack from the directory and, for
|
||||||
|
shared lines, from the invalidation recipients.
|
||||||
|
|
||||||
|
### 5.5 Coherence
|
||||||
|
|
||||||
|
The coherence protocol is the **enforcement mechanism** for
|
||||||
|
write atomicity. A memory model without coherence is undefined
|
||||||
|
(the "coherence" axiom in RVWMO and other models assumes that
|
||||||
|
each location has a single most-recent store at any time).
|
||||||
|
|
||||||
|
For XH-1, the recommended baseline is directory-based MESI.
|
||||||
|
Variants such as MOESI (for data forwarding) or MESIF (for
|
||||||
|
forwarding the response) are reasonable. The choice of variant
|
||||||
|
does not change the memory model but affects performance and
|
||||||
|
verification.
|
||||||
|
|
||||||
|
A specific concern for RVWMO: a **store that is followed by a
|
||||||
|
load to the same line** is a critical case. The store must be
|
||||||
|
visible to the load, which in a directory system means the
|
||||||
|
store must complete (be acknowledged) before the load
|
||||||
|
returns. This is enforced by the L1 store buffer: a load
|
||||||
|
misses the L1, the directory is consulted, and the request is
|
||||||
|
serviced from the L1's Modified state (the store) or from
|
||||||
|
another cache (if the L1 already wrote the line to L2).
|
||||||
|
|
||||||
|
### 5.6 Interrupts and Traps
|
||||||
|
|
||||||
|
The RISC-V memory model requires that **interrupts and traps
|
||||||
|
see a consistent view of memory**. Specifically:
|
||||||
|
|
||||||
|
- A load that traps (e.g., page fault) must see the
|
||||||
|
architectural state as if the load had not occurred (or, for
|
||||||
|
certain traps, as if the load had occurred but no later
|
||||||
|
operation had).
|
||||||
|
- The model must specify whether the load is "performed" with
|
||||||
|
respect to other harts at the time of the trap.
|
||||||
|
|
||||||
|
In RVWMO, an unsuccessful load (one that traps) is not
|
||||||
|
considered to have been performed. A successful load is
|
||||||
|
considered performed in program order relative to other loads
|
||||||
|
and stores of the same hart, under the PPO rules. The
|
||||||
|
architectural state at the time of the trap is the state
|
||||||
|
before the load (and before any subsequent operation in
|
||||||
|
program order).
|
||||||
|
|
||||||
|
For XH-1, this is implemented by the pipeline: an
|
||||||
|
exception-causing load does not write back to a register, and
|
||||||
|
its store buffer / load queue entries are marked invalid.
|
||||||
|
Subsequent instructions in program order are squashed.
|
||||||
|
|
||||||
|
A **subtle case** is the I/O fence: a FENCE before a non-
|
||||||
|
idempotent device write must be retired before the write is
|
||||||
|
issued, even if an interrupt arrives between them. RVWMO
|
||||||
|
permits the interrupt to be taken at any time but requires
|
||||||
|
the FENCE to be retired before the write is observable to the
|
||||||
|
device.
|
||||||
|
|
||||||
|
### 5.7 Operating System
|
||||||
|
|
||||||
|
The operating system relies on the memory model for
|
||||||
|
synchronization:
|
||||||
|
|
||||||
|
- **Spinlocks**: typically implemented as `LR.W.aq` /
|
||||||
|
`SC.W.rl` or as `AMOSWAP.W.aqrl`. Both are acquire/release
|
||||||
|
and are well supported by RVWMO.
|
||||||
|
- **I/O fences**: required before/after device accesses.
|
||||||
|
Linux/RISC-V includes these in its I/O macros.
|
||||||
|
- **Page table updates**: require a release fence before
|
||||||
|
updating the PTE and an acquire fence (or
|
||||||
|
load-acquire) after, so that the TLB refill sees the
|
||||||
|
updated PTE.
|
||||||
|
- **Inter-processor interrupts (IPIs)**: require a release
|
||||||
|
fence before sending the IPI and an acquire fence after
|
||||||
|
receiving it, so that the data and the IPI are ordered
|
||||||
|
correctly.
|
||||||
|
- **Boot and shutdown**: require global fences, which on
|
||||||
|
RISC-V must be composed from local FENCE + AMO + IPI
|
||||||
|
patterns. This is more complex than on x86/TSO.
|
||||||
|
|
||||||
|
The OS for XH-1 (proposed) is likely to be Linux/RISC-V or a
|
||||||
|
research kernel. Either way, the memory model requirements
|
||||||
|
are well documented in the kernel's memory-barrier
|
||||||
|
documentation, and XH-1's implementation must support them.
|
||||||
|
|
||||||
|
### 5.8 Verification
|
||||||
|
|
||||||
|
Verifying the memory model is one of the hardest parts of
|
||||||
|
modern CPU design. The standard approach is:
|
||||||
|
|
||||||
|
1. **Litmus tests**: small programs designed to expose
|
||||||
|
reordering bugs. The RISC-V memory model has a
|
||||||
|
comprehensive test suite (see
|
||||||
|
https://github.com/litmus-tests/litmus-tests-riscv).
|
||||||
|
2. **Model checking**: tools such as Herd7, diy7, and
|
||||||
|
MemAlloy can check the axiomatic model against
|
||||||
|
implementation behaviors.
|
||||||
|
3. **Random instruction stream testing (RIST)**: at the
|
||||||
|
RTL level, run randomized tests and check the
|
||||||
|
observable orderings against the model.
|
||||||
|
|
||||||
|
For 128 cores, the verification surface grows combinatorially
|
||||||
|
in the number of cores. Standard practice is to verify the
|
||||||
|
single-core ordering first, then verify the coherence
|
||||||
|
protocol (which involves multiple cores), and finally
|
||||||
|
verify the cross-core ordering at the SoC level using
|
||||||
|
co-simulation or formal methods.
|
||||||
|
|
||||||
|
XH-1 (proposed) should adopt a multi-level verification
|
||||||
|
strategy: litmus tests for unit-level, formal verification
|
||||||
|
for the coherence protocol (using a tool such as CCSS or
|
||||||
|
Murphi), and full-system testing for the SoC. The
|
||||||
|
verification of the memory model is a **major
|
||||||
|
undertaking** and is often the gating item for tape-out.
|
||||||
|
|
||||||
|
### 5.9 Performance
|
||||||
|
|
||||||
|
The memory model's impact on performance is workload-
|
||||||
|
dependent:
|
||||||
|
|
||||||
|
- **Sequential programs** (single hart): the cost is the
|
||||||
|
store-load forwarding latency (a few cycles) and the
|
||||||
|
cost of fences (which are usually rare).
|
||||||
|
- **Parallel programs**: the cost is the latency of
|
||||||
|
acquire/release operations, which is typically a few
|
||||||
|
cycles for an L1 hit and tens of cycles for a remote
|
||||||
|
cache hit.
|
||||||
|
- **Highly contended workloads**: the cost is the
|
||||||
|
latency of a failed `SC` (which may require a
|
||||||
|
pipeline flush on some implementations) and the
|
||||||
|
cost of cache-line ping-pong.
|
||||||
|
|
||||||
|
The RVWMO model is generally **faster** than TSO for
|
||||||
|
parallel workloads because it allows more reordering
|
||||||
|
between loads and between stores. The Ztso extension is
|
||||||
|
provided for software that wants TSO semantics.
|
||||||
|
|
||||||
|
For a 128-core XH-1, the dominant performance concern is
|
||||||
|
not the memory model but the **coherence traffic and
|
||||||
|
interconnect bandwidth**. The memory model is a
|
||||||
|
secondary concern, but a poorly implemented fence can
|
||||||
|
serialize the entire machine (e.g., a global fence in
|
||||||
|
the OS boot path).
|
||||||
|
|
||||||
|
## 6. Alternatives and Trade-offs
|
||||||
|
|
||||||
|
### 6.1 TSO (Ztso) vs. RVWMO
|
||||||
|
|
||||||
|
- **TSO advantages**: simpler verification, easier OS
|
||||||
|
port, smaller fence frequency.
|
||||||
|
- **TSO disadvantages**: slightly more conservative
|
||||||
|
reordering, which costs performance on workloads
|
||||||
|
with relaxed dependencies.
|
||||||
|
|
||||||
|
XH-1 (proposed) is recommended to support both: RVWMO
|
||||||
|
as the default, with Ztso as an optional hart-level
|
||||||
|
feature. (As above, the CSR-based opt-in mechanism is
|
||||||
|
**open**.)
|
||||||
|
|
||||||
|
### 6.2 Directory MESI vs. Token Coherence
|
||||||
|
|
||||||
|
- **MESI directory advantages**: well understood, simple
|
||||||
|
protocol, predictable latency.
|
||||||
|
- **MESI directory disadvantages**: directory storage
|
||||||
|
cost (16 MiB for 64 MiB L3, 128-bit sharers),
|
||||||
|
invalidation storms for shared lines.
|
||||||
|
- **Token coherence advantages**: bounded traffic,
|
||||||
|
no directory storage.
|
||||||
|
- **Token coherence disadvantages**: more complex
|
||||||
|
protocol, more messages per coherence event,
|
||||||
|
less predictable latency.
|
||||||
|
|
||||||
|
For a research project, MESI is the recommended baseline.
|
||||||
|
Token coherence is a more advanced option.
|
||||||
|
|
||||||
|
### 6.3 Coarse-Vector vs. Full-Map Directory
|
||||||
|
|
||||||
|
- **Coarse-vector advantages**: less storage, similar
|
||||||
|
performance on most workloads.
|
||||||
|
- **Coarse-vector disadvantages**: occasional broadcast
|
||||||
|
fallback for over-subscribed lines.
|
||||||
|
- **Full-map advantages**: no broadcast fallback.
|
||||||
|
- **Full-map disadvantages**: 16 MiB of directory state
|
||||||
|
for 64 MiB L3.
|
||||||
|
|
||||||
|
For XH-1 (proposed), a coarse vector with 16 pointers per
|
||||||
|
entry plus broadcast fallback is the recommended baseline.
|
||||||
|
|
||||||
|
### 6.4 In-Network Coherence vs. Traditional Coherence
|
||||||
|
|
||||||
|
- **In-network advantages**: lower latency, lower
|
||||||
|
directory storage.
|
||||||
|
- **In-network disadvantages**: requires custom routers,
|
||||||
|
harder verification, less mature tooling.
|
||||||
|
|
||||||
|
This is a research-level option and not recommended for
|
||||||
|
the first version of XH-1.
|
||||||
|
|
||||||
|
### 6.5 Fence Implementation: Strict vs. Optimistic
|
||||||
|
|
||||||
|
- **Strict**: FENCE always waits for the full
|
||||||
|
predecessor set to drain at the system level.
|
||||||
|
- **Optimistic**: FENCE may retire as soon as the
|
||||||
|
predecessor set is locally drained, and the system
|
||||||
|
guarantees that the effect of the FENCE is preserved
|
||||||
|
even if the global drain has not completed.
|
||||||
|
|
||||||
|
RVWMO does not require a global drain for FENCE; the
|
||||||
|
hardware may retire the FENCE as soon as the local
|
||||||
|
guarantees are met. The OS and synchronization libraries
|
||||||
|
that need global ordering must compose a global fence
|
||||||
|
explicitly. XH-1 (proposed) is recommended to follow
|
||||||
|
this approach.
|
||||||
|
|
||||||
|
## 7. Advantages and Disadvantages of the Proposed Approach
|
||||||
|
|
||||||
|
### Advantages
|
||||||
|
|
||||||
|
- **Standards compliance**: implements RVWMO per the
|
||||||
|
RISC-V specification.
|
||||||
|
- **Familiar**: uses conventional MESI directory, store
|
||||||
|
buffer, load queue structures.
|
||||||
|
- **Scalable**: coarse-vector directory bounds storage,
|
||||||
|
2D mesh NoC provides sufficient bandwidth.
|
||||||
|
- **Verifiable**: litmus tests, formal verification of
|
||||||
|
the protocol, full-system co-simulation.
|
||||||
|
- **OS friendly**: supports Linux/RISC-V's memory-
|
||||||
|
barrier requirements.
|
||||||
|
|
||||||
|
### Disadvantages
|
||||||
|
|
||||||
|
- **Verification cost**: 128 cores × many memory model
|
||||||
|
rules is a large verification surface.
|
||||||
|
- **Interconnect complexity**: 2D mesh with virtual
|
||||||
|
channels, per-router flow control, per-channel
|
||||||
|
ordering.
|
||||||
|
- **Directory cost**: 16 MiB of directory state for a
|
||||||
|
64 MiB L3.
|
||||||
|
- **Fence overhead**: a global fence is more
|
||||||
|
expensive than on smaller machines.
|
||||||
|
- **OS porting burden**: global fences must be
|
||||||
|
composed from local FENCE + AMO patterns.
|
||||||
|
|
||||||
|
## 8. Unresolved Design Questions
|
||||||
|
|
||||||
|
1. **Zalasr status**: the Zalasr extension is **not yet
|
||||||
|
ratified** at the time of writing. XH-1 should not
|
||||||
|
rely on it until the extension is finalized.
|
||||||
|
2. **Ztso opt-in mechanism**: should XH-1 provide a
|
||||||
|
CSR-based opt-in for TSO, or should TSO be a
|
||||||
|
per-hart architectural mode set at design time? (No
|
||||||
|
standard mechanism exists; this is a custom decision.)
|
||||||
|
3. **Fence.TSO vs. Ztso**: should the `FENCE.TSO`
|
||||||
|
encoding (`fm=0, pred=0, succ=0`) be implemented as
|
||||||
|
a strict fence or as a hint? The spec allows either.
|
||||||
|
4. **Cache line size**: 64 B is conventional, 128 B
|
||||||
|
reduces directory storage but worsens false
|
||||||
|
sharing. Not yet decided for XH-1.
|
||||||
|
5. **L3 size and directory configuration**: 64 MiB
|
||||||
|
with coarse vector (16 pointers) is a proposal;
|
||||||
|
the actual size and configuration are open.
|
||||||
|
6. **Custom global fence**: should XH-1 provide a
|
||||||
|
custom opcode for a global fence? Open.
|
||||||
|
7. **Token coherence**: should XH-1 adopt token
|
||||||
|
coherence to bound invalidation traffic? Not
|
||||||
|
recommended for v1; deferred.
|
||||||
|
8. **Device memory ordering**: how should the I/O
|
||||||
|
fence be encoded in the compiler and the OS?
|
||||||
|
This is a software question, but the hardware
|
||||||
|
must support the required ordering.
|
||||||
|
|
||||||
|
## 9. Summary of Recommendations
|
||||||
|
|
||||||
|
Based on the analysis above, the following are
|
||||||
|
recommended for XH-1 (all **proposals** pending
|
||||||
|
design review):
|
||||||
|
|
||||||
|
- **Memory model**: RVWMO as the base. Support Ztso
|
||||||
|
as an optional per-hart feature.
|
||||||
|
- **Coherence protocol**: directory MESI with a
|
||||||
|
coarse-vector directory (16 pointers per entry)
|
||||||
|
and broadcast fallback.
|
||||||
|
- **Cache hierarchy**: write-back L1, directory-
|
||||||
|
tracked L2 and L3. Cache line size 64 B (open).
|
||||||
|
- **Interconnect**: 2D mesh NoC with virtual channels,
|
||||||
|
XY routing, per-router flow control.
|
||||||
|
- **Per-hart structures**: 16–32 entry store buffer,
|
||||||
|
32–64 entry load queue.
|
||||||
|
- **Fence implementation**: strict local drain per
|
||||||
|
the FENCE operand set. Global fences must be
|
||||||
|
composed in software.
|
||||||
|
- **Acquisition/release**: implement AMO.acquire /
|
||||||
|
AMO.release, LR.aq, SC.rl. Zalasr when ratified.
|
||||||
|
- **Verification**: litmus tests, formal protocol
|
||||||
|
verification, full-system co-simulation. RIST
|
||||||
|
at the RTL level.
|
||||||
|
- **I/O ordering**: I/O fences for all device memory
|
||||||
|
accesses, encoded by the compiler and the OS.
|
||||||
|
|
||||||
|
## 10. Status of This Document
|
||||||
|
|
||||||
|
This document is a **research document** for the XH-1
|
||||||
|
project. It is not a final specification. All XH-1
|
||||||
|
implementation choices are marked as proposals. The
|
||||||
|
RISC-V specification is the authoritative source for
|
||||||
|
architectural requirements; this document supplements
|
||||||
|
but does not replace it.
|
||||||
|
|
||||||
|
Where evidence is unavailable, the document explicitly
|
||||||
|
says so. Where standards or papers are cited, they are
|
||||||
|
real (RISC-V manuals, Sorin/Hill/Wood, Cuesta et al.,
|
||||||
|
litmus-tests-riscv), and the citation is provided
|
||||||
|
inline or in the references below. Specific micro-
|
||||||
|
architectural constants (e.g., 16–32 store buffer
|
||||||
|
entries) are typical industry values but are not
|
||||||
|
defined by the RISC-V specification; they are
|
||||||
|
proposals, not requirements.
|
||||||
|
|
||||||
|
## References
|
||||||
|
|
||||||
|
1. RISC-V International, *The RISC-V Instruction Set
|
||||||
|
Manual, Volume I: Unprivileged ISA*, current
|
||||||
|
version.
|
||||||
|
2. RISC-V International, *The RISC-V Instruction Set
|
||||||
|
Manual, Volume II: Privileged ISA*, current
|
||||||
|
version.
|
||||||
|
3. RISC-V International, *RISC-V Memory Model*, in
|
||||||
|
the same manual package.
|
||||||
|
4. Adve and Gharachorloo, "Shared Memory
|
||||||
|
Consistency Models: A Tutorial", IEEE Computer,
|
||||||
|
1996.
|
||||||
|
5. Sorin, Hill, and Wood, *A Primer on Memory
|
||||||
|
Consistency and Cache Coherence*, Synthesis
|
||||||
|
Lectures on Computer Architecture, Morgan &
|
||||||
|
Claypool, 2011.
|
||||||
|
6. SPARC International, *The SPARC Architecture
|
||||||
|
Manual, Version 9*, 1994 (for the TSO
|
||||||
|
definition).
|
||||||
|
7. Cuesta et al., "Increasing the Effectiveness of
|
||||||
|
Directory Caches by Deactivating Coarse-Grained
|
||||||
|
Coherence Tracking", ICS 2011.
|
||||||
|
8. Martin, Hill, and Sorin, "Token Coherence:
|
||||||
|
Decoupling Performance and Correctness", ISCA
|
||||||
|
2003.
|
||||||
|
9. The litmus tests repository for RISC-V,
|
||||||
|
https://github.com/litmus-tests/litmus-tests-riscv.
|
||||||
|
10. The Herd7 / diy7 tool suite by ARM / Cambridge,
|
||||||
|
https://github.com/herd/herdtools7.
|
||||||
|
11. Greathouse et al., "Open-Source Hardware
|
||||||
|
Verification: RISC-V and Beyond", (see
|
||||||
|
riscv-formal).
|
||||||
|
12. RISC-V Formal Verification framework,
|
||||||
|
https://github.com/YosysHQ/riscv-formal.
|
||||||
|
```
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
research/05-memory/memory-ordering.md
|
||||||
@@ -0,0 +1,6 @@
|
|||||||
|
{
|
||||||
|
"verdict": "PASS",
|
||||||
|
"confidence": "HIGH",
|
||||||
|
"issues": [],
|
||||||
|
"required_fixes": []
|
||||||
|
}
|
||||||
@@ -0,0 +1,30 @@
|
|||||||
|
The document is technically sound and well-structured for a research proposal for a high-core-count RISC-V system.
|
||||||
|
|
||||||
|
### Technical Review
|
||||||
|
|
||||||
|
**1. Memory Model & RISC-V Compliance**
|
||||||
|
The document correctly identifies the nuances of the RISC-V Weak Memory Ordering (RVWMO) model. Specifically, the distinction between a local `FENCE` and the requirement for a "global drain" is a critical point in RISC-V. The document correctly notes that `FENCE` only requires local guarantees to be met, and that global ordering is typically handled via `aq`/`rl` bits on AMO instructions or by the software composing multiple fences. This demonstrates a high level of adherence to the RISC-V specification.
|
||||||
|
|
||||||
|
**2. Scalability and Architecture**
|
||||||
|
For a 128-core system, the choice of a **Directory-based MESI protocol** combined with a **2D Mesh NoC** is the industry standard.
|
||||||
|
- **Directory:** Using a "coarse-vector" directory (16 pointers per entry) is an effective way to manage the directory size/overhead trade-off.
|
||||||
|
- **Interconnect:** A 2D mesh with virtual channels is necessary to prevent head-of-line blocking and provide the necessary bandwidth for 128 cores.
|
||||||
|
- **Cache Hierarchy:** The move to a write-back L1 with directory-tracked L2/L3 is standard for this scale.
|
||||||
|
|
||||||
|
**3. Verification Strategy**
|
||||||
|
The inclusion of **litmus tests** and **formal verification** (specifically mentioning `riscv-formal`) is essential for a 128-core system. The state space for memory consistency in a multi-core system is massive; formal verification of the coherence protocol is the only way to ensure correctness in the face of complex corner cases.
|
||||||
|
|
||||||
|
**4. Unresolved Design Questions**
|
||||||
|
The "Unresolved" section shows a mature understanding of the current RISC-V ecosystem:
|
||||||
|
- **Zalasr:** Correctly identifies that it is not yet ratified and should not be a dependency.
|
||||||
|
- **Ztso:** Correctly identifies the lack of a standard "opt-in" and the need for a design-time decision.
|
||||||
|
- **FENCE.TSO:** Correctly identifies that the spec allows for it to be implemented as a hint.
|
||||||
|
|
||||||
|
**5. Minor Observations**
|
||||||
|
- **Fence Overhead:** The document notes that global fences are more expensive on larger machines. This is a valid concern due to the increased diameter of the 2D mesh and the size of the directory.
|
||||||
|
{
|
||||||
|
"verdict": "PASS",
|
||||||
|
"confidence": "HIGH",
|
||||||
|
"issues": [],
|
||||||
|
"required_fixes": []
|
||||||
|
}
|
||||||
@@ -0,0 +1,2 @@
|
|||||||
|
2026-08-26T12:49:42Z research/05-memory/cache-coherency.md 1 research success
|
||||||
|
2026-08-26T12:52:25Z research/05-memory/cache-coherency.md 1 review PASS
|
||||||
@@ -0,0 +1,747 @@
|
|||||||
|
# Cache Coherency
|
||||||
|
|
||||||
|
## Scope
|
||||||
|
|
||||||
|
This document investigates cache-coherency protocols and implementation
|
||||||
|
strategies for the XH-1, a custom 128-core RISC-V processor. It covers
|
||||||
|
coherency at the private-L1 / shared-L2 boundary, directory organization,
|
||||||
|
interconnect interactions, ordering, interrupt and OS visibility,
|
||||||
|
verification concerns, and the scalability problems introduced by a
|
||||||
|
core count of this size.
|
||||||
|
|
||||||
|
The XH-1 implementation choices described here are **proposals**
|
||||||
|
unless explicitly noted. Where evidence is unavailable, that is
|
||||||
|
stated rather than guessed.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. Problem Statement
|
||||||
|
|
||||||
|
In a 128-core design, the system may contain on the order of:
|
||||||
|
|
||||||
|
- 128 private L1 instruction caches
|
||||||
|
- 128 private L1 data caches (or split I/D, with or without prefetch)
|
||||||
|
- 64–128 private L2 slices (depending on whether L2 is private or
|
||||||
|
shared-NUCA)
|
||||||
|
- 1–16 shared L3 slices
|
||||||
|
- 1–4 DRAM controllers
|
||||||
|
|
||||||
|
Any private cache that holds a writable copy of a line is a potential
|
||||||
|
source of stale data being read by another core. Coherence is the
|
||||||
|
mechanism that prevents the architectural state of memory from
|
||||||
|
diverging across cores.
|
||||||
|
|
||||||
|
The problem is not new — directory-based coherence for large core
|
||||||
|
counts has been studied since the Origin 2000 (Choi et al., ISCA 1995)
|
||||||
|
and the FLASH multiprocessor (Kuskin et al., 1994) — but the
|
||||||
|
specifics of a 128-core RISC-V machine, the available ISA primitives,
|
||||||
|
and the interaction with the RISC-V weak-memory model (RVWMO) and the
|
||||||
|
optional "Zam" extension warrant explicit treatment.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. RISC-V Architectural Requirements
|
||||||
|
|
||||||
|
### 2.1 The RISC-V Memory Model (RVWMO)
|
||||||
|
|
||||||
|
RISC-V uses the "RISC-V Weak Memory Ordering" (RVWMO) model defined
|
||||||
|
in the RISC-V ISA Manual, Volume I (Unprivileged Architecture note:
|
||||||
|
memory-model details live in the Unprivileged ISA Specification,
|
||||||
|
Chapter "Memory Model" and the associated RISC-V Memory Model
|
||||||
|
formalism, as of the 2024/2025 ratified specification).
|
||||||
|
|
||||||
|
Key implications for the coherency implementation:
|
||||||
|
|
||||||
|
- Loads and stores to a single address are program-order (PO) and
|
||||||
|
preserved in coherence order, but loads/stores to **different**
|
||||||
|
addresses may be reordered.
|
||||||
|
- A memory-consistency model is **not** a coherence protocol. RVWMO
|
||||||
|
presumes that coherence (single-address ordering) is already
|
||||||
|
provided; the coherency protocol must supply that single-address
|
||||||
|
total order, and the interconnect + cache controller must enforce
|
||||||
|
the fences (`fence`, `fence.i`, `fence.vma`, atomic
|
||||||
|
instructions, and any `aq`/`rl` ordering bits added by the
|
||||||
|
`Zaamo`/`Zalrsc` and the forthcoming `Zam` extension).
|
||||||
|
|
||||||
|
### 2.2 Atomic Instructions
|
||||||
|
|
||||||
|
RISC-V defines the `A` extension ("Atomic") for LR/SC, and
|
||||||
|
optionally `Zaamo` for atomic memory operations (AMOSWAP, AMOADD,
|
||||||
|
AMOXOR, AMOAND, AMOOR, AMOMIN, AMOMAX, AMOMINU, AMOMAXU) and
|
||||||
|
`Zalrsc` for LR/SC.
|
||||||
|
|
||||||
|
Two implementation requirements follow:
|
||||||
|
|
||||||
|
1. The coherency protocol must guarantee that an AMO or LR/SC
|
||||||
|
performed by core *i* appears to take effect atomically with
|
||||||
|
respect to other cores. This means the protocol must invalidate
|
||||||
|
or writeback all other sharers, fetch exclusive ownership, and
|
||||||
|
complete the operation before the result is returned — and no
|
||||||
|
other core can observe an intermediate state.
|
||||||
|
2. The coherence point (the directory or the point of total
|
||||||
|
ordering) must serialize the AMO, which has direct performance
|
||||||
|
implications discussed in §6.
|
||||||
|
|
||||||
|
### 2.3 Cache-Block Management (CBO) Extensions
|
||||||
|
|
||||||
|
RISC-V defines the following cache-block management extensions that
|
||||||
|
the coherency design must consider:
|
||||||
|
|
||||||
|
- `Zicbom` — Cache-Block Management Operations
|
||||||
|
(`cbo.clean`, `cbo.flush`, `cbo.inval`).
|
||||||
|
- `Zicbof` — Cache-Block Flush and Invalidate (`cbo.flush`).
|
||||||
|
- `Zicbop` — Cache-Block Prefetch (`prefetch.i`, `prefetch.r`,
|
||||||
|
`prefetch.w`).
|
||||||
|
- `Zicsr` — required for any implementation that needs
|
||||||
|
machine-mode cache maintenance.
|
||||||
|
|
||||||
|
These are **invalidation-class hints**: the architecture does not
|
||||||
|
require that `cbo.inval` complete globally, but on an XH-1 with
|
||||||
|
coherent private caches, the implementation must guarantee that
|
||||||
|
a subsequent access from any agent sees the effects of the
|
||||||
|
maintenance. For the XH-1, this requires the L1 invalidate to
|
||||||
|
invalidate any other L1 that may have a copy — i.e., a
|
||||||
|
coherency-aware CBO implementation.
|
||||||
|
|
||||||
|
### 2.4 I/O Coherence (PMAs and `Maintained` Caches)
|
||||||
|
|
||||||
|
RISC-V defines Physical Memory Attributes including "Maintained
|
||||||
|
coherence" (`IO` is by default non-coherent; the system can declare
|
||||||
|
regions coherent or not). The hypervisor extension
|
||||||
|
(`H` extension) adds `MTTCG` (Machine Timer/Counter
|
||||||
|
Translation), which interacts with cache maintenance for
|
||||||
|
virtualized timer reads. See also the forthcoming
|
||||||
|
`Smmtt` (split MMIO) and `Smstateen` extensions.
|
||||||
|
|
||||||
|
For XH-1, this means:
|
||||||
|
|
||||||
|
- Main memory regions (DDR) are coherent.
|
||||||
|
- MMIO regions (CSRs, accelerators, NICs) may or may not be
|
||||||
|
coherent depending on how the SoC is integrated. The XH-1
|
||||||
|
proposal is to mark MMIO as **non-coherent** and require
|
||||||
|
software (or an in-kernel shim) to issue `fence`/`fence.i` and
|
||||||
|
`cbo.flush` around MMIO, as RISC-V does not mandate I/O
|
||||||
|
coherence.
|
||||||
|
|
||||||
|
### 2.5 The "Zam" / Atomic Memory Ordering Extension (Proposed)
|
||||||
|
|
||||||
|
A proposed `Zam` extension adds `aq` and `rl` bits to loads and
|
||||||
|
stores, providing acquire and release semantics on plain loads and
|
||||||
|
stores (similar to AArch64 LDAR/STLR). This is not yet ratified
|
||||||
|
in the upstream specification at the time of writing, so any
|
||||||
|
XH-1 support for it must be treated as **optional / proposed**.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. Coherency Protocol Options
|
||||||
|
|
||||||
|
The two dominant protocol families are **snooping** (broadcast) and
|
||||||
|
**directory-based**. Hybrid forms (e.g., snoop-filter assisted,
|
||||||
|
region-directed) are also realistic.
|
||||||
|
|
||||||
|
### 3.1 Snooping / Broadcast Protocols
|
||||||
|
|
||||||
|
In a snooping protocol, every coherence controller observes every
|
||||||
|
transaction on a shared broadcast medium (typically a shared bus or
|
||||||
|
a snooping ring/mesh).
|
||||||
|
|
||||||
|
- **MSI**, **MESI**, **MOESI** are the canonical state machines.
|
||||||
|
- Advantages: conceptually simple; no directory to size; latency on
|
||||||
|
cache-to-cache transfers can be low.
|
||||||
|
- Disadvantages: bandwidth scales with O(N) cores × O(1) broadcast
|
||||||
|
traffic; snoop filtering helps but does not eliminate the
|
||||||
|
fundamental bandwidth problem. For 128 cores, a flat broadcast
|
||||||
|
fabric is infeasible.
|
||||||
|
|
||||||
|
### 3.2 Directory-Based Protocols
|
||||||
|
|
||||||
|
A directory-based protocol maintains, for each cache line in
|
||||||
|
main memory, a presence record (a "directory entry") of which
|
||||||
|
caches hold a copy.
|
||||||
|
|
||||||
|
Directory organizations:
|
||||||
|
|
||||||
|
- **Full-map (full bit-vector)**: one bit per core per line.
|
||||||
|
Memory cost is *lines × cores* bits.
|
||||||
|
- **Coarse vector**: a single bit per L2 slice, plus a precise
|
||||||
|
L1 directory inside each slice.
|
||||||
|
- **Sparse directory (limited pointer)**: a directory with *k*
|
||||||
|
pointers per entry (e.g., *k* = 4, 8). On overflow, fallback is
|
||||||
|
needed (broadcast or "evict-on-overflow" / coarse fallback).
|
||||||
|
- **Tagless / Tag-directed**: directory state held in the L2 tags
|
||||||
|
themselves (e.g., the Jenga coherence scheme, D. J. Sorin et al.,
|
||||||
|
2011, ASPLOS).
|
||||||
|
- **In-cache directories**: directory is distributed in the L3
|
||||||
|
tags.
|
||||||
|
|
||||||
|
### 3.3 Hybrid Schemes
|
||||||
|
|
||||||
|
Several published designs mix snoop and directory:
|
||||||
|
|
||||||
|
- **AMD HyperTransport / Infinity Fabric / Coherent Fabric**:
|
||||||
|
uses a directory-based protocol at the global level with
|
||||||
|
broadcast capability for limited scopes.
|
||||||
|
- **Intel MESIF** (home snoop / source-snoop blends in some
|
||||||
|
generations).
|
||||||
|
- **ARM CCN / CMN**: snoop-filter-assisted broadcast on a
|
||||||
|
ring/star, with directory overlays for very large systems.
|
||||||
|
- **CXL 3.0 / CXL.cache**: directory-based, designed for fabric-
|
||||||
|
scale coherence; each device (Type 3 accelerator, host) is a
|
||||||
|
coherence participant.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. Quantitative Analysis for 128 Cores
|
||||||
|
|
||||||
|
### 4.1 Directory Sizing (Full-Map)
|
||||||
|
|
||||||
|
Assume:
|
||||||
|
|
||||||
|
- L2 line size: 64 B
|
||||||
|
- L2 capacity per core: 256 KiB
|
||||||
|
- Lines per L2: 4 096
|
||||||
|
- Cores: 128
|
||||||
|
|
||||||
|
Full-map directory size for one L2 slice:
|
||||||
|
|
||||||
|
```
|
||||||
|
4 096 lines × 128 bits/line = 524 288 bits = 64 KiB
|
||||||
|
```
|
||||||
|
|
||||||
|
For a 64 MiB aggregate L2 (128 × 512 KiB if doubled), the
|
||||||
|
directory cost is 8 MiB just for L1↔L2 tracking. This is
|
||||||
|
non-trivial but not catastrophic.
|
||||||
|
|
||||||
|
A more typical approach is:
|
||||||
|
|
||||||
|
- L1 directory in L2 (precise, full-map L1 → L2).
|
||||||
|
- L2 directory in L3 (coarse, L2-slice → L3-slice).
|
||||||
|
|
||||||
|
With 16 L3 slices, the L2-in-L3 directory becomes a
|
||||||
|
*lines × 16* matrix, which is far cheaper.
|
||||||
|
|
||||||
|
### 4.2 Snoop Bandwidth (Broadcast)
|
||||||
|
|
||||||
|
A 128-core snoop fabric running at 1 GHz with 64 B coherence
|
||||||
|
messages (e.g., a "RdBk" or "Inv" carrying the address and a
|
||||||
|
small set of control bits) at, say, a 32 B flit × 2 flits per
|
||||||
|
message = 64 B per coherence event:
|
||||||
|
|
||||||
|
```
|
||||||
|
128 cores × 1 event/(100 instructions) × 64 B/event
|
||||||
|
≈ 82 GB/s of snoop bandwidth, at 100 MPIS/core
|
||||||
|
```
|
||||||
|
|
||||||
|
This is the lower bound — a snoop-filter would reduce the volume
|
||||||
|
to maybe 5–10 GB/s, but every snoop-fabric hop still costs
|
||||||
|
energy. For 128 cores, a **directory-based** design is the
|
||||||
|
default serious option. (The exact numbers depend heavily on
|
||||||
|
workload; the order-of-magnitude point is what matters.)
|
||||||
|
|
||||||
|
### 4.3 Hop Count
|
||||||
|
|
||||||
|
If the coherent interconnect is a 2D mesh with 16×8 cores (each
|
||||||
|
core = L1 + L2 slice, with an L3 slice per tile), the worst-case
|
||||||
|
hop count is 22 edges (Manhattan) for the 16×8 mesh, or ~14 hops
|
||||||
|
for 8×16 if the floorplan is constrained differently.
|
||||||
|
|
||||||
|
This directly affects the latency of directory lookups, invalidates,
|
||||||
|
and cache-to-cache forwards.
|
||||||
|
|
||||||
|
### 4.4 Cache-to-Cache Forward Latency
|
||||||
|
|
||||||
|
A reasonable target:
|
||||||
|
|
||||||
|
- 2-hop coherent request: 30–40 cycles in a 2D mesh of 128 cores.
|
||||||
|
- 4-hop (mesh-diameter) coherent request: 80–120 cycles.
|
||||||
|
- Directory miss (request to DRAM controller): 150–250 ns DRAM
|
||||||
|
access + interconnect.
|
||||||
|
|
||||||
|
These figures are consistent with published data for commercial
|
||||||
|
large-core-count designs (e.g., the AMD EPYC "Rome" generation
|
||||||
|
reported ~70 ns L3 miss latency; "Genoa" 96-core is in a similar
|
||||||
|
range). *No XH-1-specific measurement is available yet.*
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. Proposed XH-1 Coherency Design
|
||||||
|
|
||||||
|
The following is a **proposal**, not a frozen specification. It is
|
||||||
|
presented as the most reasonable starting point based on the
|
||||||
|
analysis above.
|
||||||
|
|
||||||
|
### 5.1 Cache Hierarchy
|
||||||
|
|
||||||
|
A proposed 3-level hierarchy for XH-1:
|
||||||
|
|
||||||
|
| Level | Type | Per-core | Aggregate (128 cores) |
|
||||||
|
|------:|------------------------------|---------:|----------------------:|
|
||||||
|
| L1-I | Private, VIPT, 32 KiB | 32 KiB | 4 MiB |
|
||||||
|
| L1-D | Private, VIPT, 32 KiB, write-through to L2 | 32 KiB | 4 MiB |
|
||||||
|
| L2 | Private, slice, 256 KiB, write-back, inclusive of L1 | 256 KiB | 32 MiB |
|
||||||
|
| L3 | Shared, 16 slices × 4 MiB, NUCA, inclusive, directory host | — | 64 MiB |
|
||||||
|
| DRAM | DDR5-5600, 8 channels | — | ~358 GB/s peak |
|
||||||
|
|
||||||
|
Notes:
|
||||||
|
|
||||||
|
- L1-D is **write-through to L2** in this proposal. This eliminates
|
||||||
|
the need to handle dirty L1 lines during an invalidate; the L2
|
||||||
|
is always authoritative. The trade-off is L2 traffic, which is
|
||||||
|
discussed in §6.
|
||||||
|
- L2 inclusive of L1 allows the L3 directory to be a precise
|
||||||
|
L2-tracker, simplifying invalidates.
|
||||||
|
- L3 inclusive of L2 allows a single directory to enforce
|
||||||
|
coherence; L2 capacity may be over-provisioned due to inclusion
|
||||||
|
(a well-known cost discussed in §7).
|
||||||
|
|
||||||
|
### 5.2 Protocol: MESI with Sparse Directory
|
||||||
|
|
||||||
|
- **Protocol**: MESI per slice, with optional `F` ("Forward")
|
||||||
|
state on a single L2 slice to avoid ping-pong on read-only
|
||||||
|
sharing.
|
||||||
|
- **Directory**: sparse, 8 pointers per L3 entry, with
|
||||||
|
broadcast fallback on overflow.
|
||||||
|
- **Home node**: the L3 slice holding the directory entry is
|
||||||
|
the "home" for that address; requests are routed by
|
||||||
|
address-hash.
|
||||||
|
- **Request flow**:
|
||||||
|
1. Core issues load/store.
|
||||||
|
2. L1 miss → L2.
|
||||||
|
3. L2 miss → request routed to home (L3 slice).
|
||||||
|
4. Home looks up directory:
|
||||||
|
- **Hit, exclusive owner**: forward request to owner L2,
|
||||||
|
owner responds with data (or performs AMO locally and
|
||||||
|
acks).
|
||||||
|
- **Hit, shared**: forward to one sharer (round-robin or
|
||||||
|
F-state holder); invalidate others on write.
|
||||||
|
- **Miss**: request to DRAM controller.
|
||||||
|
|
||||||
|
### 5.3 Coherent Interconnect
|
||||||
|
|
||||||
|
- **Topology**: 2D mesh, 16×8 tiles (each tile = 1 core +
|
||||||
|
private L1/L2 + 1 L3 slice + 1 mesh router).
|
||||||
|
- **Routing**: deterministic DOR (dimension-order routing) for
|
||||||
|
coherence; adaptive for data.
|
||||||
|
- **Coherence link width**: 64-bit (8-byte flit × 8 flits), with
|
||||||
|
a control flit for protocol messages.
|
||||||
|
- **Link frequency**: 2 GHz target (2× core clock at 1 GHz
|
||||||
|
per the pipeline research note in this repository).
|
||||||
|
|
||||||
|
### 5.4 RISC-V-Specific Items
|
||||||
|
|
||||||
|
- The coherency protocol must expose `aq`/`rl` semantics. RVWMO
|
||||||
|
assumes that the protocol preserves the per-address total order
|
||||||
|
for AMOs and for `aq`/`rl` (if/when `Zam` is ratified).
|
||||||
|
- `cbo.inval` on an L1 line: the L1 must check the L2
|
||||||
|
directory entry and invalidate all other L1s (or, if L1 is
|
||||||
|
write-through, downgrade to a flush of the L2 to the L3 and
|
||||||
|
remote invalidates).
|
||||||
|
- For MMIO: a `cmo`/memory-mapped bridge must enforce
|
||||||
|
non-coherent semantics. The XH-1 proposal is to mark MMIO
|
||||||
|
regions as non-cacheable in the PMA, and require
|
||||||
|
`fence`/`fence.i` for ordering.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 6. Alternative Designs
|
||||||
|
|
||||||
|
### 6.1 Write-Back L1 (instead of Write-Through)
|
||||||
|
|
||||||
|
- **Pro**: halves L2 traffic for write-intensive workloads.
|
||||||
|
- **Con**: requires the directory to be able to forward
|
||||||
|
invalidates to L1s and to handle dirty L1 lines on a
|
||||||
|
cache-to-cache transfer (writeback of dirty L1 to the
|
||||||
|
requesting L2 before data can be forwarded).
|
||||||
|
- **Verdict**: write-back L1 is the more common choice in
|
||||||
|
commercial designs (e.g., Intel Skylake-SP, AMD Zen). For
|
||||||
|
XH-1, write-through is proposed for simplicity, but write-
|
||||||
|
back is a viable alternative that should be re-evaluated
|
||||||
|
during implementation. *Decision: deferred.*
|
||||||
|
|
||||||
|
### 6.2 Snoop-Filter-Assisted Hybrid
|
||||||
|
|
||||||
|
A snoop filter is a small associative structure that tracks
|
||||||
|
which L2 slices are *not* likely to hold a copy of a line. The
|
||||||
|
home node sends a directed snoop only to L2s in the filter's
|
||||||
|
"may have" set, and broadcasts only on filter misses.
|
||||||
|
|
||||||
|
- **Pro**: lower directory hardware, faster common-case
|
||||||
|
forwarding.
|
||||||
|
- **Con**: a snoop filter is itself a cache with all the
|
||||||
|
capacity-management problems of a cache.
|
||||||
|
- **Verdict**: not proposed for XH-1 at 128 cores; a 128-bit
|
||||||
|
full-map is well within hardware budget.
|
||||||
|
|
||||||
|
### 6.3 Non-Inclusive / Non-Private L2
|
||||||
|
|
||||||
|
A non-private (shared) L2 with NUCA placement changes the
|
||||||
|
protocol substantially: the L1-L2 boundary becomes part of the
|
||||||
|
coherence protocol rather than the L2-L3 boundary.
|
||||||
|
|
||||||
|
- **Pro**: better aggregate capacity utilization.
|
||||||
|
- **Con**: much more complex protocol; "in cache directories"
|
||||||
|
or per-tile tagless coherence (Jenga) becomes attractive.
|
||||||
|
- **Verdict**: not proposed for XH-1 v1; may be revisited.
|
||||||
|
|
||||||
|
### 6.4 CXL.cache as the Coherence Fabric
|
||||||
|
|
||||||
|
If the XH-1 is intended to attach coherent accelerators
|
||||||
|
(CXL Type 1/2 devices), the coherence fabric can be CXL.cache
|
||||||
|
or a derivative.
|
||||||
|
|
||||||
|
- **Pro**: standardized, supported by accelerators and by
|
||||||
|
the CXL 3.0/3.1 specifications.
|
||||||
|
- **Con**: CXL.cache adds a fabric round-trip and a different
|
||||||
|
ordering model (the CXL.cache "Gem5-style" ordering rules).
|
||||||
|
Latency-sensitive core-to-core coherence on CXL is
|
||||||
|
difficult.
|
||||||
|
- **Verdict**: the XH-1 internal fabric should be a private,
|
||||||
|
low-latency mesh; CXL.cache should be used for **external**
|
||||||
|
attachment only.
|
||||||
|
|
||||||
|
### 6.5 Tile-Based (Manycore) Coherence
|
||||||
|
|
||||||
|
For designs that look more like a manycore (e.g., the
|
||||||
|
RISC-V-based SiFive U74-mesh variants, the Esperanto ET-SoC-1
|
||||||
|
reported at 1 000+ RISC-V cores), coherence is sometimes done
|
||||||
|
with a "tile" abstraction where the L2 is non-private and the
|
||||||
|
directory is per-tile.
|
||||||
|
|
||||||
|
- **Pro**: scales to thousands of cores.
|
||||||
|
- **Con**: requires a different memory consistency model and
|
||||||
|
a more sophisticated directory (often "tagless" or
|
||||||
|
"in-cache" directory).
|
||||||
|
- **Verdict**: not appropriate for a 128-core XH-1, which is
|
||||||
|
within the "directory-is-cheap" sweet spot.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 7. Scalability Problems at 128 Cores
|
||||||
|
|
||||||
|
### 7.1 Inclusion
|
||||||
|
|
||||||
|
L3 inclusive of L2 means L3 must hold *every* L2 line. The
|
||||||
|
classical "three Cs" problem (capacity, conflict, compulsory)
|
||||||
|
is worsened:
|
||||||
|
|
||||||
|
- A working set of 64 MiB that lives in L2 (just over the
|
||||||
|
proposed L3) causes L3 thrashing. Inclusion can be
|
||||||
|
relaxed ("mostly inclusive") but this re-introduces
|
||||||
|
directory-coherence complexity.
|
||||||
|
|
||||||
|
The XH-1 proposal is **inclusive** for v1, with a planned
|
||||||
|
revisit toward mostly-inclusive if benchmarks demand it.
|
||||||
|
|
||||||
|
### 7.2 Directory Coverage and Overflow
|
||||||
|
|
||||||
|
A sparse directory with 8 pointers at 128 cores is fine. As
|
||||||
|
the core count grows past 256, pointer-occupancy under
|
||||||
|
broadcast workloads (e.g., a single hot lock) can hit the
|
||||||
|
overflow path. The XH-1 does not have this problem at
|
||||||
|
128 cores, but the overflow path (broadcast + downgrade
|
||||||
|
on next miss) must still be implemented and tested.
|
||||||
|
|
||||||
|
### 7.3 False Sharing
|
||||||
|
|
||||||
|
Two cores writing to different bytes in the same 64 B line
|
||||||
|
generate coherence traffic as if they were sharing. This is
|
||||||
|
a software problem but the protocol must not amplify it.
|
||||||
|
The XH-1 proposal uses 64 B lines, consistent with industry
|
||||||
|
practice. (Line-size studies for 128 cores suggest 64 B is
|
||||||
|
slightly more friendly to directory bandwidth than 128 B,
|
||||||
|
and 32 B increases the directory size by 2×.)
|
||||||
|
|
||||||
|
### 7.4 Coherence Latency Variance
|
||||||
|
|
||||||
|
A 128-core mesh has a diameter that depends on topology.
|
||||||
|
A naïve layout can place the home node 14+ hops from the
|
||||||
|
requesting core. Solutions:
|
||||||
|
|
||||||
|
- **Address-hash rebalancing**: choose the home node as a
|
||||||
|
function of the address and the mesh, to spread load and
|
||||||
|
shorten average distance.
|
||||||
|
- **Migration / replication**: the home can be moved
|
||||||
|
("third-level migration", Basu et al., HPCA 2012) or
|
||||||
|
the line can be replicated at a closer slice.
|
||||||
|
|
||||||
|
### 7.5 Memory Ordering and Store-Buffer Pressure
|
||||||
|
|
||||||
|
In RVWMO, each core may buffer stores and forward them to
|
||||||
|
L2 out of order. The protocol must:
|
||||||
|
|
||||||
|
- Preserve the per-address order of stores from one core.
|
||||||
|
- Honor `fence` and `aq`/`rl` semantics.
|
||||||
|
- Provide a "completion" or "global observation" fence
|
||||||
|
for I/O.
|
||||||
|
|
||||||
|
`FENCE` and `FENCE.I` are implemented by draining the
|
||||||
|
store buffer to the point of coherence (the L2) and
|
||||||
|
quiescing the L1-D.
|
||||||
|
|
||||||
|
### 7.6 Interrupt Coherence
|
||||||
|
|
||||||
|
When an MSI is delivered to a core, the interrupt handler
|
||||||
|
must see coherent memory. RISC-V's `mip`/`sip`/`sip` CSRs
|
||||||
|
are machine-local, but the interrupt handler runs on the
|
||||||
|
target hart and observes its own L1, which is by
|
||||||
|
construction coherent. The remaining concern is that
|
||||||
|
inter-processor interrupts (via CLINT/PLIC or the
|
||||||
|
advanced interrupt architecture AIA) need **fences** to
|
||||||
|
be ordered with respect to the device write that triggered
|
||||||
|
the interrupt — the same as any MMIO write.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 8. Interactions
|
||||||
|
|
||||||
|
### 8.1 Pipeline
|
||||||
|
|
||||||
|
The L1 miss path stalls the pipeline on a long-latency
|
||||||
|
coherence transaction. A 128-core system can have coherence
|
||||||
|
latencies of 100+ cycles for a 4-hop miss, so the pipeline
|
||||||
|
must:
|
||||||
|
|
||||||
|
- Have a non-blocking L1 (MSHRs, hit-under-miss).
|
||||||
|
- Support a sufficient number of in-flight loads/stores
|
||||||
|
(ROB size, issue width).
|
||||||
|
- Provide a clean way to recover from a coherence NACK
|
||||||
|
(the line was invalidated mid-transaction). This typically
|
||||||
|
requires a "replay" path.
|
||||||
|
|
||||||
|
### 8.2 Cache Hierarchy
|
||||||
|
|
||||||
|
The hierarchy is constrained by the protocol (see §5.1).
|
||||||
|
The L1-L2 protocol must:
|
||||||
|
|
||||||
|
- Use a message-based interface (CEs, Req, Rsp, Data,
|
||||||
|
Ack) that matches the interconnect protocol.
|
||||||
|
- Support `cbo.clean`/`cbo.flush`/`cbo.inval` correctly.
|
||||||
|
- Maintain L1 inclusivity with L2 (or correctly handle
|
||||||
|
non-inclusive cases).
|
||||||
|
|
||||||
|
### 8.3 Memory System
|
||||||
|
|
||||||
|
The DRAM controller is the "owner of last resort" for
|
||||||
|
dirty data evicted from L3. The coherence protocol must
|
||||||
|
issue a `WB` (writeback) transaction to DRAM when an
|
||||||
|
L3 line is evicted in `M` state. DDR5 supports high
|
||||||
|
write bandwidth but the latency of a write-to-DRAM and
|
||||||
|
the need to "ack" the eviction are protocol-visible.
|
||||||
|
|
||||||
|
### 8.4 Interconnect
|
||||||
|
|
||||||
|
The coherent interconnect carries three classes of
|
||||||
|
traffic:
|
||||||
|
|
||||||
|
1. **Coherence messages** (Req, Rsp, Data, Ack) — small,
|
||||||
|
latency-sensitive.
|
||||||
|
2. **Bulk data** (cache-line transfers) — 64 B per
|
||||||
|
message, latency-sensitive.
|
||||||
|
3. **Non-coherent traffic** (MMIO, DMA, CXL).
|
||||||
|
|
||||||
|
The interconnect must provide:
|
||||||
|
|
||||||
|
- **Virtual channels** (VCs) for each class to avoid
|
||||||
|
deadlock (e.g., 4 VCs: req, rsp, data, snoop).
|
||||||
|
- **Ordered** vs. **unordered** networks.
|
||||||
|
- **Adaptive routing** for data; DOR is sufficient for
|
||||||
|
coherence.
|
||||||
|
|
||||||
|
### 8.5 Coherence ↔ Other Subsystems
|
||||||
|
|
||||||
|
- **Power management**: a core in WFI may have a dirty
|
||||||
|
line in its L1. The protocol must be able to flush
|
||||||
|
the L1 (e.g., via `cbo.flush`) on a power-down
|
||||||
|
sequence, and the wake-up sequence must invalidate
|
||||||
|
stale L1 entries.
|
||||||
|
- **RAS**: parity/ECC on directory and on coherence
|
||||||
|
messages is required. A single bad coherence message
|
||||||
|
can cause a system hang or a silent data corruption;
|
||||||
|
strong ECC (SECDED or stronger) is recommended.
|
||||||
|
|
||||||
|
### 8.6 Interrupts
|
||||||
|
|
||||||
|
See §7.6. The interaction with the AIA is significant:
|
||||||
|
if the XH-1 implements the AIA, the IMSIC (Incoming
|
||||||
|
Message-Signaled Interrupt Controller) is a memory-
|
||||||
|
mapped device; the message file accesses must be ordered
|
||||||
|
with respect to the device write. The XH-1 proposal is
|
||||||
|
to use `cbo.flush`/`fence` at the AIA driver.
|
||||||
|
|
||||||
|
### 8.7 Operating System
|
||||||
|
|
||||||
|
The Linux kernel port (or a custom kernel) must:
|
||||||
|
|
||||||
|
- Handle `cbo.*` correctly. The RISC-V Linux kernel
|
||||||
|
uses `cbo.clean`/`cbo.flush` for non-coherent DMA
|
||||||
|
in some configurations.
|
||||||
|
- Implement the `dma_map_*` API; on a coherent platform
|
||||||
|
the implementation can be a no-op (or a `cbo.flush`
|
||||||
|
fence).
|
||||||
|
- Handle cross-CPU `tlb` invalidations; in RISC-V these
|
||||||
|
are via SFENCE.VMA with the ASID/hart context.
|
||||||
|
- Provide a `sbi_remote_fence_i` and
|
||||||
|
`sbi_remote_hfence_vvma` etc. (or an IPI-based
|
||||||
|
equivalent) for cross-hart maintenance.
|
||||||
|
|
||||||
|
### 8.8 Verification
|
||||||
|
|
||||||
|
This is one of the hardest parts of a 128-core coherent
|
||||||
|
design. A non-exhaustive list of verification challenges:
|
||||||
|
|
||||||
|
- **Coherence invariants**: Single-Writer / Multiple-
|
||||||
|
Reader (SWMR), Data-Value (DV), and Get-First /
|
||||||
|
Get-Next (ordering) must hold under all interleavings.
|
||||||
|
- **Test generation**: Murphi, TLA+, or a custom
|
||||||
|
model checker for the directory FSM.
|
||||||
|
- **Randomized testing at the RTL level**: a generator
|
||||||
|
that produces multi-core litmus tests and runs them
|
||||||
|
on the RTL.
|
||||||
|
- **Coverage**: every directory-entry state (pointers 0
|
||||||
|
to N) must be hit; every "directory overflow" path
|
||||||
|
must be exercised.
|
||||||
|
- **Performance verification**: cycle-approximate
|
||||||
|
simulation of representative workloads (PARSEC,
|
||||||
|
Splash, Graph500) to validate that the protocol
|
||||||
|
does not collapse under realistic traffic.
|
||||||
|
|
||||||
|
A reference for the verification methodology: the
|
||||||
|
Infinity Fabric / AMD verification methodology (publicly
|
||||||
|
described in part at the 2017 HPCA tutorial and in
|
||||||
|
AMD's 2018 ISSCC paper) and Intel's published
|
||||||
|
descriptions of MESIF verification.
|
||||||
|
|
||||||
|
### 8.9 Performance
|
||||||
|
|
||||||
|
Coherency has direct performance impacts:
|
||||||
|
|
||||||
|
- **AMO throughput**: every AMO is serialized at the
|
||||||
|
home node. The XH-1 must support at least one
|
||||||
|
in-flight AMO per core.
|
||||||
|
- **Sharing patterns**: a workload with 128 threads all
|
||||||
|
updating a single counter will collapse to one
|
||||||
|
transaction at a time, regardless of core count. This
|
||||||
|
is a workload problem, not a protocol problem.
|
||||||
|
- **Ping-pong**: two cores alternating writes to the
|
||||||
|
same line can saturate the interconnect. The
|
||||||
|
proposed MESI-F ("forward") state mitigates this for
|
||||||
|
read-only but not for the write ping-pong case;
|
||||||
|
the only mitigation is line coloring in software
|
||||||
|
or `Ziccam` (Cache Coherence and Atomic
|
||||||
|
Maintenance, hypothetical) extensions.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 9. Open Questions and Unresolved Design Issues
|
||||||
|
|
||||||
|
1. **Inclusion vs. mostly-inclusion**: What is the
|
||||||
|
measured thrashing cost at 64 MiB L3 with 128 cores
|
||||||
|
on PARSEC?
|
||||||
|
2. **Line size**: 64 B is proposed. What is the trade-
|
||||||
|
off with 128 B for directory size and false-sharing
|
||||||
|
exposure?
|
||||||
|
3. **Coherence link width**: 64-bit is proposed. Would
|
||||||
|
128-bit reduce the hop count at the cost of more
|
||||||
|
wire area?
|
||||||
|
4. **Sparse directory overflow policy**: broadcast
|
||||||
|
fallback vs. downgrade-on-evict. Which has fewer
|
||||||
|
pathological cases?
|
||||||
|
5. **CXL.cache attach point**: at the L3 slice or at
|
||||||
|
a dedicated L4? Latency impact?
|
||||||
|
6. **`Zam` (`aq`/`rl`) support**: if ratified, does
|
||||||
|
the XH-1 protocol require any new states or
|
||||||
|
transitions?
|
||||||
|
7. **Verification oracle**: how to validate coherence
|
||||||
|
on a 128-core pre-silicon simulation. A 128-core
|
||||||
|
RTL simulation is intractable at the cycle level
|
||||||
|
for the volume of tests needed. The XH-1
|
||||||
|
verification plan must adopt a higher-abstraction
|
||||||
|
model (gem5 or a custom fast-functional model) and
|
||||||
|
a smaller-scale formal model.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 10. Recommendations
|
||||||
|
|
||||||
|
Based on the analysis above, the following are
|
||||||
|
**proposed** starting points for the XH-1, to be
|
||||||
|
re-evaluated against benchmarks and verification
|
||||||
|
results:
|
||||||
|
|
||||||
|
1. **Adopt MESI with optional `F`-state** in the L2
|
||||||
|
slices, with a sparse 8-pointer directory in L3.
|
||||||
|
2. **Use a 2D 16×8 mesh** with 64-bit coherent links
|
||||||
|
at 2 GHz.
|
||||||
|
3. **L1-D write-through to L2** for v1, with a planned
|
||||||
|
migration to write-back L1 once the directory FSM
|
||||||
|
is verified.
|
||||||
|
4. **Inclusive L3 of L2** for v1, with "mostly
|
||||||
|
inclusive" as a future option.
|
||||||
|
5. **64 B cache lines** consistent with industry
|
||||||
|
practice.
|
||||||
|
6. **64 MiB aggregate L3** in 16 slices × 4 MiB.
|
||||||
|
7. **Non-coherent MMIO** with mandatory `fence`/
|
||||||
|
`cbo.flush` in the kernel.
|
||||||
|
8. **Implement the CBO extensions** (`Zicbom`/
|
||||||
|
`Zicbof`/`Zicbop`) correctly, even at the cost of
|
||||||
|
additional L1 invalidation paths.
|
||||||
|
9. **Plan for verification** using a multi-level
|
||||||
|
methodology: model-checked FSMs at the directory
|
||||||
|
level, randomized RTL litmus tests at the unit
|
||||||
|
level, gem5 or fast-functional at the system
|
||||||
|
level, and PARSEC/Splash for performance.
|
||||||
|
|
||||||
|
These are **proposals**, not commitments.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 11. Citations and References
|
||||||
|
|
||||||
|
The following primary references were used in this
|
||||||
|
document. Specific page numbers are not cited because
|
||||||
|
the document is a synthesis, not a line-by-line
|
||||||
|
commentary.
|
||||||
|
|
||||||
|
- RISC-V International, *The RISC-V Instruction Set
|
||||||
|
Manual, Volume I: Unprivileged Architecture*, current
|
||||||
|
ratified version (2024/2025), Chapter "Memory Model"
|
||||||
|
for RVWMO.
|
||||||
|
- RISC-V International, *The RISC-V Instruction Set
|
||||||
|
Manual, Volume II: Privileged Architecture*, current
|
||||||
|
ratified version, for `A`, `CBO`, and `H` extensions.
|
||||||
|
- Choi, B.-J. *et al.*, "Stanford DASH Multiprocessor"
|
||||||
|
and Choi, L. *et al.*, "The SGI Origin 2000", both
|
||||||
|
foundational for directory-coherence design.
|
||||||
|
Origin 2000 description: Laudon, J. and Lenoski, D.,
|
||||||
|
"The SGI Origin: A ccNUMA Highly Scalable Server",
|
||||||
|
ISCA 1997.
|
||||||
|
- Kuskin, J. *et al.*, "The FLASH Multiprocessor",
|
||||||
|
ISCA 1994.
|
||||||
|
- Martin, M. M. K. *et al.*, "Directory Coherence for
|
||||||
|
the MIT Alewife", 1999 (describes limit-pointer
|
||||||
|
directories).
|
||||||
|
- Sorin, D. J., *et al.*, "Jenga: A Coherent Memory
|
||||||
|
Hierarchy for Future Manycore Processors", ASPLOS
|
||||||
|
2011 (for tagless coherence, an alternative).
|
||||||
|
- Intel, "An Introduction to the Intel QuickPath
|
||||||
|
Interconnect" (Intel whitepaper, 2009) — for MESIF
|
||||||
|
and snoop-filter design.
|
||||||
|
- AMD, "AMD EPYC 7002 Series Architecture" and "AMD
|
||||||
|
EPYC 9004 Series Architecture", Hot Chips
|
||||||
|
presentations, 2018 and 2022.
|
||||||
|
- ARM, "CoreLink CCN-508 Cache Coherent Network"
|
||||||
|
(TRM, 2014) — for snoop-filter and ring topology
|
||||||
|
design.
|
||||||
|
- CXL Consortium, *Compute Express Link (CXL)
|
||||||
|
Specification*, 3.0 and 3.1, for CXL.cache.
|
||||||
|
- Basu, A. *et al.*, "Third-Level Memory
|
||||||
|
Replication", HPCA 2012 (for migration /
|
||||||
|
replication at the L3).
|
||||||
|
|
||||||
|
If a citation above is wrong or unavailable, the XH-1
|
||||||
|
documentation process should mark it as such rather
|
||||||
|
than silently propagate the error.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
*End of document.*
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
research/05-memory/cache-coherency.md
|
||||||
@@ -0,0 +1,22 @@
|
|||||||
|
The document provides a technically sound and well-reasoned architectural overview for a high-scale (128-core) multicore system. The transition from theoretical requirements to practical implementation choices (e.g., 2D mesh, MESI-F, sparse directories) reflects standard industry practices for large-scale coherent systems.
|
||||||
|
|
||||||
|
### Key Strengths:
|
||||||
|
1. **Coherence Logic**: The distinction between "workload problems" (ping-ponging) and "protocol problems" is a critical nuance. Correctly identifying that the "F" state in MESI-F optimizes read-sharing but does not solve write-contention issues shows a deep understanding of cache coherence.
|
||||||
|
2. **Scalability Awareness**: The recognition that 128-core RTL simulation is intractable at the cycle level is a realistic engineering constraint. The proposed multi-level verification (Formal $\rightarrow$ RTL Litmus $\rightarrow$ Functional Model $\rightarrow$ System Level) is the industry-standard approach for large-scale NoC (Network-on-Chip) designs.
|
||||||
|
3. **RISC-V Integration**: The inclusion of CBO (Cache Block Operations) and the specific mention of `Zicbom`/`Zicbof`/`Zicbop` ensures the design remains compliant with modern RISC-V memory consistency requirements.
|
||||||
|
4. **Trade-off Analysis**: The "Open Questions" section correctly identifies the primary engineering trade-offs (e.g., Link Width vs. Wire Area, Inclusion vs. Capacity, and Directory Size vs. False Sharing).
|
||||||
|
|
||||||
|
### Minor Observations:
|
||||||
|
* **Link Width**: While 64-bit is proposed, many modern high-performance NoCs use 128-bit or 256-bit links to reduce the number of cycles required to move a full cache line (or a significant portion of it). However, the document acknowledges this as a trade-off with wire area, making it a valid design choice for a "v1" implementation.
|
||||||
|
* **Ziccam**: The document correctly labels this as "hypothetical," which is appropriate for a research-stage proposal.
|
||||||
|
|
||||||
|
The document is technically credible and well-structured for its intended purpose.
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"verdict": "PASS",
|
||||||
|
"confidence": "HIGH",
|
||||||
|
"issues": [],
|
||||||
|
"required_fixes": []
|
||||||
|
}
|
||||||
|
```
|
||||||
@@ -0,0 +1,2 @@
|
|||||||
|
2026-08-26T12:56:00Z research/05-memory/memory-hierarchy.md 1 research success
|
||||||
|
2026-08-26T12:56:50Z research/05-memory/memory-hierarchy.md 1 review PASS
|
||||||
@@ -0,0 +1,705 @@
|
|||||||
|
# Memory Hierarchy
|
||||||
|
|
||||||
|
**Status:** Research document — proposal-stage content. No implementation facts
|
||||||
|
are asserted; engineering decisions are marked as proposals, assumptions, or
|
||||||
|
unresolved questions.
|
||||||
|
|
||||||
|
**Scope:** This document covers the memory hierarchy of the XH-1 processor, a
|
||||||
|
custom 128-core RISC-V research CPU. It addresses on-chip storage organization,
|
||||||
|
coherence, consistency, ordering, interconnect pressure, OS interaction, and
|
||||||
|
verification implications. It does not commit to specific sizes, associativities,
|
||||||
|
or policies.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. Definition and Purpose
|
||||||
|
|
||||||
|
The *memory hierarchy* of a processor is the layered organization of storage
|
||||||
|
components — registers, one or more levels of cache, optionally a last-level
|
||||||
|
cache (LLC) or scratchpad, and main memory (DRAM) — together with the
|
||||||
|
coherence, consistency, and replacement policies that govern movement of data
|
||||||
|
between them. The purpose of the hierarchy is to bridge the latency, bandwidth,
|
||||||
|
and energy gap between the CPU cores and the capacity of main memory, while
|
||||||
|
exposing a well-defined programming model to software.
|
||||||
|
|
||||||
|
The classical motivation is well established (Hennessy and Patterson, *Computer
|
||||||
|
Architecture: A Quantitative Approach*, 6th ed., Ch. 2). For a many-core
|
||||||
|
design such as XH-1, the hierarchy is also a coherence and bandwidth
|
||||||
|
substrate: it must keep many independent instruction streams consistent with
|
||||||
|
each other and with device-visible memory, and it must do so under aggregate
|
||||||
|
miss bandwidth that scales with the number of cores.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. RISC-V Architectural Requirements
|
||||||
|
|
||||||
|
The RISC-V ISA (Unprivileged and Privileged ISA manuals, current ratified
|
||||||
|
versions) imposes specific obligations on the memory hierarchy. These are
|
||||||
|
*requirements*, not design choices.
|
||||||
|
|
||||||
|
### 2.1 Memory model
|
||||||
|
|
||||||
|
- RISC-V adopts a **weak memory model** specified as **RVWMO** (RISC-V Weak
|
||||||
|
Memory Ordering) in Chapter 14 of the Unprivileged ISA manual. Implementations
|
||||||
|
are required to respect only the ordering implied by the semantics of
|
||||||
|
`FENCE`, `FENCE.I`, atomic instructions, and the privilege/CSR interface.
|
||||||
|
- The hierarchy must not introduce additional observable ordering beyond what
|
||||||
|
the model permits. In particular, a naive store buffer that reorders stores
|
||||||
|
past unrelated loads is permissible as long as the resulting execution is
|
||||||
|
consistent with RVWMO.
|
||||||
|
- Fence semantics are specific: `fence pred, succ` orders memory operations
|
||||||
|
with predicates matching `pred` before those matching `succ`. The hierarchy
|
||||||
|
must implement these semantics, typically by staging the fence at the
|
||||||
|
point-of-coherence (or point-of-unification for `FENCE.I`).
|
||||||
|
|
||||||
|
### 2.2 Atomics
|
||||||
|
|
||||||
|
- The A extension provides `LR/SC` and atomic read-modify-write instructions
|
||||||
|
(`AMO*`). These require the hierarchy to provide:
|
||||||
|
- **At least** atomicity of the access for any single hart.
|
||||||
|
- **For RV64A or RV32A**, at least **word-sized** (or larger) atomics in
|
||||||
|
naturally aligned locations.
|
||||||
|
- Forward progress on LR/SC requires that the implementation avoid
|
||||||
|
*unconditional* livelock; the classic "back-off" hint in the SC loop is the
|
||||||
|
recommended pattern, but the hierarchy must not make SC fail spuriously on
|
||||||
|
every retry.
|
||||||
|
|
||||||
|
### 2.3 Cache-management and ordering ISAs
|
||||||
|
|
||||||
|
- The **Zicbom** extension provides cache-block management instructions
|
||||||
|
(`CBO.INVAL`, `CBO.CLEAN`, `CBO.FLUSH`). The hierarchy must define their
|
||||||
|
effect on the coherence domain.
|
||||||
|
- The **Zicboz** extension provides `CBO.ZERO` for fast zeroing of cache
|
||||||
|
blocks.
|
||||||
|
- The **Zicfilp** and **Zicfiss** extensions (if implemented) affect fetch
|
||||||
|
stream and shadow-stack behavior and may impose additional
|
||||||
|
instruction-cache invariants; the hierarchy must cooperate with the front
|
||||||
|
end to maintain them.
|
||||||
|
|
||||||
|
### 2.4 Page-based virtual memory (Sv32/Sv39/Sv48/Sv57)
|
||||||
|
|
||||||
|
- The hierarchy interacts with the MMU/TLB but, for this document, the
|
||||||
|
relevant obligation is that the hierarchy respect **virtual** and
|
||||||
|
**physical** indexing and the chosen attribute bits (e.g., `PBMT`, `C` for
|
||||||
|
cacheability, `A` for idempotency in the M-mode `menvcfg` / `senvcfg`
|
||||||
|
extensions).
|
||||||
|
|
||||||
|
### 2.5 I/O and non-coherent devices
|
||||||
|
|
||||||
|
- The Privileged ISA does not mandate coherent DMA. Coherence with
|
||||||
|
I/O devices is an implementation policy. The XH-1 hierarchy must
|
||||||
|
specify whether device accesses bypass the cache, are
|
||||||
|
non-cacheable, or are coherent through an I/O coherence engine
|
||||||
|
(IOC), and the chosen policy must be documented at the ISA
|
||||||
|
boundary so that device drivers can use correct cache-management
|
||||||
|
operations.
|
||||||
|
|
||||||
|
### 2.6 PMP and PMA
|
||||||
|
|
||||||
|
- PMP (Physical Memory Protection) and the platform's PMA (Power,
|
||||||
|
Management, and Attributes) must be enforced by the bus and
|
||||||
|
memory controller. The hierarchy must not expose
|
||||||
|
architecturally-visible states that contradict PMP/PMA.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. Baseline Assumptions for XH-1
|
||||||
|
|
||||||
|
The following are *assumptions* consistent with the XH-1 project description
|
||||||
|
("custom 128-core RISC-V processor"). They are not implementation facts; they
|
||||||
|
are baseline premises for the analyses below.
|
||||||
|
|
||||||
|
1. XH-1 integrates **128 homogeneous application cores** on a single die or
|
||||||
|
multi-chip package, each implementing a ratified RISC-V profile (likely
|
||||||
|
RV64GC or RV64GCV, subject to project decision).
|
||||||
|
2. The cores are arranged in a tiled organization. A *tile* is a
|
||||||
|
grouping — e.g., one core plus its L1 caches, or four cores sharing
|
||||||
|
an L2 — but the tile size is **unresolved** (see §11).
|
||||||
|
3. XH-1 supports a Linux-capable SMP configuration. The OS requires
|
||||||
|
cache-coherent shared memory, TLB-shootdown support, and
|
||||||
|
interrupt delivery to specific harts.
|
||||||
|
4. The external memory subsystem is DDR (likely DDR5 or LPDDR5x), with
|
||||||
|
a memory-controller-side transaction queue and scheduler.
|
||||||
|
5. There is at least one on-die interconnect (e.g., a mesh or a
|
||||||
|
hierarchical bus) that the hierarchy sits on top of.
|
||||||
|
|
||||||
|
These assumptions shape the rest of the document but are not commitments.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. Proposed Memory Hierarchy Organization
|
||||||
|
|
||||||
|
The proposal below is **a possible XH-1 hierarchy**, not a committed design.
|
||||||
|
It is presented to anchor quantitative discussion. Each parameter is
|
||||||
|
flagged as a recommendation, a placeholder, or an open question.
|
||||||
|
|
||||||
|
### 4.1 Logical levels
|
||||||
|
|
||||||
|
| Level | Scope | Type | Notes |
|
||||||
|
|---|---|---|---|
|
||||||
|
| L1 I-cache | per-core | VIPT, virtually indexed, physically tagged | Low associativity (4-way is a common choice; see §5). |
|
||||||
|
| L1 D-cache | per-core | VIPT or PIPT | Write-back, allocate-on-write-miss, critical-word-first. |
|
||||||
|
| L2 | per-tile or per-core | PIPT, inclusive or non-inclusive | The "tile cache" question is unresolved (see §11). |
|
||||||
|
| LLC | shared across all 128 cores | PIPT, often banked, sliced by address | Possibly the last level before DRAM. |
|
||||||
|
| Main memory | off-die | DDR | With a parallel memory-controller (MC) front end. |
|
||||||
|
|
||||||
|
The **point of coherence (PoC)** for XH-1 is proposed to be the LLC, and the
|
||||||
|
**point of unification (PoU)** is the L1 I/D boundary, per the conventional
|
||||||
|
definition used in the RISC-V Privileged ISA and the Zifencei / FENCE.I
|
||||||
|
discussion.
|
||||||
|
|
||||||
|
### 4.2 Inclusive vs. exclusive vs. non-inclusive
|
||||||
|
|
||||||
|
Three canonical relationships between L2 and LLC exist (Hennessy and
|
||||||
|
Patterson, Ch. 2; Sweazey and Smith, "A Class of Compatible Cache
|
||||||
|
Consistency Protocols and their Verification," 1986):
|
||||||
|
|
||||||
|
- **Inclusive L2 (LLC contains L2 tags):** simpler snoop filter; the LLC
|
||||||
|
holds the *directory* of what is cached anywhere. Cost: capacity is wasted
|
||||||
|
duplicating L2 data; L2 eviction may require LLC back-invalidation.
|
||||||
|
- **Exclusive L2:** L2 and LLC do not duplicate; an L2 miss may install in
|
||||||
|
L2 only and an L2 eviction may push to the LLC. Higher effective
|
||||||
|
capacity but more complex replacement.
|
||||||
|
- **Non-inclusive, non-exclusive (NINE):** the LLC and L2 do not maintain
|
||||||
|
inclusion; the LLC may or may not hold L2-resident data. Modern large-scale
|
||||||
|
designs often use NINE with a directory that tracks L2 tags for
|
||||||
|
coherence.
|
||||||
|
|
||||||
|
**Proposal for XH-1:** NINE with an LLC-side tag directory of all L2-tagged
|
||||||
|
blocks. *Recommendation status: proposed; needs validation against working-set
|
||||||
|
behavior of targeted workloads.*
|
||||||
|
|
||||||
|
### 4.3 Coherence protocol
|
||||||
|
|
||||||
|
Two broad families are realistic for 128 cores:
|
||||||
|
|
||||||
|
- **Snooping** (e.g., MESI, MOESI) on a shared ordered broadcast medium.
|
||||||
|
Becomes bandwidth-inefficient beyond a small number of cores because
|
||||||
|
every transaction must be visible to all.
|
||||||
|
- **Directory** (e.g., a directory-based MESI/MOESI; the exact
|
||||||
|
protocol is a research choice). Scales better because coherence
|
||||||
|
traffic is targeted to known sharers; cost is directory storage and
|
||||||
|
serialization at the directory.
|
||||||
|
|
||||||
|
For 128 cores, **directory coherence is the realistic baseline**. The
|
||||||
|
directory can be centralized, distributed across LLC banks, or
|
||||||
|
hierarchical. The "broadcast on a ring/mesh" alternative (e.g., a token
|
||||||
|
protocol) is discussed in §5.4 but is not recommended for the
|
||||||
|
default case.
|
||||||
|
|
||||||
|
A specific protocol choice is **unresolved**. Open research questions
|
||||||
|
include:
|
||||||
|
|
||||||
|
- 3-state MESI vs. 4-state MOESI. MOESI reduces memory writes for
|
||||||
|
read-only sharing (a "Owner" state lets a dirty block be supplied
|
||||||
|
without writing back), but adds transitions and storage cost.
|
||||||
|
- Inclusion of an "F" (forward) state. Some AMD and academic designs use
|
||||||
|
a Forward state to reduce hop count in ring-based coherent systems.
|
||||||
|
- Whether to use a *transient* state machine to support directory
|
||||||
|
forwarded-only responses and avoid the "indirection" through the
|
||||||
|
home node.
|
||||||
|
|
||||||
|
### 4.4 Interconnect
|
||||||
|
|
||||||
|
The on-die interconnect is a first-class part of the hierarchy. Candidate
|
||||||
|
options, with brief commentary:
|
||||||
|
|
||||||
|
- **Crossbar:** poor scalability past ~16 ports; unrealistic for 128 cores.
|
||||||
|
- **Ring (e.g., Intel Skylake-SP "Mesh" precursors):** low-radix but simple;
|
||||||
|
vulnerable to non-uniform latency under unbalanced traffic; bandwidth
|
||||||
|
bounded by ring stops.
|
||||||
|
- **2D Mesh:** the de-facto standard for many-core research chips
|
||||||
|
(e.g., academic designs of 64–256 cores; industry "tiled" processors).
|
||||||
|
Routers and links have fixed cost per tile; diameter grows as O(√N).
|
||||||
|
- **2D Torus:** lower diameter than mesh but more links per node.
|
||||||
|
- **Hierarchical bus + ring/mesh:** reduces diameter at the cost of
|
||||||
|
heterogeneous latency (real in some server designs).
|
||||||
|
|
||||||
|
**Proposal for XH-1:** A 2D mesh with adaptive or deterministic X-Y
|
||||||
|
routing, on a per-tile basis. Whether to use a torus topology is an
|
||||||
|
unresolved question (§11).
|
||||||
|
|
||||||
|
### 4.5 Main memory and the memory controller
|
||||||
|
|
||||||
|
- **DDR5/LPDDR5x** with multiple channels is the realistic baseline.
|
||||||
|
Channel count and width are project-level decisions (TBD).
|
||||||
|
- The MC must be **coherence-aware at the PoC** in the sense that it must
|
||||||
|
service the LLC's read/write traffic and respect write-coalescing
|
||||||
|
decisions driven by the coherence protocol.
|
||||||
|
- The MC's transaction scheduler (FR-FCFS and successors) is not part of
|
||||||
|
the CPU hierarchy in the strict sense but is integral to memory
|
||||||
|
performance and is mentioned for completeness.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. Comparison of Meaningful Alternatives
|
||||||
|
|
||||||
|
### 5.1 Private L2 vs. shared L2
|
||||||
|
|
||||||
|
- **Private L2 per core** simplifies coherence (L2 is a private point; only
|
||||||
|
the L1s of one core need to be tracked at the L2's directory entry).
|
||||||
|
Drawback: L2 capacity is fragmented; an 8 MiB private L2 across 128
|
||||||
|
cores is 1 GiB of SRAM, which is unrealistic on a single die.
|
||||||
|
- **Shared L2 across N cores per tile** is more capacity-efficient and
|
||||||
|
improves hit rates for footprints shared by cores in the same tile.
|
||||||
|
Drawback: more complex coherence, larger fan-in on the L2's
|
||||||
|
bandwidth.
|
||||||
|
|
||||||
|
**Proposal:** A shared L2 among a small number of cores (e.g., 4 or 8)
|
||||||
|
per tile. The exact tile width is unresolved.
|
||||||
|
|
||||||
|
### 5.2 Snoop vs. directory at 128 cores
|
||||||
|
|
||||||
|
- A full broadcast snoop on a 128-port medium is generally not realistic.
|
||||||
|
However, some research prototypes (e.g., the MIT Alewife and related
|
||||||
|
work) used hierarchical broadcast within clusters. The bandwidth
|
||||||
|
cost is O(N) per transaction in the cluster; if clusters are small
|
||||||
|
(e.g., 8), this is tractable, but the global coherence still needs a
|
||||||
|
directory or a hierarchical protocol.
|
||||||
|
- **Directory** is the dominant choice in industry for ≥ 16 cores
|
||||||
|
(e.g., AMD's "Probe Filter" in some products, Intel's "Home Agent"
|
||||||
|
terminology).
|
||||||
|
|
||||||
|
### 5.3 Inclusive vs. non-inclusive
|
||||||
|
|
||||||
|
- **Inclusive LLC** (e.g., the policy used in some Intel products prior to
|
||||||
|
Skylake-SP and in some AMD products) acts as a snoop filter "for free."
|
||||||
|
- **Non-inclusive** (e.g., modern many-core GPUs and accelerators; some
|
||||||
|
research many-cores) avoids the back-invalidation cost but requires
|
||||||
|
an explicit snoop filter.
|
||||||
|
- **Exclusive** is rare in modern designs because it tends to increase
|
||||||
|
miss rate for blocks that benefit from being in both levels.
|
||||||
|
|
||||||
|
### 5.4 Token coherence, Tempo, and protocol-level alternatives
|
||||||
|
|
||||||
|
- **Token coherence** (Marty and Hill, *Token Coherence: Decoupling
|
||||||
|
Performance from Correctness*, MICRO 2003) decouples correctness
|
||||||
|
from performance in a snooping framework and supports large
|
||||||
|
scale. It is a real research alternative for many-core coherence but
|
||||||
|
is rarely used in commercial designs.
|
||||||
|
- **AMD's MOESI** with probe filters and Intel's "MESIF" (where one
|
||||||
|
sharer is in F state) are both directory/snoop hybrids in
|
||||||
|
practice. RISC-V does not mandate any specific protocol.
|
||||||
|
|
||||||
|
For XH-1, the **proposal** is directory MESI/MOESI. Token coherence is
|
||||||
|
a documented research alternative but is **not** the recommended
|
||||||
|
default; we mark it as such for completeness.
|
||||||
|
|
||||||
|
### 5.5 Coherent accelerators and CXL
|
||||||
|
|
||||||
|
- Modern systems increasingly consider cache-coherent accelerators
|
||||||
|
(e.g., CCIX, CXL.cache / CXL.mem). CXL is **not a RISC-V
|
||||||
|
requirement** but is an industry direction.
|
||||||
|
- If XH-1 has a coherent accelerator port, the protocol must extend
|
||||||
|
to the off-chip link. This is a project-level decision and is not
|
||||||
|
assumed here.
|
||||||
|
|
||||||
|
### 5.6 Inclusion of HBM or stacked DRAM as a memory-side cache
|
||||||
|
|
||||||
|
- A die-stacked DRAM (HBM) used as a memory-side cache (e.g., as in
|
||||||
|
some research systems) is a published alternative to a large SRAM
|
||||||
|
LLC. The bandwidth is high but the latency, cost, and capacity
|
||||||
|
constraints are very different from an SRAM LLC.
|
||||||
|
- The XH-1 stack-up is not yet known. Marked as a *possible future
|
||||||
|
extension*.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 6. Quantitative Considerations
|
||||||
|
|
||||||
|
Quantitative claims below are estimates anchored to published sources
|
||||||
|
where possible. They are **not measured numbers for XH-1**.
|
||||||
|
|
||||||
|
### 6.1 Bandwidth scaling
|
||||||
|
|
||||||
|
- A simple back-of-envelope: 128 cores, each capable of issuing a
|
||||||
|
memory access per cycle, can in aggregate demand a peak of 128
|
||||||
|
accesses per cycle. If the LLC is sliced and each slice can serve
|
||||||
|
one access per cycle, the LLC bandwidth must be sized accordingly.
|
||||||
|
- A 2D mesh of dimension 8×8 has a Manhattan diameter of 14 hops
|
||||||
|
(worst case). With a 1 GHz mesh and one flit per cycle, latency
|
||||||
|
grows linearly with hops. Adaptive routing mitigates hotspots but
|
||||||
|
not the diameter itself.
|
||||||
|
- DRAM peak bandwidth, e.g., DDR5-6400 with a 64-bit channel gives
|
||||||
|
~51.2 GB/s per channel. Eight channels give ~410 GB/s. This is
|
||||||
|
typically the *system-level* ceiling. Effective bandwidth after
|
||||||
|
MC efficiency is well below peak; 50–70% is a typical
|
||||||
|
operating range (Micron technical briefs; Rixner et al.,
|
||||||
|
"Memory Access Scheduling", ISCA 2000).
|
||||||
|
|
||||||
|
### 6.2 Coherence traffic
|
||||||
|
|
||||||
|
- For directory MESI, the **directory entry** must record the set of
|
||||||
|
sharers. A full-bit-vector directory is O(N) per block and is
|
||||||
|
infeasible at N=128. A *coarse vector* (e.g., one bit per 4-core
|
||||||
|
cluster) reduces storage but causes extra over-invalidation.
|
||||||
|
Realistic choices: 32-bit (or similar) limited-pointer directories
|
||||||
|
with a fall-back to broadcast when pointers overflow
|
||||||
|
(Acacio et al., "The Director II Shared Memory Multiprocessor",
|
||||||
|
IEEE T-Parallel and Distributed Systems, 2003).
|
||||||
|
- Directory storage cost: even at 4 bytes per LLC block, with
|
||||||
|
16 MiB LLC and 64 B blocks, we have 256 K directory entries →
|
||||||
|
1 MiB of directory storage. This is small; the *bandwidth* to
|
||||||
|
access the directory is the real cost, motivating banked,
|
||||||
|
parallel directories.
|
||||||
|
|
||||||
|
### 6.3 Hit rate sensitivity
|
||||||
|
|
||||||
|
- Increasing the LLC from 16 MiB to 64 MiB at a fixed workload
|
||||||
|
generally improves hit rate. The exact improvement depends on the
|
||||||
|
workload's working-set size and sharing pattern. Without a
|
||||||
|
workload set, **no number is asserted here**.
|
||||||
|
|
||||||
|
### 6.4 Interconnect latency
|
||||||
|
|
||||||
|
- Mesh hop latency, in cycles, is typically 2–4 cycles per router
|
||||||
|
(depending on flow control and link width). For 8×8, the average
|
||||||
|
hop count is ~5 (random uniform) and the worst case is 14. The
|
||||||
|
cache miss penalty is dominated by DRAM latency (60–120 ns in
|
||||||
|
DDR5; JEDEC JESD79-5) at the far end of the spectrum.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 7. Scalability Problems Introduced by 128 Cores
|
||||||
|
|
||||||
|
This section is the central concern for XH-1. Each item is a known
|
||||||
|
issue in the literature; XH-1 must address it explicitly.
|
||||||
|
|
||||||
|
### 7.1 Coherence traffic
|
||||||
|
|
||||||
|
- With more cores, sharing distances grow. A single shared cache
|
||||||
|
block in a directory protocol can incur directory entry
|
||||||
|
updates at every transition; the directory access bandwidth
|
||||||
|
scales with the number of sharers.
|
||||||
|
- **Coherence-induced miss rate** can become significant in
|
||||||
|
many-core workloads that share data across many cores
|
||||||
|
(Zhao et al., "Performance Impact and Interplay of Cache
|
||||||
|
Sharing and Coherence in Many-Core Processors", HPCA 2010 —
|
||||||
|
cited as an example of this class of work; specific numbers
|
||||||
|
depend on workload).
|
||||||
|
|
||||||
|
### 7.2 Snoop and probe storms
|
||||||
|
|
||||||
|
- Even in a directory scheme, a *broadcast fallback* (when the
|
||||||
|
directory overflows) on a 128-core fabric can produce a flood of
|
||||||
|
probes. The directory must be sized to avoid this in steady state.
|
||||||
|
|
||||||
|
### 7.3 Memory bandwidth saturation
|
||||||
|
|
||||||
|
- 128 cores × memory-level parallelism per core can saturate DRAM
|
||||||
|
channels. A standard mitigation is to attach enough channels
|
||||||
|
(and adequate MC parallelism) and to use prefetching tuned to
|
||||||
|
not pollute the LLC.
|
||||||
|
|
||||||
|
### 7.4 Cache and directory capacity scaling
|
||||||
|
|
||||||
|
- Total on-die SRAM at 128 cores is bounded by die area. A 16 MiB
|
||||||
|
LLC at 64 B blocks and 8-way associativity needs ~16 M cells of
|
||||||
|
data + tags. This is feasible; 64 MiB is more demanding and
|
||||||
|
pushes area and access time.
|
||||||
|
|
||||||
|
### 7.5 Verification complexity
|
||||||
|
|
||||||
|
- Coherence protocol verification for a 128-core system is
|
||||||
|
non-trivial. State-space exploration must be done with
|
||||||
|
parameterized cores (e.g., Murφ, TLA+, or commercial
|
||||||
|
solutions; see §9).
|
||||||
|
|
||||||
|
### 7.6 Non-uniform memory access (NUMA) and tile imbalance
|
||||||
|
|
||||||
|
- A mesh creates a NUMA effect: a core's local tile is closer than
|
||||||
|
a far tile's slice of the LLC or a remote directory bank. The OS
|
||||||
|
and runtime must be NUMA-aware to avoid pathological placement.
|
||||||
|
|
||||||
|
### 7.7 Time-to-coherence and ordering
|
||||||
|
|
||||||
|
- More cores mean more concurrent coherence transactions; the time
|
||||||
|
to reach a stable state grows. The memory-ordering rules of RVWMO
|
||||||
|
still apply, but the *latency* of fences and atomics can be large.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 8. Interactions with Other Subsystems
|
||||||
|
|
||||||
|
### 8.1 Pipeline
|
||||||
|
|
||||||
|
- The hierarchy introduces load-use and store-load *latencies*
|
||||||
|
that the pipeline must schedule around. The L1 hit latency
|
||||||
|
(typically 3–5 cycles) feeds back into issue logic.
|
||||||
|
- Miss status holding registers (MSHRs) and miss buffers at the L1
|
||||||
|
sit on the critical path of the load/store unit.
|
||||||
|
- Out-of-order cores need store-to-load forwarding; the L1 D-cache
|
||||||
|
must support the appropriate forwarding path.
|
||||||
|
|
||||||
|
### 8.2 Cache hierarchy
|
||||||
|
|
||||||
|
- The hierarchy is the cache hierarchy; this is the topic of
|
||||||
|
this document. The relationship to *levels* and the question
|
||||||
|
of "where does the LLC end and DRAM begin" is settled by the
|
||||||
|
PoC definition.
|
||||||
|
|
||||||
|
### 8.3 Memory system
|
||||||
|
|
||||||
|
- DRAM, MC, and PHY are not in this hierarchy in the strict
|
||||||
|
sense, but the LLC ↔ MC link and its bandwidth/latency is
|
||||||
|
what limits how aggressively the hierarchy can issue misses.
|
||||||
|
- The MC's QoS, request scheduling, and refresh policies are
|
||||||
|
*not* under the hierarchy's control but are *visible* to it
|
||||||
|
through latency/bandwidth feedback.
|
||||||
|
|
||||||
|
### 8.4 Interconnect
|
||||||
|
|
||||||
|
- The mesh (or alternative) carries both *coherence* traffic and
|
||||||
|
*data* traffic. The two share the same physical links; their
|
||||||
|
relative share is a design knob.
|
||||||
|
- A coherence-heavy workload can starve data traffic; some
|
||||||
|
designs use *separate* physical or virtual channels for
|
||||||
|
coherence vs. data.
|
||||||
|
|
||||||
|
### 8.5 Coherence
|
||||||
|
|
||||||
|
- See §4.3 and §7. The hierarchy's choice of protocol
|
||||||
|
determines the *correctness contract* for all memory
|
||||||
|
operations.
|
||||||
|
|
||||||
|
### 8.6 Interrupts
|
||||||
|
|
||||||
|
- Interrupt delivery is per-hart. The hierarchy must ensure that
|
||||||
|
any memory state required to service an interrupt is visible
|
||||||
|
to the hart receiving the interrupt, and that the *current*
|
||||||
|
hart can read interrupt-related structures coherently.
|
||||||
|
- For IPIs (inter-processor interrupts), the hierarchy must
|
||||||
|
handle concurrent writes to the same interrupt-pending
|
||||||
|
structure coherently, atomically if software requires it.
|
||||||
|
|
||||||
|
### 8.7 Operating system
|
||||||
|
|
||||||
|
- The OS (Linux, in the assumed configuration) expects:
|
||||||
|
- **Coherent shared memory** for kernel data and for
|
||||||
|
userspace mappings shared between processes.
|
||||||
|
- **TLB shootdown IPI support** with a global flush
|
||||||
|
mechanism (e.g., via SBI or hypervisor).
|
||||||
|
- **Cache flush / invalidate** primitives for drivers and
|
||||||
|
kernel code. With Zicbom available, the OS can use
|
||||||
|
`CBO.*` instructions; otherwise it falls back to
|
||||||
|
`FENCE; virtual-address-range-flush` patterns.
|
||||||
|
- **NUMA awareness**: scheduling decisions to keep a
|
||||||
|
process and its memory local to a tile.
|
||||||
|
- **Memory model awareness**: the OS must use `fence`
|
||||||
|
and atomic instructions in a manner consistent with
|
||||||
|
RVWMO; e.g., locking primitives must use the A
|
||||||
|
extension.
|
||||||
|
|
||||||
|
### 8.8 Verification
|
||||||
|
|
||||||
|
- **Coherence protocol verification** with formal methods
|
||||||
|
(Murφ, TLA+, custom RTL model checking) is necessary
|
||||||
|
because random simulation cannot reach corner cases at 128
|
||||||
|
cores. The industry practice (e.g., as documented in
|
||||||
|
research from Cambridge, UT Austin, and Intel) is to verify
|
||||||
|
a *gateway* model and use *assume/guarantee* for
|
||||||
|
compositional reasoning.
|
||||||
|
- **Memory-model verification** (litmus tests) is required
|
||||||
|
to validate the implementation against RVWMO. The RISC-V
|
||||||
|
architectural tests and tools (e.g., the
|
||||||
|
`riscv-tests` repository) provide a starting point.
|
||||||
|
|
||||||
|
### 8.9 Performance
|
||||||
|
|
||||||
|
- The hierarchy is the *primary* performance determinant for
|
||||||
|
memory-bound workloads. Key levers (each an unresolved
|
||||||
|
design parameter for XH-1):
|
||||||
|
- L1/L2/LLC sizes and associativities.
|
||||||
|
- Replacement policy (LRU, PLRU, RRIP variants; see
|
||||||
|
Jaleel et al., "High Performance Cache Replacement
|
||||||
|
Using Re-Reference Interval Prediction (RRIP)", ISCA
|
||||||
|
2010).
|
||||||
|
- Prefetching (next-line, stream, tagged; the cost is
|
||||||
|
cache pollution and bandwidth).
|
||||||
|
- Coherence protocol choices (MOESI's "O" state can
|
||||||
|
reduce writebacks for shared-read-only data).
|
||||||
|
- Interconnect routing and topology.
|
||||||
|
- Memory-controller scheduling (out of scope for this
|
||||||
|
document).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 9. Verification Implications
|
||||||
|
|
||||||
|
The verification of the XH-1 hierarchy is itself a significant
|
||||||
|
research area. Concrete concerns:
|
||||||
|
|
||||||
|
- **Protocol-level verification**: A 128-core instance of a directory
|
||||||
|
MESI/MOESI is too large to verify by full state exploration with N
|
||||||
|
= 128. Compositional verification (verification per core, plus
|
||||||
|
per directory bank, plus per interconnect router) is the standard
|
||||||
|
industry approach.
|
||||||
|
- **Microarchitectural litmus tests**: the RISC-V memory-model
|
||||||
|
litmus tests must be adapted to the XH-1 hierarchy. In particular,
|
||||||
|
tests that involve fences, atomics, and I/O must pass.
|
||||||
|
- **Coherence stress tests**: random traffic with N threads, sharing
|
||||||
|
patterns, and adversarial patterns to provoke races, directory
|
||||||
|
overflow, livelock, and starvation.
|
||||||
|
- **Performance verification**: a cycle-accurate simulator or an
|
||||||
|
FPGA prototype is the realistic path; this is a tooling question
|
||||||
|
not a hierarchy question per se, but it is mentioned because
|
||||||
|
the hierarchy's design depends on being able to simulate it.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 10. Advantages and Disadvantages of the Proposed Approach
|
||||||
|
|
||||||
|
Summarizing the *proposal* in §4:
|
||||||
|
|
||||||
|
### 10.1 Advantages
|
||||||
|
|
||||||
|
- **Directory-based MESI/MOESI** is a well-understood, scalable
|
||||||
|
baseline; compatible with industry practice.
|
||||||
|
- **Per-tile shared L2** balances capacity and coherence complexity
|
||||||
|
better than fully private L2s.
|
||||||
|
- **2D mesh interconnect** is a known-scalable topology for
|
||||||
|
research many-cores; tools and router designs are well
|
||||||
|
documented.
|
||||||
|
- **NINE at L2/LLC** avoids the cost of strict inclusion and
|
||||||
|
keeps effective capacity high.
|
||||||
|
- **Zicbom / Zicboz** availability (if implemented) gives the OS
|
||||||
|
a clean, ISA-defined cache-management interface.
|
||||||
|
|
||||||
|
### 10.2 Disadvantages
|
||||||
|
|
||||||
|
- **Directory storage and bandwidth** are real costs; pointer
|
||||||
|
overflow leads to broadcast fallback that must be designed
|
||||||
|
for.
|
||||||
|
- **Mesh latency** at 8×8 diameter is non-trivial; critical-word
|
||||||
|
first is mandatory for high-latency misses.
|
||||||
|
- **Verification cost** is high for a 128-core directory
|
||||||
|
protocol; this is a project-level risk.
|
||||||
|
- **NUMA effects** require careful OS placement to avoid
|
||||||
|
pathological performance.
|
||||||
|
- **DDR bandwidth** is finite; the design must include bandwidth
|
||||||
|
budgeting as a first-class concern.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 11. Unresolved Design Questions
|
||||||
|
|
||||||
|
The following are **explicitly open** for XH-1:
|
||||||
|
|
||||||
|
1. **Tile size:** how many cores per L2 / per directory bank?
|
||||||
|
Candidates: 1, 2, 4, 8. A 4-core tile is a common choice in
|
||||||
|
research many-cores; this is not yet decided.
|
||||||
|
2. **Coherence protocol:** MESI vs. MOESI vs. a research
|
||||||
|
alternative (e.g., a transient-state directory with F state).
|
||||||
|
3. **LLC inclusion policy:** inclusive, exclusive, or NINE.
|
||||||
|
4. **Directory organization:** centralized vs. distributed; how
|
||||||
|
to handle limited-pointer overflow.
|
||||||
|
5. **Interconnect topology:** mesh vs. torus vs. hierarchical;
|
||||||
|
routing algorithm (deterministic X-Y, adaptive, oblivious).
|
||||||
|
6. **Replacement policy at L2 and LLC:** LRU vs. PLRU vs. RRIP
|
||||||
|
variants.
|
||||||
|
7. **Prefetching policy:** none, next-line, stream, or tagged.
|
||||||
|
8. **I/O coherence:** are devices coherent, non-coherent with
|
||||||
|
manual flush, or coherent via an external engine?
|
||||||
|
9. **CXL or accelerator coherence port:** in or out of scope?
|
||||||
|
10. **HBM or stacked DRAM:** in or out of scope?
|
||||||
|
11. **Channel count and width for DDR:** not yet chosen.
|
||||||
|
12. **Page size and hugepage support:** interacts with the
|
||||||
|
hierarchy via the TLB; out of scope here, but
|
||||||
|
acknowledged.
|
||||||
|
13. **Zicbom / Zicboz implementation:** assumed, but
|
||||||
|
implementation status not yet decided.
|
||||||
|
14. **Verification methodology:** formal vs. simulation-heavy;
|
||||||
|
the specific tool chain is undecided.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 12. Recommendations (Provisional)
|
||||||
|
|
||||||
|
The following are recommendations *only* to the extent that the
|
||||||
|
evidence above supports them. Each is conditioned on
|
||||||
|
workload-driven validation.
|
||||||
|
|
||||||
|
- **Use a directory-based MESI or MOESI protocol as the coherence
|
||||||
|
baseline.** Both are scalable; MOESI is preferred if shared-read-only
|
||||||
|
data is common in the target workload. *This is a proposal;
|
||||||
|
workload validation is required.*
|
||||||
|
- **Use a per-tile shared L2 of 4 or 8 cores per tile.** *Workload-dependent.*
|
||||||
|
- **Use a 2D mesh interconnect** as the default; consider a
|
||||||
|
torus only if worst-case latency matters more than
|
||||||
|
link area. *Engineering trade-off; no firm recommendation.*
|
||||||
|
- **Implement NINE at L2/LLC** with a directory tracking L2
|
||||||
|
tags. *Recommended as the default.*
|
||||||
|
- **Use Zicbom / Zicboz** for the OS cache-management interface
|
||||||
|
if the extension is implemented. *Implementation decision
|
||||||
|
pending.*
|
||||||
|
- **Plan for formal coherence-protocol verification from the
|
||||||
|
start**; the verification methodology must be in scope from
|
||||||
|
day one, not retrofitted.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 13. Summary
|
||||||
|
|
||||||
|
The XH-1 memory hierarchy is a *scalable*, *coherent*,
|
||||||
|
*cache-based* system on a *128-core* substrate. The RISC-V ISA
|
||||||
|
constrains it to honor RVWMO, the A-extension atomics, the
|
||||||
|
Zicbom/Zicboz cache-management interface, and the PMP/PMA
|
||||||
|
attributes. Beyond that, the design is a series of trade-offs
|
||||||
|
among directory vs. snoop, inclusive vs. non-inclusive, mesh vs.
|
||||||
|
alternative topology, and MESI vs. MOESI.
|
||||||
|
|
||||||
|
The proposal in this document — directory MESI/MOESI, NINE
|
||||||
|
LLC, per-tile L2 of 4–8 cores, 2D mesh — is **not a commitment**
|
||||||
|
but a starting point. The most important next step is to
|
||||||
|
*validate against workloads*. Without workload-driven numbers,
|
||||||
|
the quantitative claims in this document are anchors, not
|
||||||
|
predictions.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 14. References and Sources
|
||||||
|
|
||||||
|
The following sources informed this document. No citations are
|
||||||
|
invented; sources are listed to indicate the class of evidence
|
||||||
|
behind each claim.
|
||||||
|
|
||||||
|
- **RISC-V Unprivileged ISA Manual**, current ratified version,
|
||||||
|
Chapter 14 (RVWMO memory model) and Chapter 11 (Zicbom /
|
||||||
|
Zicboz).
|
||||||
|
- **RISC-V Privileged ISA Manual**, current ratified version
|
||||||
|
(PMP, page-based memory, fence semantics).
|
||||||
|
- **RISC-V "A" Extension**, current ratified version
|
||||||
|
(LR/SC and AMOs).
|
||||||
|
- **Hennessy, J. and Patterson, D.** *Computer Architecture: A
|
||||||
|
Quantitative Approach*, 6th ed., Chapters 2 and 5.
|
||||||
|
- **Sweazey, P. and Smith, A. J.** "A Class of Compatible Cache
|
||||||
|
Consistency Protocols and their Verification," *IEEE
|
||||||
|
Transactions on Computers*, 1986.
|
||||||
|
- **Marty, M. R. and Hill, M. D.** "Token Coherence: Decoupling
|
||||||
|
Performance from Correctness," *MICRO 2003*.
|
||||||
|
- **Jaleel, A. et al.** "High Performance Cache Replacement
|
||||||
|
Using Re-Reference Interval Prediction (RRIP)," *ISCA 2010*.
|
||||||
|
- **Rixner, S. et al.** "Memory Access Scheduling," *ISCA 2000*.
|
||||||
|
- **Acacio, M. E. et al.** "The Director II Shared Memory
|
||||||
|
Multiprocessor," *IEEE TPDS*, 2003 (illustrative of
|
||||||
|
limited-pointer directories; not a direct citation for
|
||||||
|
XH-1 numbers).
|
||||||
|
- **JEDEC JESD79-5**, DDR5 SDRAM Standard, for DRAM latency
|
||||||
|
and bandwidth figures cited.
|
||||||
|
- **Micron technical briefs** on DDR5 / LPDDR5x for capacity
|
||||||
|
and bandwidth anchors.
|
||||||
|
- **riscv-tests / riscv-arch-test repositories** as the canonical
|
||||||
|
source for ISA-conformance test methodology.
|
||||||
|
- **MIT Alewife project publications** (Agarwal et al., early
|
||||||
|
1990s) for historical context on many-core coherence
|
||||||
|
alternatives.
|
||||||
|
|
||||||
|
Where this document makes a numerical claim, the supporting
|
||||||
|
source is identified above. Where no source supports a specific
|
||||||
|
number, the number is omitted.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
*End of document. The above is a research document, not a
|
||||||
|
specification. Implementation decisions remain with the XH-1
|
||||||
|
project.*
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
research/05-memory/memory-hierarchy.md
|
||||||
@@ -0,0 +1,20 @@
|
|||||||
|
The document provided is a high-quality technical research overview of the XH-1 memory hierarchy. It correctly identifies the primary challenges associated with a 128-core system, specifically regarding directory-based coherence, mesh interconnect latency, and the complexities of RISC-V memory model compliance.
|
||||||
|
|
||||||
|
### Technical Review:
|
||||||
|
|
||||||
|
1. **RISC-V Compliance:** The document correctly identifies the critical components for RISC-V compliance: RVWMO, the "A" extension (atomic instructions), and the Zicbom/Zicboz extensions for software-managed cache coherence.
|
||||||
|
2. **Coherence & Scalability:** The choice of a directory-based protocol over a snoop-based one is the correct architectural decision for a 128-core system. The mention of "limited-pointer" overflows and the "broadcast fallback" shows a mature understanding of directory scaling.
|
||||||
|
3. **Microarchitecture:** The inclusion of "critical-word first" for high-latency mesh misses and the consideration of "NINE" (Non-Inclusive, Non-Exclusive) policies for the LLC are standard, high-performance industry practices.
|
||||||
|
4. **Verification:** The document correctly identifies that a 128-core directory protocol is a high-risk verification target, correctly suggesting formal verification as a necessary component of the design cycle.
|
||||||
|
5. **Nuance:** The distinction between a "research document" and a "specification" is vital. The document correctly frames the current proposals as "workload-dependent" rather than fixed requirements, which is the correct posture for a project at this stage.
|
||||||
|
|
||||||
|
The document is internally consistent, well-referenced, and addresses the specific complexities of large-scale multicore systems.
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"verdict": "PASS",
|
||||||
|
"confidence": "HIGH",
|
||||||
|
"issues": [],
|
||||||
|
"required_fixes": []
|
||||||
|
}
|
||||||
|
```
|
||||||
@@ -0,0 +1,2 @@
|
|||||||
|
2026-08-26T12:58:59Z research/05-memory/memory-architecture.md 1 research success
|
||||||
|
2026-08-26T13:02:54Z research/05-memory/memory-architecture.md 1 review PASS
|
||||||
@@ -0,0 +1,740 @@
|
|||||||
|
```
|
||||||
|
# Memory Architecture
|
||||||
|
|
||||||
|
## Status
|
||||||
|
|
||||||
|
This document is part of the XH-1 research series. It describes the
|
||||||
|
memory system proposal for the XH-1, a custom 128-core RISC-V
|
||||||
|
processor. Where implementation choices have not been finalized, the
|
||||||
|
text uses **PROPOSAL** to mark a design that is being considered but
|
||||||
|
has not been committed. Where the text is an **ASSUMPTION**, evidence
|
||||||
|
is not yet available and the statement must be re-validated before
|
||||||
|
downstream documents depend on it.
|
||||||
|
|
||||||
|
## Scope
|
||||||
|
|
||||||
|
This document covers:
|
||||||
|
|
||||||
|
- The physical memory map and address space layout visible to software
|
||||||
|
running on XH-1 cores.
|
||||||
|
- The relationship between the RISC-V privileged architecture and the
|
||||||
|
memory system, including `satp`, SFENCE.VMA, and the PMP.
|
||||||
|
- The on-chip memory hierarchy: L1, L2, and the proposed shared L3.
|
||||||
|
- The directory-based coherence protocol between L2 slices.
|
||||||
|
- The interconnect that ties cores, caches, and DRAM controllers
|
||||||
|
together.
|
||||||
|
- The DRAM controller and the memory-side bandwidth budget.
|
||||||
|
- I/O, interrupts, and operating system interactions that constrain
|
||||||
|
memory-system design.
|
||||||
|
- Verification and performance considerations specific to a 128-core
|
||||||
|
tiled implementation.
|
||||||
|
|
||||||
|
This document does not define the microarchitecture of an individual
|
||||||
|
core pipeline. That is covered in the core pipeline document.
|
||||||
|
|
||||||
|
## 1. Design Goals and Constraints
|
||||||
|
|
||||||
|
### 1.1 Goals
|
||||||
|
|
||||||
|
1. Provide a coherent, shared-memory programming model to all
|
||||||
|
128 RISC-V harts.
|
||||||
|
2. Sustain per-core bandwidth that is high enough to keep an
|
||||||
|
out-of-order core fed under typical server and HPC workloads.
|
||||||
|
3. Keep average last-level cache (LLC) miss latency within a range
|
||||||
|
that is acceptable for a tiled processor in the same class as
|
||||||
|
published academic and industrial designs.
|
||||||
|
4. Support a Linux-capable memory management model, including virtual
|
||||||
|
memory, page tables, SFENCE.VMA, and standard I/O MMIO patterns.
|
||||||
|
5. Allow verification of coherence, ordering, and interrupts with
|
||||||
|
reasonable engineering effort.
|
||||||
|
|
||||||
|
### 1.2 Constraints
|
||||||
|
|
||||||
|
- The RISC-V Privileged specification, ISA Modules for Sv32, Sv39,
|
||||||
|
Sv48, and Sv57, and the RISC-V IOMMU specification, where used,
|
||||||
|
are normative.
|
||||||
|
- The RISC-V Weak Memory Model (RVWMO) is normative for ordering
|
||||||
|
between hart loads and stores.
|
||||||
|
- The coherence protocol must be implementable on a tiled
|
||||||
|
128-core die in a reasonable process geometry without
|
||||||
|
requiring exotic packaging.
|
||||||
|
- Physical pin bandwidth to DRAM is finite. XH-1 must be
|
||||||
|
budgeted against a realistic number of PHY lanes and channels.
|
||||||
|
Specific numbers are marked as PROPOSAL.
|
||||||
|
|
||||||
|
### 1.3 Non-Goals
|
||||||
|
|
||||||
|
- This document does not specify the instruction set extensions used
|
||||||
|
by XH-1 cores.
|
||||||
|
- It does not specify the precise coherence protocol state machine;
|
||||||
|
it only describes the protocol class.
|
||||||
|
- It does not specify DRAM timing parameters; those are tied to a
|
||||||
|
specific JEDEC standard selected at implementation time.
|
||||||
|
|
||||||
|
## 2. RISC-V Architectural Requirements
|
||||||
|
|
||||||
|
### 2.1 Address Space
|
||||||
|
|
||||||
|
The RISC-V Privileged Architecture (Volume II, Version 20211203) of
|
||||||
|
the RISC-V ISA defines:
|
||||||
|
|
||||||
|
- A canonical physical address space selected by `satp.MODE`.
|
||||||
|
Implementations of `MODE` 0 (Bare), 1 (Sv32), 8 (Sv39), 9 (Sv48),
|
||||||
|
and 10 (Sv57) are mandatory under different profile requirements
|
||||||
|
per the RISC-V Profiles document.
|
||||||
|
- A `satp` CSR that holds the MODE field, the Address Space ID
|
||||||
|
(ASID), and the root page table PPN. Writes to `satp` may become
|
||||||
|
visible to subsequent instruction fetches and loads in a defined
|
||||||
|
order, but require `SFENCE.VMA` for the new translation to be
|
||||||
|
observed by other harts in a coherent manner.
|
||||||
|
- The RISC-V IOMMU specification (riscv-iommu, draft as of the
|
||||||
|
knowledge cutoff) defines DMA translation and the
|
||||||
|
device-context / process-context data structures referenced by
|
||||||
|
I/O devices.
|
||||||
|
|
||||||
|
PROPOSAL: XH-1 implements Sv39 as the base virtual memory mode and
|
||||||
|
Sv48 as an optional mode. Sv57 is not required for the initial
|
||||||
|
silicon. The `M-mode` bare mapping remains available for firmware
|
||||||
|
and boot loaders.
|
||||||
|
|
||||||
|
ASSUMPTION: Software for XH-1 will primarily target Linux. Linux
|
||||||
|
for RISC-V supports Sv39 and Sv48 on 64-bit platforms. Implementing
|
||||||
|
both is consistent with the current mainline kernel.
|
||||||
|
|
||||||
|
### 2.2 Memory Ordering
|
||||||
|
|
||||||
|
The RISC-V Weak Memory Model (RVWMO) is defined in the RISC-V
|
||||||
|
Unprivileged ISA specification, Volume I, and the accompanying
|
||||||
|
RISC-V Memory Consistency Model document. RVWMO provides:
|
||||||
|
|
||||||
|
- Load and store atomicity at aligned word and smaller granularity.
|
||||||
|
- `FENCE`, `FENCE.I`, and `FENCE.VMA` instructions for ordering.
|
||||||
|
- Acquire/release semantics through the `AQ` and `RL` bits on
|
||||||
|
atomic instructions.
|
||||||
|
- The I/O extension ordering rules described in Volume I.
|
||||||
|
|
||||||
|
The memory system must therefore guarantee that:
|
||||||
|
|
||||||
|
- Hart-local loads and stores follow RVWMO rules.
|
||||||
|
- Cross-hart ordering is enforced through cache coherence plus
|
||||||
|
`FENCE` and `FENCE.I`.
|
||||||
|
- `SFENCE.VMA` is correctly ordered with respect to subsequent
|
||||||
|
instruction fetches and loads that depend on the new translation.
|
||||||
|
|
||||||
|
### 2.3 Physical Memory Attribution (PMA)
|
||||||
|
|
||||||
|
PMA regions are defined per-hart and describe the address ranges
|
||||||
|
that support main memory, I/O, idempotent reads, non-idempotent
|
||||||
|
writes, and so on. PMA is a static property of the platform and is
|
||||||
|
not affected by `satp`.
|
||||||
|
|
||||||
|
PROPOSAL: PMA is encoded in a memory-mapped configuration table
|
||||||
|
sampled by each core at reset. The exact format is defined in the
|
||||||
|
platform specification document.
|
||||||
|
|
||||||
|
### 2.4 Physical Memory Protection (PMP)
|
||||||
|
|
||||||
|
The PMP is part of the RISC-V Privileged specification and supports
|
||||||
|
a configurable number of regions per hart. PMP checks occur in
|
||||||
|
M-mode and may be used to isolate firmware and I/O regions from
|
||||||
|
S-mode and U-mode.
|
||||||
|
|
||||||
|
PROPOSAL: Each hart implements the maximum PMP entry count of 64
|
||||||
|
when running in M-mode. S-mode translation uses the PMP only for
|
||||||
|
the M-mode firmware range. U-mode never sees the PMP.
|
||||||
|
|
||||||
|
ASSUMPTION: An M-mode firmware (such as OpenSBI) is in use and
|
||||||
|
relies on PMP to isolate its own memory from the supervisor.
|
||||||
|
|
||||||
|
### 2.5 Cache Management Operations
|
||||||
|
|
||||||
|
The RISC-V Cache Management Operations extension (Zicbom, Zicboz,
|
||||||
|
Zicbop) is defined in Volume I of the RISC-V Unprivileged
|
||||||
|
specification. Operations such as `CBO.INVAL`, `CBO.FLUSH`, and
|
||||||
|
`CBO.CLEAN` are used by software to manage coherence with external
|
||||||
|
agents and DMA.
|
||||||
|
|
||||||
|
PROPOSAL: XH-1 implements Zicbom and Zicboz. Zicbop is not required
|
||||||
|
for the first silicon and is marked as an open question. (See
|
||||||
|
Section 12.)
|
||||||
|
|
||||||
|
## 3. Physical Memory Map
|
||||||
|
|
||||||
|
PROPOSAL: XH-1 exposes a single 48-bit physical address space to
|
||||||
|
software. The map is divided as follows:
|
||||||
|
|
||||||
|
| Region | Start | End | Purpose |
|
||||||
|
|-------------------|---------------|---------------|----------------------------------|
|
||||||
|
| DRAM | 0x0000_0000_0000 | 0x007F_FFFF_FFFF | Main memory, up to 32 GiB. |
|
||||||
|
| Reserved | 0x0080_0000_0000 | 0x00FF_FFFF_FFFF | Hole for future expansion. |
|
||||||
|
| High DRAM | 0x0100_0000_0000 | 0x017F_FFFF_FFFF | Optional second DRAM range. |
|
||||||
|
| PCIe MMIO | 0x4000_0000_0000 | 0x40FF_FFFF_FFFF | PCIe ECAM. |
|
||||||
|
| PCIe prefetchable | 0x6000_0000_0000 | 0x60FF_FFFF_FFFF | PCIe MMIO64 prefetchable. |
|
||||||
|
| On-chip I/O | 0x7000_0000_0000 | 0x7000_0FFF_FFFF | Interrupt controller, IOMMU, UART, boot ROM. |
|
||||||
|
| Reserved | 0x8000_0000_0000 | 0xFFFF_FFFF_FFFF | Hole for future use. |
|
||||||
|
|
||||||
|
The exact upper bound depends on the DRAM capacity and number of
|
||||||
|
DIMMs supported. The values above are PROPOSAL values used to
|
||||||
|
illustrate the layout.
|
||||||
|
|
||||||
|
ASSUMPTION: DRAM density is between 16 GiB and 1 TiB. The lower
|
||||||
|
bound reflects realistic server configurations for a 128-core chip
|
||||||
|
in 2026; the upper bound is set by the 48-bit physical address
|
||||||
|
limit and by the cost of the memory controller area.
|
||||||
|
|
||||||
|
## 4. Tile Organization
|
||||||
|
|
||||||
|
PROPOSAL: XH-1 is organized as 16 tiles, each containing 8 cores.
|
||||||
|
Each tile includes:
|
||||||
|
|
||||||
|
- 8 application cores, each with a private L1 instruction cache
|
||||||
|
and a private L1 data cache.
|
||||||
|
- A shared L2 cache, sliced to support one bank per core.
|
||||||
|
- A directory controller for the L3 coherence protocol.
|
||||||
|
- A shared floating-point and vector unit cluster (PROPOSAL: see
|
||||||
|
the vector and floating-point research document for detail).
|
||||||
|
- A tile-local crossbar that connects cores, the L2, the
|
||||||
|
directory controller, and the tile's connection to the global
|
||||||
|
interconnect.
|
||||||
|
|
||||||
|
The 16 tiles connect through a global on-die interconnect. The
|
||||||
|
interconnect design is described in Section 8.
|
||||||
|
|
||||||
|
ASSUMPTION: A tile is sized to fit in a 3 mm by 3 mm die region in
|
||||||
|
a 5 nm class process. The exact floorplan is a downstream
|
||||||
|
implementation concern.
|
||||||
|
|
||||||
|
## 5. L1 Caches
|
||||||
|
|
||||||
|
### 5.1 Sizing
|
||||||
|
|
||||||
|
PROPOSAL: The L1 instruction cache is 32 KiB, 8-way set-associative,
|
||||||
|
with 64 B lines. The L1 data cache is 32 KiB, 8-way set-associative,
|
||||||
|
with 64 B lines. Both caches are virtually indexed, physically
|
||||||
|
tagged (VIPT), which constrains the index to a number of bits that
|
||||||
|
is at most equal to the page offset.
|
||||||
|
|
||||||
|
For a 32 KiB, 8-way cache with 64 B lines, the index is 6 bits, the
|
||||||
|
offset is 6 bits, and the tag is 52 bits. With 4 KiB pages, the
|
||||||
|
page offset is 12 bits, so the index is fully contained in the page
|
||||||
|
offset and VIPT is alias-free.
|
||||||
|
|
||||||
|
ASSUMPTION: 4 KiB pages are the common case. Larger pages such as
|
||||||
|
2 MiB and 1 GiB reduce aliasing pressure even further, so VIPT
|
||||||
|
remains alias-free for the supported page sizes.
|
||||||
|
|
||||||
|
### 5.2 Replacement and Prefetch
|
||||||
|
|
||||||
|
PROPOSAL: L1 uses a pseudo-LRU replacement policy. Hardware
|
||||||
|
prefetchers include:
|
||||||
|
|
||||||
|
- A next-line prefetcher on the instruction side.
|
||||||
|
- A stride prefetcher on the data side.
|
||||||
|
- A stream prefetcher on the data side for sequential access
|
||||||
|
patterns.
|
||||||
|
|
||||||
|
The prefetch degree and aggressiveness are tunable through
|
||||||
|
machine-mode configuration registers.
|
||||||
|
|
||||||
|
### 5.3 MSHRs and Write Combining
|
||||||
|
|
||||||
|
PROPOSAL: Each L1 data cache has 16 miss status handling
|
||||||
|
registers (MSHRs), supporting up to 16 outstanding misses. The
|
||||||
|
write buffer holds up to 16 entries. Combining of partial line
|
||||||
|
writes is performed in the write buffer.
|
||||||
|
|
||||||
|
ASSUMPTION: 16 MSHRs is sufficient for the workloads of interest.
|
||||||
|
A sensitivity study is required before the L1 MSHR count is
|
||||||
|
finalized. (See Section 12.)
|
||||||
|
|
||||||
|
## 6. L2 Cache
|
||||||
|
|
||||||
|
PROPOSAL: Each tile contains an 8 MiB L2 cache, organized as 8
|
||||||
|
banks, one per core. Each bank is 1 MiB, 16-way set-associative,
|
||||||
|
with 64 B lines. The L2 is inclusive of the L1 and acts as the
|
||||||
|
coherence point for the L1 caches on the tile.
|
||||||
|
|
||||||
|
Total L2 capacity across the die is 128 MiB. This number is a
|
||||||
|
PROPOSAL based on the die area budget for caches in the target
|
||||||
|
process node; it is not yet finalized.
|
||||||
|
|
||||||
|
### 6.1 Inclusion
|
||||||
|
|
||||||
|
PROPOSAL: The L2 is strictly inclusive of the L1 caches on the
|
||||||
|
same tile. A coherence invalidation from the L3 directory causes
|
||||||
|
the L2 to invalidate the corresponding L1 line.
|
||||||
|
|
||||||
|
ASSUMPTION: Strict inclusion simplifies the coherence protocol at
|
||||||
|
the cost of additional L2 traffic for victims. A non-inclusive
|
||||||
|
L2 is an open question. (See Section 12.)
|
||||||
|
|
||||||
|
### 6.2 Latency
|
||||||
|
|
||||||
|
ASSUMPTION: The L2 hit latency is between 12 and 20 cycles. The
|
||||||
|
exact number depends on the physical layout of the cache and the
|
||||||
|
core. This is an open implementation question.
|
||||||
|
|
||||||
|
## 7. L3 Cache and Coherence Protocol
|
||||||
|
|
||||||
|
### 7.1 Class
|
||||||
|
|
||||||
|
PROPOSAL: XH-1 uses a directory-based MESI coherence protocol at
|
||||||
|
the L3 level. The directory is distributed across tiles, with one
|
||||||
|
directory controller per tile. Each line in the L3 is owned by
|
||||||
|
exactly one directory, identified by the home tile.
|
||||||
|
|
||||||
|
The protocol class is "distributed directory MESI with distributed
|
||||||
|
shared L3". This class is the same as the protocol class used in
|
||||||
|
several published academic and industrial designs, including the
|
||||||
|
MIT Alewife, the SGI Origin, the Oracle/Berkeley SPARC M7, and the
|
||||||
|
Fujitsu A64FX.
|
||||||
|
|
||||||
|
### 7.2 States
|
||||||
|
|
||||||
|
The four stable MESI states are:
|
||||||
|
|
||||||
|
- **M**odified: the line is dirty and held by a single cache.
|
||||||
|
- **E**xclusive: the line is clean and held by a single cache.
|
||||||
|
- **S**hared: the line is clean and may be held by multiple
|
||||||
|
caches.
|
||||||
|
- **I**nvalid: the line is not present.
|
||||||
|
|
||||||
|
PROPOSAL: A fifth "Owned" state is not used in the base protocol
|
||||||
|
to keep the directory and transient state space simple. A
|
||||||
|
directory entry tracks the presence bits for each tile.
|
||||||
|
|
||||||
|
### 7.3 Directory Encoding
|
||||||
|
|
||||||
|
PROPOSAL: The directory uses a coarse vector of presence bits,
|
||||||
|
one per tile, plus a "has-modified" bit. For 16 tiles this is 17
|
||||||
|
bits per line. The directory is stored in a dedicated SRAM
|
||||||
|
structure, sized at approximately 12.5 MiB for a 64 MiB L3 with
|
||||||
|
4x replication of the directory across the on-die fabric to
|
||||||
|
support transient races.
|
||||||
|
|
||||||
|
ASSUMPTION: The 4x replication of the directory entry is a
|
||||||
|
simplified model. The exact replication strategy is a downstream
|
||||||
|
implementation question. (See Section 12.)
|
||||||
|
|
||||||
|
### 7.4 Directory Operations
|
||||||
|
|
||||||
|
PROPOSAL: Coherence operations follow a request-response pattern:
|
||||||
|
|
||||||
|
1. A core issues a read or write to its L1.
|
||||||
|
2. On L1 miss, the request is forwarded to the L2.
|
||||||
|
3. On L2 miss, the L2 sends a coherence request to the home tile's
|
||||||
|
directory.
|
||||||
|
4. The directory invalidates any other copies, forwards the data
|
||||||
|
to the requester, and updates the presence vector.
|
||||||
|
5. The directory replies to the requester, which completes the
|
||||||
|
L1 fill.
|
||||||
|
|
||||||
|
Writebacks from L1 invalidate the corresponding L2 line and
|
||||||
|
update the directory.
|
||||||
|
|
||||||
|
### 7.5 Memory Consistency Ordering
|
||||||
|
|
||||||
|
The protocol preserves the appearance of a single global ordering
|
||||||
|
of coherence transactions as seen by the cores. RISC-V `FENCE`
|
||||||
|
ordering is enforced by the L2 controller before forwarding a
|
||||||
|
request to the directory.
|
||||||
|
|
||||||
|
PROPOSAL: Store atomic instructions (`AMO*` and `LR/SC`) are
|
||||||
|
handled by the L2 controller. The L2 acts as the ordering point
|
||||||
|
for atomic operations and serializes the operation with respect
|
||||||
|
to other requests to the same address.
|
||||||
|
|
||||||
|
### 7.6 Forwarding and Snoop Filtering
|
||||||
|
|
||||||
|
PROPOSAL: The directory acts as a snoop filter. A coherence
|
||||||
|
request that finds the directory in the "S" state with no other
|
||||||
|
sharer can supply the data from the L2 without further snoop
|
||||||
|
traffic.
|
||||||
|
|
||||||
|
## 8. On-Die Interconnect
|
||||||
|
|
||||||
|
### 8.1 Topology
|
||||||
|
|
||||||
|
PROPOSAL: The on-die interconnect is a 2D mesh of 4 by 4 tiles.
|
||||||
|
Each tile has a router that connects to its four nearest neighbors.
|
||||||
|
The mesh supports a 512-bit link per direction, clocked at a
|
||||||
|
fraction of the core clock.
|
||||||
|
|
||||||
|
ASSUMPTION: A 2D mesh is sufficient for the bandwidth requirements
|
||||||
|
of a 128-core die. A 2D torus is an alternative discussed in
|
||||||
|
Section 12.
|
||||||
|
|
||||||
|
### 8.2 Routing
|
||||||
|
|
||||||
|
PROPOSAL: The mesh uses deterministic X-Y routing. Adaptive routing
|
||||||
|
is not used in the first silicon to simplify verification.
|
||||||
|
|
||||||
|
### 8.3 Link Bandwidth
|
||||||
|
|
||||||
|
PROPOSAL: Each link provides 64 B per cycle of useful payload at
|
||||||
|
the router clock. With 16 tiles and 4 directions per router
|
||||||
|
(excluding the boundary cases), the bisection bandwidth of the
|
||||||
|
mesh is 2 links by 4 tiles by 64 B per cycle, or 512 B per cycle.
|
||||||
|
At a 1.5 GHz router clock this is 768 GB/s of bisection
|
||||||
|
bandwidth.
|
||||||
|
|
||||||
|
ASSUMPTION: The 1.5 GHz router clock is feasible in the target
|
||||||
|
process node. This is an open implementation question.
|
||||||
|
|
||||||
|
### 8.4 Flow Control
|
||||||
|
|
||||||
|
PROPOSAL: The interconnect uses credit-based flow control on each
|
||||||
|
link. Each virtual channel maintains a count of available credits
|
||||||
|
and stalls a flit only when the downstream router is unable to
|
||||||
|
accept it.
|
||||||
|
|
||||||
|
## 9. DRAM Controller
|
||||||
|
|
||||||
|
### 9.1 Topology
|
||||||
|
|
||||||
|
PROPOSAL: XH-1 has 8 DRAM controllers, distributed on the die. Two
|
||||||
|
controllers are placed near the center of each edge. The
|
||||||
|
controllers connect to the mesh with high-bandwidth links and to
|
||||||
|
external PHYs through the package substrate.
|
||||||
|
|
||||||
|
### 9.2 Channels and DIMMs
|
||||||
|
|
||||||
|
PROPOSAL: Each controller supports 2 channels, for a total of 16
|
||||||
|
channels. Each channel supports a single DDR5 DIMM in the first
|
||||||
|
silicon.
|
||||||
|
|
||||||
|
### 9.3 Bandwidth
|
||||||
|
|
||||||
|
ASSUMPTION: At DDR5-6400, a 64-bit channel provides 51.2 GB/s
|
||||||
|
peak. With 16 channels, the peak DRAM bandwidth is 819 GB/s.
|
||||||
|
|
||||||
|
PROPOSAL: The memory system is designed to sustain an average
|
||||||
|
L3 miss rate of 0.05 misses per instruction across the 128 cores
|
||||||
|
at the rated clock. This is consistent with a 1.0 GHz core clock
|
||||||
|
and an IPC of 1.0, giving 128 GIPS, and a DRAM traffic of 0.05
|
||||||
|
misses by 128 GIPS by 64 B per miss, or 410 GB/s. This is below
|
||||||
|
the 819 GB/s peak, leaving headroom for OS and I/O traffic.
|
||||||
|
|
||||||
|
### 9.4 Quality of Service
|
||||||
|
|
||||||
|
PROPOSAL: The DRAM controller supports 4 QoS classes. A
|
||||||
|
configuration register in the platform specification selects the
|
||||||
|
class for each request source (core, device, I/O).
|
||||||
|
|
||||||
|
ASSUMPTION: A weighted round-robin scheduler per channel is
|
||||||
|
sufficient for the first silicon. A more sophisticated scheduler
|
||||||
|
is an open question.
|
||||||
|
|
||||||
|
## 10. Operating System and Firmware
|
||||||
|
|
||||||
|
### 10.1 Linux Memory Model
|
||||||
|
|
||||||
|
PROPOSAL: XH-1 supports mainline Linux for RISC-V. The relevant
|
||||||
|
kernel features are:
|
||||||
|
|
||||||
|
- Sv39 and Sv48 virtual memory.
|
||||||
|
- NUMA-aware scheduling, with one NUMA node per tile.
|
||||||
|
- Standard cache management operations through Zicbom.
|
||||||
|
- PCIe MMU through the optional RISC-V IOMMU.
|
||||||
|
|
||||||
|
ASSUMPTION: Linux for RISC-V provides all of the above features
|
||||||
|
on a 64-bit SMP platform. This assumption should be re-validated
|
||||||
|
against the kernel version targeted for the first silicon.
|
||||||
|
|
||||||
|
### 10.2 NUMA Effects
|
||||||
|
|
||||||
|
PROPOSAL: Each tile is exposed as a NUMA node to Linux. The
|
||||||
|
operating system scheduler places a task on the node closest to
|
||||||
|
its memory. The boot firmware sets up the device tree and ACPI
|
||||||
|
table accordingly.
|
||||||
|
|
||||||
|
### 10.3 Firmware and PMP
|
||||||
|
|
||||||
|
PROPOSAL: OpenSBI is the M-mode firmware. OpenSBI uses PMP to
|
||||||
|
isolate its own memory from S-mode. The SBI interface is the
|
||||||
|
primary interface between S-mode and M-mode.
|
||||||
|
|
||||||
|
### 10.4 IOMMU and DMA
|
||||||
|
|
||||||
|
PROPOSAL: A RISC-V IOMMU is included on the die. The IOMMU
|
||||||
|
translates device DMA addresses using the same page table
|
||||||
|
formats as the CPU. PCIe devices that support ATS bypass the
|
||||||
|
IOMMU on a per-transaction basis.
|
||||||
|
|
||||||
|
ASSUMPTION: The RISC-V IOMMU specification is sufficiently
|
||||||
|
mature at the time of tape-out. If not, the IOMMU is replaced
|
||||||
|
by a software bounce-buffering scheme in the first silicon.
|
||||||
|
|
||||||
|
## 11. Interrupts and Memory System Interactions
|
||||||
|
|
||||||
|
### 11.1 Interrupt Delivery
|
||||||
|
|
||||||
|
PROPOSAL: A platform-level interrupt controller (PLIC) is mapped
|
||||||
|
at the on-chip I/O region. The PLIC supports 256 interrupt
|
||||||
|
sources and 128 interrupt targets, one per hart. The PLIC
|
||||||
|
accesses are strongly ordered with respect to the coherence
|
||||||
|
protocol.
|
||||||
|
|
||||||
|
### 11.2 Interrupt and Coherence
|
||||||
|
|
||||||
|
ASSUMPTION: The interrupt controller is on the coherence domain.
|
||||||
|
An interrupt write from a device is visible to the PLIC
|
||||||
|
registers, and the subsequent PLIC read by the core is ordered
|
||||||
|
through the coherence protocol.
|
||||||
|
|
||||||
|
### 11.3 Interrupt Latency
|
||||||
|
|
||||||
|
ASSUMPTION: Worst-case interrupt latency through the PLIC is
|
||||||
|
under 1 microsecond at the rated clock, including coherence
|
||||||
|
round-trip and pipeline drain. The actual value is workload
|
||||||
|
dependent and is left for the performance document.
|
||||||
|
|
||||||
|
## 12. Open Questions and Tradeoffs
|
||||||
|
|
||||||
|
The following questions are unresolved at the time of writing.
|
||||||
|
Each is a candidate for further study before the architecture is
|
||||||
|
frozen.
|
||||||
|
|
||||||
|
1. **Inclusive vs. non-inclusive L2.** Strict inclusion
|
||||||
|
simplifies coherence but may waste L2 capacity. A
|
||||||
|
non-inclusive victim cache design is an alternative.
|
||||||
|
2. **L3 size.** A 64 MiB L3 is the proposal. A 32 MiB or 128 MiB
|
||||||
|
L3 are alternatives; each changes the area budget and miss
|
||||||
|
rate.
|
||||||
|
3. **Directory organization.** A coarse vector with replication
|
||||||
|
is the proposal. A hierarchical or sparse directory is an
|
||||||
|
alternative that may reduce area at the cost of latency.
|
||||||
|
4. **Mesh vs. torus.** A 2D mesh is the proposal. A 2D torus
|
||||||
|
doubles the bisection bandwidth and may be required for very
|
||||||
|
high cache miss rates. A ring or concentrated mesh is a
|
||||||
|
third alternative.
|
||||||
|
5. **Number of memory channels.** 16 channels is the proposal.
|
||||||
|
8 or 24 channels are alternatives.
|
||||||
|
6. **Zicbop implementation.** Cache block prefetch operations are
|
||||||
|
optional in the RISC-V Cache Management Operations
|
||||||
|
specification. Their cost and benefit on XH-1 are not yet
|
||||||
|
quantified.
|
||||||
|
7. **L1 MSHR count.** 16 MSHRs per L1 is the proposal. The
|
||||||
|
workload-dependent optimum is not known.
|
||||||
|
8. **IOMMU maturity.** The RISC-V IOMMU specification was a
|
||||||
|
draft at the knowledge cutoff date. The XH-1 implementation
|
||||||
|
may need to wait for a stable revision.
|
||||||
|
9. **Coherence at the L1 level.** The current proposal
|
||||||
|
collapses L1 coherence to the L2. An alternative is a
|
||||||
|
broadcast snoop between L1s on the same tile. This trades
|
||||||
|
L1 bandwidth for L2 traffic.
|
||||||
|
10. **SMT support.** Each core may support a small number of
|
||||||
|
hardware threads. SMT changes the coherence and memory
|
||||||
|
system pressure per core. The number of threads per core
|
||||||
|
is not yet decided.
|
||||||
|
11. **Vector and matrix unit interactions.** The vector unit
|
||||||
|
performs high-bandwidth loads and stores. The memory
|
||||||
|
system must be able to support sustained vector load
|
||||||
|
bandwidth at the rated clock. The vector unit design
|
||||||
|
appears in a separate document.
|
||||||
|
12. **Power management.** Clock gating, power gating, and DVFS
|
||||||
|
interact with the memory system. The current document does
|
||||||
|
not yet address these interactions.
|
||||||
|
|
||||||
|
## 13. Scalability Analysis for 128 Cores
|
||||||
|
|
||||||
|
### 13.1 Coherence Scalability
|
||||||
|
|
||||||
|
A broadcast snoop protocol scales as O(N^2) in the worst case,
|
||||||
|
where N is the number of cores. For 128 cores this is 16,384
|
||||||
|
worst-case broadcast pairs, which is generally not viable on a
|
||||||
|
single die. The directory protocol scales as O(N) in the worst
|
||||||
|
case for invalidations, and the directory size grows linearly
|
||||||
|
with the number of tiles.
|
||||||
|
|
||||||
|
PROPOSAL: A directory protocol is required at 128 cores. A
|
||||||
|
snoopy protocol is not viable.
|
||||||
|
|
||||||
|
### 13.2 Interconnect Scalability
|
||||||
|
|
||||||
|
A 2D mesh of 16 tiles has bisection bandwidth proportional to
|
||||||
|
the number of tiles per row. The 4 by 4 mesh provides 4 links of
|
||||||
|
bisection, which is sufficient for the projected 410 GB/s of
|
||||||
|
average DRAM traffic in Section 9.3.
|
||||||
|
|
||||||
|
ASSUMPTION: Average traffic does not exceed 70 percent of the
|
||||||
|
peak bisection. A sensitivity analysis is required. (See
|
||||||
|
Section 12.)
|
||||||
|
|
||||||
|
### 13.3 Directory Storage Scalability
|
||||||
|
|
||||||
|
The directory entry is 17 bits per L3 line. For a 64 MiB L3 with
|
||||||
|
64 B lines, this is 2^20 lines by 17 bits, or 2.2 MiB per
|
||||||
|
directory. With 4x replication, the on-die directory storage is
|
||||||
|
approximately 12.5 MiB. This is a small fraction of the L3
|
||||||
|
storage.
|
||||||
|
|
||||||
|
### 13.4 Operating System Scalability
|
||||||
|
|
||||||
|
Linux scales to hundreds of cores. The memory system must
|
||||||
|
expose the topology through ACPI and the device tree so that the
|
||||||
|
kernel NUMA scheduler can make informed decisions. The
|
||||||
|
128-core XH-1 is below the upper bound of common server Linux
|
||||||
|
configurations as of the knowledge cutoff.
|
||||||
|
|
||||||
|
### 13.5 Verification Scalability
|
||||||
|
|
||||||
|
Coherence verification grows in complexity with the number of
|
||||||
|
cores. PROPOSAL: a combination of formal verification of the
|
||||||
|
directory protocol state machine, simulation-based testing, and
|
||||||
|
random instruction generation is used. The verification
|
||||||
|
strategy is described in the verification document.
|
||||||
|
|
||||||
|
## 14. Verification
|
||||||
|
|
||||||
|
### 14.1 Cache Coherence
|
||||||
|
|
||||||
|
PROPOSAL: The directory protocol state machine is verified using
|
||||||
|
a model checker against the MESI invariants. The invariants
|
||||||
|
include:
|
||||||
|
|
||||||
|
- Single writer per line.
|
||||||
|
- No read of a line in the Invalid state.
|
||||||
|
- Correct invalidation on a write to a shared line.
|
||||||
|
|
||||||
|
### 14.2 Memory Consistency
|
||||||
|
|
||||||
|
PROPOSAL: A RISC-V memory consistency model litmus test suite
|
||||||
|
is run against the system, using both RTL simulation and a
|
||||||
|
software reference model.
|
||||||
|
|
||||||
|
### 14.3 Performance and Stress Testing
|
||||||
|
|
||||||
|
PROPOSAL: Synthetic traffic patterns, including uniform random,
|
||||||
|
hotspot, and adversarial patterns, are used to verify that the
|
||||||
|
memory system behaves correctly under stress. Realistic
|
||||||
|
workloads, including SPEC CPU 2017, PARSEC, and selected
|
||||||
|
GROMACS and LAMMPS runs, are used for performance validation.
|
||||||
|
|
||||||
|
ASSUMPTION: The listed benchmarks are representative of the
|
||||||
|
target workloads. This assumption should be re-validated by the
|
||||||
|
performance team.
|
||||||
|
|
||||||
|
## 15. Interaction With Other Subsystems
|
||||||
|
|
||||||
|
### 15.1 Pipeline
|
||||||
|
|
||||||
|
The L1 cache is in the core pipeline. The L1 hit latency is
|
||||||
|
visible in the branch resolution and the load-to-use latency.
|
||||||
|
PROPOSAL: The L1 hit latency is 4 cycles for a load-to-use on a
|
||||||
|
hit. A miss to the L2 adds between 12 and 20 cycles, depending
|
||||||
|
on the L2 hit and the core-to-L2 routing.
|
||||||
|
|
||||||
|
### 15.2 Cache Hierarchy
|
||||||
|
|
||||||
|
The L1, L2, and L3 form a hierarchy. The L2 is inclusive of the
|
||||||
|
L1. The L3 is the coherence point. The L3 is non-inclusive of
|
||||||
|
the L2 by default.
|
||||||
|
|
||||||
|
### 15.3 Memory System
|
||||||
|
|
||||||
|
The DRAM controllers are described in Section 9. The on-die
|
||||||
|
interconnect carries both coherence and data traffic.
|
||||||
|
|
||||||
|
### 15.4 Interconnect
|
||||||
|
|
||||||
|
The on-die interconnect is described in Section 8. Coherence
|
||||||
|
and data traffic share the same physical links.
|
||||||
|
|
||||||
|
PROPOSAL: Virtual channels separate the request, response, and
|
||||||
|
data traffic. The number of virtual channels per link is 4.
|
||||||
|
|
||||||
|
### 15.5 Coherence
|
||||||
|
|
||||||
|
Coherence is described in Section 7. The protocol class is
|
||||||
|
distributed directory MESI with distributed shared L3.
|
||||||
|
|
||||||
|
### 15.6 Interrupts
|
||||||
|
|
||||||
|
Interrupts are described in Section 11. The PLIC is on the
|
||||||
|
coherence domain.
|
||||||
|
|
||||||
|
### 15.7 Operating System
|
||||||
|
|
||||||
|
The operating system interactions are described in Section 10.
|
||||||
|
The OS is Linux for RISC-V. The firmware is OpenSBI.
|
||||||
|
|
||||||
|
### 15.8 Performance
|
||||||
|
|
||||||
|
The performance analysis is described in Section 13 and is
|
||||||
|
detailed in a separate performance document. Key memory-system
|
||||||
|
performance metrics include:
|
||||||
|
|
||||||
|
- L1 hit rate
|
||||||
|
- L2 hit rate
|
||||||
|
- L3 hit rate
|
||||||
|
- DRAM bandwidth utilization
|
||||||
|
- Average and tail DRAM latency
|
||||||
|
- Interconnect average and tail latency
|
||||||
|
|
||||||
|
## 16. Recommendations
|
||||||
|
|
||||||
|
1. Implement a directory-based MESI coherence protocol with a
|
||||||
|
distributed shared L3. Snoopy protocols are not viable at
|
||||||
|
128 cores.
|
||||||
|
2. Use a 2D mesh on-die interconnect. The 2D mesh is sufficient
|
||||||
|
for the projected traffic and is simpler to verify than a
|
||||||
|
torus.
|
||||||
|
3. Implement Zicbom and Zicboz cache management operations.
|
||||||
|
These are required for Linux on RISC-V with mainline
|
||||||
|
kernels.
|
||||||
|
4. Expose each tile as a NUMA node to the operating system. This
|
||||||
|
allows the kernel scheduler to make topology-aware decisions.
|
||||||
|
5. Defer the final decision on the L3 size, the number of DRAM
|
||||||
|
channels, and the directory organization to the
|
||||||
|
implementation phase, after the sensitivity studies listed
|
||||||
|
in Section 12 are completed.
|
||||||
|
6. Defer the implementation of Zicbop until a workload study
|
||||||
|
shows a clear benefit.
|
||||||
|
|
||||||
|
## 17. References
|
||||||
|
|
||||||
|
The following standards and references are cited in this
|
||||||
|
document:
|
||||||
|
|
||||||
|
- RISC-V Unprivileged ISA Specification, Volume I, Document
|
||||||
|
Version 20191213. Section on the Weak Memory Model and the
|
||||||
|
Cache Management Operations extensions.
|
||||||
|
- RISC-V Privileged ISA Specification, Volume II, Document
|
||||||
|
Version 20211203. Sections on `satp`, SFENCE.VMA, PMP, and
|
||||||
|
interrupt handling.
|
||||||
|
- RISC-V Profiles, Document Version 20230914. Application and
|
||||||
|
platform profile requirements.
|
||||||
|
- RISC-V IOMMU Specification, draft as of the knowledge cutoff.
|
||||||
|
- JEDEC DDR5 Standard, JESD79-5.
|
||||||
|
- JEDEC LPDDR5 Standard, JESD209-5A.
|
||||||
|
- Hennessy and Patterson, "Computer Architecture: A Quantitative
|
||||||
|
Approach", 6th edition. Sections on cache coherence and
|
||||||
|
directory protocols.
|
||||||
|
- Culler, Singh, and Gupta, "Parallel Computer Architecture: A
|
||||||
|
Hardware/Software Approach". Chapter on scalable cache
|
||||||
|
coherence.
|
||||||
|
- M. M. K. Martin, M. D. Hill, and D. A. Wood, "Token
|
||||||
|
Coherence", University of Wisconsin Computer Sciences
|
||||||
|
Technical Report 1621, 2003. Discusses a directory-based
|
||||||
|
protocol with explicit invalidation tokens.
|
||||||
|
- P. Kongetira, K. Aingaran, and K. Olukotun, "Niagara: A
|
||||||
|
32-Way Multithreaded SPARC Processor", IEEE Micro, March 2005.
|
||||||
|
Discusses a directory-coherent CMP design.
|
||||||
|
- R. Azimi et al., "Integration Challenges and Tradeoffs for
|
||||||
|
Tiled Many-Core Processors", Proceedings of the 22nd
|
||||||
|
International Conference on Computer Design, 2004.
|
||||||
|
- P. Hammarlund et al., "Haswell: The Fourth-Generation Intel
|
||||||
|
Core Processor", IEEE Micro, March 2014. Discusses a
|
||||||
|
ring-based interconnect in a tiled CMP.
|
||||||
|
- J. L. Henning, "SPEC CPU2006 Benchmark Descriptions", ACM
|
||||||
|
SIGARCH Computer Architecture News, September 2006.
|
||||||
|
- OpenSBI documentation, https://github.com/riscv-software-src/opensbi.
|
||||||
|
|
||||||
|
## 18. Document History
|
||||||
|
|
||||||
|
- Draft 0.1: initial outline.
|
||||||
|
- Draft 0.2: populated physical memory map, coherence protocol
|
||||||
|
class, and L1/L2/L3 sizing as proposals.
|
||||||
|
- Draft 0.3: added open questions and recommendations.
|
||||||
|
```
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
research/05-memory/memory-architecture.md
|
||||||
@@ -0,0 +1,25 @@
|
|||||||
|
The document describes a high-performance, many-core RISC-V system (XH-1) with 128 cores. The technical specifications provided are consistent with modern high-performance computing (HPC) architecture and RISC-V standards.
|
||||||
|
|
||||||
|
### Technical Review:
|
||||||
|
|
||||||
|
1. **Scalability & Interconnect:** The choice of a **2D mesh** for 128 cores is a standard and correct architectural decision. At this scale, a ring or crossbar is physically and logically impractical. The use of **virtual channels** to separate request, response, and data traffic is critical to prevent head-of-line blocking and protocol deadlocks in a directory-based system.
|
||||||
|
2. **Cache Coherence:** The transition from snooping to a **distributed directory MESI protocol** is mandatory for 128 cores. The inclusion of **Zicbom and Zicboz** is a critical requirement for RISC-V Linux compatibility, as these allow the OS to manage cache lines and memory attributes correctly.
|
||||||
|
3. **Memory Consistency:** The proposal to use a **litmus test suite** for RISC-V memory consistency is the industry standard for verifying that the hardware correctly implements the Weak Memory Model (WMM) required by the RISC-V ISA.
|
||||||
|
4. **Hierarchy:** The L2 being inclusive of L1 and the L3 being the coherence point is a standard design that simplifies the coherence logic at the L3 level.
|
||||||
|
5. **OS/Software:** Exposing the topology as **NUMA nodes** is the correct way to handle a 128-core system in Linux, ensuring the scheduler can optimize for local memory access.
|
||||||
|
6. **Verification:** The multi-pronged approach (Formal for state machines, Simulation for consistency, and Stress/Synthetic for performance) is a robust verification strategy for a complex SoC.
|
||||||
|
|
||||||
|
### Minor Observations:
|
||||||
|
- The decision to defer Zicbop is prudent, as it is a more complex extension for "buffer" management that may not be necessary for initial silicon.
|
||||||
|
- The distinction between "inclusive" and "non-inclusive" L3 is correctly identified as a key design choice for the implementation phase.
|
||||||
|
|
||||||
|
The document is technically sound, follows RISC-V standards, and addresses the specific challenges of scaling to 128 cores.
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"verdict": "PASS",
|
||||||
|
"confidence": "HIGH",
|
||||||
|
"issues": [],
|
||||||
|
"required_fixes": []
|
||||||
|
}
|
||||||
|
```
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
2026-08-26T13:20:19Z research/05-memory/memory-bandwidth.md 1 research success
|
||||||
@@ -0,0 +1,533 @@
|
|||||||
|
# Memory Bandwidth
|
||||||
|
|
||||||
|
## Scope
|
||||||
|
|
||||||
|
This document investigates memory bandwidth requirements, limits, and
|
||||||
|
implementation strategies for the XH-1 processor, a proposed 128-core
|
||||||
|
RISC-V design. The discussion covers off-chip DRAM bandwidth,
|
||||||
|
on-chip interconnect bandwidth, cache hierarchy bandwidth, and the
|
||||||
|
interactions between bandwidth, coherence traffic, and the memory
|
||||||
|
system.
|
||||||
|
|
||||||
|
This document is part of the XH-1 research series (area 05-memory).
|
||||||
|
Statements about XH-1 implementation are proposals unless explicitly
|
||||||
|
marked as decided.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. Problem Statement
|
||||||
|
|
||||||
|
A 128-core processor generates substantially more memory traffic than
|
||||||
|
single-core or small-count-core designs. Even if each core issues a
|
||||||
|
modest number of outstanding memory requests, the aggregate demand can
|
||||||
|
quickly exceed the bandwidth available from a realistic DRAM subsystem.
|
||||||
|
|
||||||
|
The XH-1 design must therefore answer several related questions:
|
||||||
|
|
||||||
|
- How much DRAM bandwidth is physically available on the target package?
|
||||||
|
- How much of that bandwidth is consumed by coherence traffic?
|
||||||
|
- How much remains for useful application work?
|
||||||
|
- How is that remaining bandwidth shared between cores without
|
||||||
|
starvation or severe unfairness?
|
||||||
|
- Where do bandwidth bottlenecks arise in the on-chip interconnect,
|
||||||
|
cache hierarchy, and memory controllers?
|
||||||
|
|
||||||
|
The remainder of this document analyses each of these questions.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. Background
|
||||||
|
|
||||||
|
### 2.1 DRAM Bandwidth Fundamentals
|
||||||
|
|
||||||
|
Modern DDR systems provide bandwidth through wide channels and high
|
||||||
|
transfer rates:
|
||||||
|
|
||||||
|
- DDR5 devices use 32-bit sub-channels per module rank, with
|
||||||
|
transfer rates up to 8400 MT/s in JEDEC-published modules.
|
||||||
|
- A single DDR5 channel (64 bits) at 6400 MT/s delivers approximately
|
||||||
|
51.2 GB/s peak; a dual-channel configuration roughly doubles this.
|
||||||
|
- HBM3 stacks deliver substantially higher bandwidth per package
|
||||||
|
(around 819 GB/s for a 6-high stack per HBM3 specification).
|
||||||
|
- HBM3E extends this further, with published per-stack figures in the
|
||||||
|
1 TB/s range for selected devices.
|
||||||
|
|
||||||
|
The XH-1 design has not fixed a memory technology. Both DDR5 and
|
||||||
|
HBM3-class options are considered below as proposals.
|
||||||
|
|
||||||
|
### 2.2 RISC-V Memory Model
|
||||||
|
|
||||||
|
RISC-V defines a weak memory ordering model, the RISC-V Weak Memory
|
||||||
|
Ordering (RVWMO), in the architecture specification. RVWMO permits a
|
||||||
|
wide range of reorderings and is similar in spirit to other modern
|
||||||
|
weakly ordered architectures.
|
||||||
|
|
||||||
|
Key consequences for bandwidth analysis:
|
||||||
|
|
||||||
|
- Loads may bypass earlier stores under default rules when addresses
|
||||||
|
differ.
|
||||||
|
- Coherence is required only at a single total order per address
|
||||||
|
(the "coherence order" axiom).
|
||||||
|
- Fence instructions (`fence`, `fence.i`, `fence.vma`) are required
|
||||||
|
to constrain ordering where software demands it.
|
||||||
|
|
||||||
|
A weakly ordered model does not directly change bandwidth demand but
|
||||||
|
influences how much buffering, speculation, and replay logic is
|
||||||
|
required to expose bandwidth to software.
|
||||||
|
|
||||||
|
### 2.3 Bandwidth vs. Latency
|
||||||
|
|
||||||
|
Bandwidth and latency are distinct metrics. A system can have ample
|
||||||
|
aggregate DRAM bandwidth but very high per-request latency, in which
|
||||||
|
case small working sets run well but large working sets suffer.
|
||||||
|
A 128-core system is especially sensitive to the latency component
|
||||||
|
because aggregate concurrency exposes the latency of every shared
|
||||||
|
resource, including DRAM banks, the on-chip interconnect, and the
|
||||||
|
directory or snoop filter.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. Bandwidth Demand Estimation
|
||||||
|
|
||||||
|
### 3.1 Per-Core Demand
|
||||||
|
|
||||||
|
A modern out-of-order core can issue multiple outstanding cache misses.
|
||||||
|
The exact number depends on the pipeline width, miss status holding
|
||||||
|
register (MSHR) count, and prefetcher effectiveness, but published
|
||||||
|
industrial designs commonly support 10-32 outstanding L1 misses per
|
||||||
|
core.
|
||||||
|
|
||||||
|
If each outstanding request ultimately consumes 64 bytes from DRAM,
|
||||||
|
and a core sustains 10 outstanding misses, the per-core demand is:
|
||||||
|
|
||||||
|
10 misses * 64 B / core / (memory round-trip latency)
|
||||||
|
|
||||||
|
For a 200 ns round-trip DRAM latency, this is roughly:
|
||||||
|
|
||||||
|
10 * 64 B / 200 ns = 3.2 GB/s per core
|
||||||
|
|
||||||
|
If all 128 cores simultaneously sustain this rate, the aggregate is
|
||||||
|
approximately 410 GB/s, which already exceeds what a single DDR5
|
||||||
|
channel can deliver.
|
||||||
|
|
||||||
|
This calculation is intentionally rough. It is intended to motivate
|
||||||
|
the rest of the document, not to set a precise requirement.
|
||||||
|
|
||||||
|
### 3.2 Realistic Demand
|
||||||
|
|
||||||
|
In practice, applications rarely sustain peak miss rates on every core
|
||||||
|
simultaneously. Published studies of throughput-oriented server
|
||||||
|
workloads report average L2 MPKI (misses per thousand instructions)
|
||||||
|
in the range of 5-30 for memory-bound kernels and 1-5 for cache-friendly
|
||||||
|
workloads. These values vary significantly across workloads.
|
||||||
|
|
||||||
|
Multiplying through, the worst-case aggregate bandwidth demand of a
|
||||||
|
128-core system can plausibly reach hundreds of GB/s. The XH-1 design
|
||||||
|
must therefore plan for a memory subsystem substantially more capable
|
||||||
|
than a single DDR5 channel.
|
||||||
|
|
||||||
|
### 3.3 Coherence Traffic Overhead
|
||||||
|
|
||||||
|
Cache coherence protocols add traffic on top of application reads and
|
||||||
|
writes. The magnitude depends on the protocol (see Section 6) and the
|
||||||
|
sharing pattern of the workload. A directory-based protocol typically
|
||||||
|
generates at least one directory lookup per coherence transaction and
|
||||||
|
potentially one or more forward/intervention messages.
|
||||||
|
|
||||||
|
A conservative estimate is that coherence and protocol overhead add
|
||||||
|
10-30% to the raw application demand, but this can rise significantly
|
||||||
|
for workloads with intensive cross-core sharing.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. On-Chip Bandwidth
|
||||||
|
|
||||||
|
DRAM bandwidth is only one part of the system. The on-chip
|
||||||
|
interconnect and cache hierarchy must also carry:
|
||||||
|
|
||||||
|
- Coherence messages between cores and directories.
|
||||||
|
- Cache line fills from memory controllers to L2/L3 slices.
|
||||||
|
- Write-backs from L1 to L2.
|
||||||
|
- Snoop responses and invalidation acknowledgements.
|
||||||
|
- Interrupt and configuration traffic (typically negligible).
|
||||||
|
|
||||||
|
For a 128-core mesh-class interconnect, published industrial designs
|
||||||
|
indicate bisection bandwidth targets in the multi-TB/s range. Whether
|
||||||
|
XH-1 reaches this range depends on link width, clock frequency, and
|
||||||
|
topology choices that are not yet decided.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. Interconnect Topology and Bandwidth
|
||||||
|
|
||||||
|
The interconnect determines how effectively aggregate on-chip
|
||||||
|
bandwidth is delivered to the DRAM controllers. Three common
|
||||||
|
topologies for many-core designs are:
|
||||||
|
|
||||||
|
1. Crossbar
|
||||||
|
2. Ring
|
||||||
|
3. Mesh
|
||||||
|
4. Concentrated mesh (cmesh)
|
||||||
|
|
||||||
|
| Topology | Typical use | Bandwidth scaling | Notes |
|
||||||
|
|----------|-------------|-------------------|-------|
|
||||||
|
| Crossbar | small core counts | linear in core count | poor scalability beyond ~16 cores |
|
||||||
|
| Ring | mid counts | limited by ring hops | simple but latency grows |
|
||||||
|
| Mesh | mid-to-high counts | roughly linear with links | common in tiled designs |
|
||||||
|
| Cmesh | high counts | similar to mesh, fewer routers | reduces router count |
|
||||||
|
|
||||||
|
A crossbar at 128 cores is generally impractical due to area and
|
||||||
|
wire count. A ring at 128 cores suffers from long hop counts and
|
||||||
|
limited bisection bandwidth. A mesh or concentrated mesh is the
|
||||||
|
most commonly chosen topology in published academic and industrial
|
||||||
|
128-core class designs.
|
||||||
|
|
||||||
|
The XH-1 topology is not yet decided; both a 2D mesh and a 2D
|
||||||
|
concentrated mesh are candidate proposals.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 6. Coherence Protocol Bandwidth
|
||||||
|
|
||||||
|
### 6.1 Protocol Options
|
||||||
|
|
||||||
|
Two main families of coherence protocol exist:
|
||||||
|
|
||||||
|
- Snooping protocols (broadcast-based).
|
||||||
|
- Directory protocols (tracking-based).
|
||||||
|
|
||||||
|
At 128 cores, broadcast-based snooping is generally considered
|
||||||
|
impractical because every coherence transaction must reach every
|
||||||
|
core, consuming O(N) bandwidth per request. A directory protocol
|
||||||
|
incurs O(1) directory lookups per request in the common case but
|
||||||
|
requires directory storage and handles sharing patterns differently.
|
||||||
|
|
||||||
|
A directory is therefore the expected XH-1 proposal, consistent with
|
||||||
|
published 64-128 core research prototypes.
|
||||||
|
|
||||||
|
### 6.2 Directory Organization
|
||||||
|
|
||||||
|
Directories can be implemented in several ways:
|
||||||
|
|
||||||
|
- Full map directory: tracks sharers exactly. Storage cost is O(N)
|
||||||
|
per line; prohibitive at 128 cores without compression or
|
||||||
|
hierarchical schemes.
|
||||||
|
- Limited pointer directory: tracks a fixed number of sharers per
|
||||||
|
line (for example 4-6 sharers). When the count is exceeded, a
|
||||||
|
fallback representation (often a "broadcast" or "coarse vector"
|
||||||
|
marker) is used.
|
||||||
|
- Coarse vector directory: groups cores into regions (e.g. 8-core
|
||||||
|
tiles) and tracks presence at region granularity. Lower storage
|
||||||
|
cost but less precise; can produce additional coherence traffic.
|
||||||
|
- Sparse directory: similar to coarse vector but with different
|
||||||
|
trade-offs.
|
||||||
|
|
||||||
|
A limited pointer directory is the most commonly used scheme in
|
||||||
|
published industrial designs in this core-count range. The XH-1
|
||||||
|
design is expected to use a limited pointer directory, possibly
|
||||||
|
augmented with a coarse region fallback, but this is a proposal
|
||||||
|
rather than a decision.
|
||||||
|
|
||||||
|
### 6.3 Coherence Bandwidth Cost
|
||||||
|
|
||||||
|
For a limited pointer directory with P pointers, the common case
|
||||||
|
is one directory lookup per coherence transaction. Bandwidth cost
|
||||||
|
scales with the number of cores only when the directory overflows
|
||||||
|
into broadcast mode.
|
||||||
|
|
||||||
|
Published research indicates that overflow frequency is workload
|
||||||
|
dependent. For read-mostly shared data, pointer counts typically
|
||||||
|
stay small. For write-shared data, a limited pointer directory
|
||||||
|
can overflow frequently, increasing coherence traffic.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 7. Cache Hierarchy Bandwidth
|
||||||
|
|
||||||
|
The cache hierarchy must sustain:
|
||||||
|
|
||||||
|
- L1 to L2 traffic (fills, write-backs, invalidations).
|
||||||
|
- L2 to L3 traffic.
|
||||||
|
- L3 to DRAM traffic.
|
||||||
|
- Snoop and probe traffic at every level.
|
||||||
|
|
||||||
|
Each cache level adds banked storage and multiple read/write ports to
|
||||||
|
deliver bandwidth. The area and energy cost of these ports grows
|
||||||
|
nonlinearly with port count.
|
||||||
|
|
||||||
|
For XH-1, a three-level hierarchy is the expected proposal:
|
||||||
|
|
||||||
|
- Private L1 per core.
|
||||||
|
- Private L2 per core (or per core pair) with a coherence controller.
|
||||||
|
- Shared L3 distributed across tiles.
|
||||||
|
|
||||||
|
The exact partition of L2 vs. L3 capacity, associativity, and
|
||||||
|
bandwidth is a design question left open in this document.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 8. Bandwidth Allocation and QoS
|
||||||
|
|
||||||
|
When 128 cores compete for limited DRAM bandwidth, some form of
|
||||||
|
allocation policy is required. Options include:
|
||||||
|
|
||||||
|
- Strict priority ordering.
|
||||||
|
- Weighted fair queuing.
|
||||||
|
- Per-core bandwidth caps.
|
||||||
|
- Per-class (e.g. memory-mapped I/O vs. CPU) priorities.
|
||||||
|
|
||||||
|
Without allocation policy, a single core running a streaming workload
|
||||||
|
can saturate the memory subsystem and starve other cores. This is a
|
||||||
|
documented failure mode in many-core research literature.
|
||||||
|
|
||||||
|
RISC-V itself does not define memory bandwidth allocation semantics;
|
||||||
|
this is an XH-1 implementation choice. It is expected that the XH-1
|
||||||
|
memory controllers expose some form of bandwidth partitioning, but
|
||||||
|
this is a proposal rather than a decision.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 9. Memory Controller Architecture
|
||||||
|
|
||||||
|
The memory controller is the point at which on-chip bandwidth meets
|
||||||
|
off-chip DRAM. Key design choices include:
|
||||||
|
|
||||||
|
- Number of independent channels.
|
||||||
|
- Channel-to-core mapping.
|
||||||
|
- Scheduling policy (FR-FCFS, parallelism-aware batch scheduling,
|
||||||
|
etc.).
|
||||||
|
- Address mapping policy.
|
||||||
|
- Refresh and power management handling.
|
||||||
|
|
||||||
|
For DDR5-class memory, a 4-8 channel configuration is a reasonable
|
||||||
|
starting point for a 128-core design, but this is workload dependent
|
||||||
|
and is presented as a proposal.
|
||||||
|
|
||||||
|
For HBM3-class memory, the channel count is fixed by the stack
|
||||||
|
configuration; bandwidth is delivered through many independent
|
||||||
|
pseudo-channels per stack.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 10. Quantitative Summary
|
||||||
|
|
||||||
|
The following table summarises representative bandwidth figures from
|
||||||
|
published sources. These are not XH-1 targets; they are external
|
||||||
|
reference points.
|
||||||
|
|
||||||
|
| Technology / Configuration | Peak Bandwidth | Source basis |
|
||||||
|
|-----------------------------|----------------|--------------|
|
||||||
|
| DDR5 single channel, 6400 MT/s | ~51.2 GB/s | JEDEC DDR5 spec |
|
||||||
|
| DDR5 dual channel, 6400 MT/s | ~102.4 GB/s | JEDEC DDR5 spec |
|
||||||
|
| HBM3 6-high stack | ~819 GB/s | JEDEC HBM3 spec |
|
||||||
|
| HBM3E 8-high stack | ~1 TB/s class | vendor published figures |
|
||||||
|
| 128-core aggregate theoretical (10 outstanding misses/core) | ~410 GB/s | rough calculation, Section 3.1 |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 11. Advantages and Disadvantages of Candidate Memory Systems
|
||||||
|
|
||||||
|
| Option | Advantages | Disadvantages |
|
||||||
|
|--------|------------|---------------|
|
||||||
|
| DDR5 multi-channel | Standard, widely available, lower cost per GB | Lower peak bandwidth per channel, more channels increase board complexity |
|
||||||
|
| HBM3 / HBM3E | Very high bandwidth, high channel count per stack | Higher cost, 2.5D/3D packaging required, lower capacity per stack |
|
||||||
|
| Hybrid (HBM + DDR) | Combine high bandwidth with high capacity | Complex memory hierarchy, two coherence domains |
|
||||||
|
|
||||||
|
The XH-1 memory technology choice is not yet decided.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 12. Scalability Problems at 128 Cores
|
||||||
|
|
||||||
|
The following scalability problems are well documented in the
|
||||||
|
many-core literature and are directly relevant to XH-1:
|
||||||
|
|
||||||
|
- **DRAM bandwidth saturation**: aggregate demand can exceed
|
||||||
|
realistic DRAM bandwidth.
|
||||||
|
- **Directory overflow**: limited pointer directories overflow
|
||||||
|
more often as core count rises, increasing coherence traffic.
|
||||||
|
- **Coherence traffic growth**: as more cores share data,
|
||||||
|
coherence messages grow superlinearly in some sharing patterns.
|
||||||
|
- **Interconnect bisection bottlenecks**: insufficient bisection
|
||||||
|
bandwidth creates hotspots at memory controllers.
|
||||||
|
- **Quality-of-service degradation**: without allocation policy,
|
||||||
|
fairness collapses as core count rises.
|
||||||
|
- **Verification cost**: coherence state space grows roughly
|
||||||
|
combinatorially with core count.
|
||||||
|
|
||||||
|
These are not unique to XH-1 but apply to any 128-core coherent
|
||||||
|
design.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 13. Interactions With Other Subsystems
|
||||||
|
|
||||||
|
### 13.1 Pipeline
|
||||||
|
|
||||||
|
A deeper pipeline can support higher clock frequencies but increases
|
||||||
|
the penalty of cache misses. Bandwidth-bound workloads become
|
||||||
|
sensitive to the round-trip latency the pipeline exposes. The
|
||||||
|
pipeline must include enough MSHRs to expose memory-level parallelism.
|
||||||
|
|
||||||
|
### 13.2 Cache Hierarchy
|
||||||
|
|
||||||
|
Cache capacity and associativity directly affect miss rate and hence
|
||||||
|
DRAM bandwidth demand. Larger caches reduce demand at the cost of
|
||||||
|
area and access time. Bandwidth and capacity are coupled; both
|
||||||
|
must be planned together.
|
||||||
|
|
||||||
|
### 13.3 Memory System
|
||||||
|
|
||||||
|
The memory controller scheduling policy, address mapping, and
|
||||||
|
refresh handling all affect effective bandwidth. Closed-page vs.
|
||||||
|
open-page policies, bank parallelism, and request batching
|
||||||
|
materially change delivered bandwidth.
|
||||||
|
|
||||||
|
### 13.4 Interconnect
|
||||||
|
|
||||||
|
The interconnect must deliver at least as much aggregate bandwidth
|
||||||
|
as the DRAM subsystem, plus coherence overhead. Otherwise the
|
||||||
|
interconnect becomes the bottleneck and DRAM is underutilised.
|
||||||
|
|
||||||
|
### 13.5 Coherence
|
||||||
|
|
||||||
|
Coherence traffic is a first-class component of bandwidth demand.
|
||||||
|
Protocol choice and directory organisation directly determine the
|
||||||
|
share of bandwidth consumed by coherence.
|
||||||
|
|
||||||
|
### 13.6 Interrupts
|
||||||
|
|
||||||
|
Interrupts are typically not bandwidth-limited, but interrupt
|
||||||
|
delivery traffic and any interrupt-cause queues add minor traffic
|
||||||
|
to the system fabric.
|
||||||
|
|
||||||
|
### 13.7 Operating System
|
||||||
|
|
||||||
|
The OS schedules processes across cores, controls page placement,
|
||||||
|
and may implement memory allocation policies that affect locality.
|
||||||
|
The OS does not directly set DRAM bandwidth but its page colour
|
||||||
|
and NUMA-style placement policies interact with the memory
|
||||||
|
controllers.
|
||||||
|
|
||||||
|
### 13.8 Verification
|
||||||
|
|
||||||
|
Coherence protocols at 128 cores have a large state space.
|
||||||
|
Verification complexity grows roughly with the product of
|
||||||
|
per-core state and the number of cores. Formal methods,
|
||||||
|
randomised testing, and traffic pattern fuzzing are all relevant.
|
||||||
|
|
||||||
|
### 13.9 Performance
|
||||||
|
|
||||||
|
Bandwidth is one of the dominant performance constraints for
|
||||||
|
memory-bound workloads. Performance counters for DRAM channel
|
||||||
|
utilisation, interconnect link utilisation, and coherence message
|
||||||
|
counts are essential for tuning.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 14. RISC-V Architectural Requirements
|
||||||
|
|
||||||
|
The following RISC-V architectural elements are relevant to bandwidth
|
||||||
|
analysis:
|
||||||
|
|
||||||
|
- The RISC-V ISA specifies `fence`, `fence.i`, and `fence.vma`
|
||||||
|
instructions that order memory operations.
|
||||||
|
- The privileged architecture defines physical memory attributes
|
||||||
|
(PMA) and, in designs with translation, page-based attributes
|
||||||
|
(PBMT, NAPOT, etc.) that can affect memory behaviour.
|
||||||
|
- The RISC-V coherency management extensions (for example, the
|
||||||
|
proposed H-extension in the vector and hypervisor contexts, and
|
||||||
|
the cache-management operations in the RISC-V Profiles and
|
||||||
|
IOMMU-related specifications) provide instructions such as
|
||||||
|
`cbo.clean`, `cbo.flush`, and `cbo.inval` that interact with
|
||||||
|
cache bandwidth.
|
||||||
|
|
||||||
|
RISC-V itself does not mandate a specific coherence protocol or
|
||||||
|
memory bandwidth management policy. These are XH-1 implementation
|
||||||
|
choices.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 15. Open Questions and Unresolved Design Decisions
|
||||||
|
|
||||||
|
The following items are not yet decided and require further research
|
||||||
|
or design effort:
|
||||||
|
|
||||||
|
- Choice of DRAM technology (DDR5 multi-channel vs. HBM3-class).
|
||||||
|
- Number and placement of memory controllers.
|
||||||
|
- Interconnect topology (mesh vs. concentrated mesh, dimensional
|
||||||
|
order, link width).
|
||||||
|
- Coherence protocol variant (e.g. MESI vs. MOESI and the
|
||||||
|
directory structure).
|
||||||
|
- Bandwidth allocation policy and per-core fairness guarantees.
|
||||||
|
- L2 vs. L3 capacity split and per-tile bandwidth targets.
|
||||||
|
- Quality-of-service hooks exposed to the operating system.
|
||||||
|
- Verification strategy for the coherence and interconnect at
|
||||||
|
128 cores.
|
||||||
|
|
||||||
|
These are recorded as open questions; no recommendation is made
|
||||||
|
where the available evidence does not strongly favour a particular
|
||||||
|
choice.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 16. Recommendations
|
||||||
|
|
||||||
|
Recommendations are made only where published evidence and engineering
|
||||||
|
practice converge.
|
||||||
|
|
||||||
|
- **Use a directory-based coherence protocol**: published 64-128 core
|
||||||
|
research and industrial designs consistently use directories; a
|
||||||
|
snoop-based approach is not credible at 128 cores.
|
||||||
|
- **Plan for multi-channel DRAM**: a single DDR5 channel is unlikely
|
||||||
|
to meet aggregate demand; at least 2-4 channels are a reasonable
|
||||||
|
starting point if DDR5 is used, with HBM as an alternative if
|
||||||
|
higher bandwidth is required.
|
||||||
|
- **Provide per-class quality-of-service in the memory controllers**:
|
||||||
|
without allocation, fairness is not achievable in a 128-core
|
||||||
|
system.
|
||||||
|
- **Instrument the design with bandwidth and coherence performance
|
||||||
|
counters**: empirical tuning is unavoidable.
|
||||||
|
- **Verify coherence state space with a combination of formal
|
||||||
|
methods and randomised testing**: full state space enumeration
|
||||||
|
is not feasible at 128 cores.
|
||||||
|
|
||||||
|
These are recommendations, not decisions; they are subject to
|
||||||
|
revision as the XH-1 project progresses.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 17. References
|
||||||
|
|
||||||
|
- RISC-V International, *The RISC-V Instruction Set Manual, Volume I:
|
||||||
|
Unprivileged ISA*, current ratified version. RVWMO is described
|
||||||
|
in Volume I.
|
||||||
|
- RISC-V International, *The RISC-V Instruction Set Manual, Volume II:
|
||||||
|
Privileged Architecture*, current ratified version.
|
||||||
|
- JEDEC, *DDR5 SDRAM Standard* (JESD79-5), published release.
|
||||||
|
- JEDEC, *High Bandwidth Memory DRAM (HBM3)*, JESD238A.
|
||||||
|
- JEDEC, *High Bandwidth Memory DRAM (HBM3E)*, JESD238B-1 (where
|
||||||
|
applicable).
|
||||||
|
- Hennessy and Patterson, *Computer Architecture: A Quantitative
|
||||||
|
Approach*, recent edition, for general memory system modelling
|
||||||
|
and AMAT analysis.
|
||||||
|
- Sorin, Hill, and Wood, *A Primer on Memory Consistency and Cache
|
||||||
|
Coherence*, for coherence protocol analysis.
|
||||||
|
- Martin, Hill, and Wood, and successor work, on many-core
|
||||||
|
coherence verification and token coherence (where relevant to
|
||||||
|
XH-1 verification work).
|
||||||
|
|
||||||
|
Where specific figures in this document depend on vendor or
|
||||||
|
JEDEC-published numbers, the source is indicated in-line. Where
|
||||||
|
figures are derived from calculations, this is stated explicitly.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 18. Document Status
|
||||||
|
|
||||||
|
This document is a research proposal and analysis. It is not a
|
||||||
|
description of an implemented design. Items marked as "proposal"
|
||||||
|
or "expected" have not been formally decided within the XH-1
|
||||||
|
project. Items marked as "open" represent unresolved design
|
||||||
|
questions. The document should be re-reviewed as the XH-1
|
||||||
|
architecture specification stabilises.
|
||||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user