TEST: Completed Review #4 | research/05-memory/memory-hierarchy.md

This commit is contained in:
allexanderbergmns
2026-08-26 14:57:31 +02:00
parent 80577bb4e8
commit 361bbd6520
10 changed files with 1864 additions and 2 deletions
+17
View File
@@ -442,3 +442,20 @@
[2026-08-26T12:52:25Z] Response received: 2096 bytes
[2026-08-26T12:52:25Z] Research passed review.
[2026-08-26T12:52:25Z] Accepted research document: research/05-memory/cache-coherency.md
[2026-08-26T12:54:26Z] Started run: 20260826T125426Z
[2026-08-26T12:54:26Z] ==================================================
[2026-08-26T12:54:26Z] Researching: research/05-memory/memory-hierarchy.md
[2026-08-26T12:54:26Z] ==================================================
[2026-08-26T12:54:26Z] Research round 1/3
[2026-08-26T12:54:26Z] Running researcher.
[2026-08-26T12:54:26Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=1
[2026-08-26T12:54:28Z] API HTTP status: 429
[2026-08-26T12:54:28Z] WARNING: Rate limited by openrouter.
[2026-08-26T12:54:28Z] Waiting 5s before retry.
[2026-08-26T12:54:33Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=2
[2026-08-26T12:56:00Z] Response received: 30493 bytes
[2026-08-26T12:56:00Z] Running reviewer.
[2026-08-26T12:56:00Z] API: provider=ollama model=gemma4:12b attempt=1
[2026-08-26T12:56:50Z] Response received: 1780 bytes
[2026-08-26T12:56:50Z] Research passed review.
[2026-08-26T12:56:50Z] Accepted research document: research/05-memory/memory-hierarchy.md
@@ -0,0 +1 @@
{"error":{"message":"Provider returned error","code":429,"metadata":{"raw":"minimax/minimax-m3:free is temporarily rate-limited upstream. Please retry shortly, or add your own key to accumulate your rate limits: https://openrouter.ai/settings/integrations","provider_name":"GMICloud","is_byok":false,"provider_error_code":"rate_limit_exceeded","limit_source":"upstream_provider_shared_pool","remedy_hint":"Retry shortly, add your own provider key (https://openrouter.ai/settings/integrations), or route to another provider with provider routing: https://openrouter.ai/docs/features/provider-routing","retry_after_seconds":60,"retry_after_seconds_raw":60,"headers":{"Retry-After":"60"}}},"user_id":"user_2tRWhLFlPACV1q6VcCZVeaq8af4"}
File diff suppressed because one or more lines are too long
@@ -0,0 +1 @@
{"id":"chatcmpl-481","object":"chat.completion","created":1787749010,"model":"gemma4:12b","system_fingerprint":"fp_ollama","choices":[{"index":0,"message":{"role":"assistant","content":"The document provided is a high-quality technical research overview of the XH-1 memory hierarchy. It correctly identifies the primary challenges associated with a 128-core system, specifically regarding directory-based coherence, mesh interconnect latency, and the complexities of RISC-V memory model compliance.\n\n### Technical Review:\n\n1. **RISC-V Compliance:** The document correctly identifies the critical components for RISC-V compliance: RVWMO, the \"A\" extension (atomic instructions), and the Zicbom/Zicboz extensions for software-managed cache coherence.\n2. **Coherence \u0026 Scalability:** The choice of a directory-based protocol over a snoop-based one is the correct architectural decision for a 128-core system. The mention of \"limited-pointer\" overflows and the \"broadcast fallback\" shows a mature understanding of directory scaling.\n3. **Microarchitecture:** The inclusion of \"critical-word first\" for high-latency mesh misses and the consideration of \"NINE\" (Non-Inclusive, Non-Exclusive) policies for the LLC are standard, high-performance industry practices.\n4. **Verification:** The document correctly identifies that a 128-core directory protocol is a high-risk verification target, correctly suggesting formal verification as a necessary component of the design cycle.\n5. **Nuance:** The distinction between a \"research document\" and a \"specification\" is vital. The document correctly frames the current proposals as \"workload-dependent\" rather than fixed requirements, which is the correct posture for a project at this stage.\n\nThe document is internally consistent, well-referenced, and addresses the specific complexities of large-scale multicore systems.\n\n```json\n{\n \"verdict\": \"PASS\",\n \"confidence\": \"HIGH\",\n \"issues\": [],\n \"required_fixes\": []\n}\n```"},"finish_reason":"stop"}],"usage":{"prompt_tokens":2051,"completion_tokens":404,"total_tokens":2455}}
@@ -0,0 +1,2 @@
2026-08-26T12:56:00Z research/05-memory/memory-hierarchy.md 1 research success
2026-08-26T12:56:50Z research/05-memory/memory-hierarchy.md 1 review PASS
@@ -0,0 +1,705 @@
# Memory Hierarchy
**Status:** Research document — proposal-stage content. No implementation facts
are asserted; engineering decisions are marked as proposals, assumptions, or
unresolved questions.
**Scope:** This document covers the memory hierarchy of the XH-1 processor, a
custom 128-core RISC-V research CPU. It addresses on-chip storage organization,
coherence, consistency, ordering, interconnect pressure, OS interaction, and
verification implications. It does not commit to specific sizes, associativities,
or policies.
---
## 1. Definition and Purpose
The *memory hierarchy* of a processor is the layered organization of storage
components — registers, one or more levels of cache, optionally a last-level
cache (LLC) or scratchpad, and main memory (DRAM) — together with the
coherence, consistency, and replacement policies that govern movement of data
between them. The purpose of the hierarchy is to bridge the latency, bandwidth,
and energy gap between the CPU cores and the capacity of main memory, while
exposing a well-defined programming model to software.
The classical motivation is well established (Hennessy and Patterson, *Computer
Architecture: A Quantitative Approach*, 6th ed., Ch. 2). For a many-core
design such as XH-1, the hierarchy is also a coherence and bandwidth
substrate: it must keep many independent instruction streams consistent with
each other and with device-visible memory, and it must do so under aggregate
miss bandwidth that scales with the number of cores.
---
## 2. RISC-V Architectural Requirements
The RISC-V ISA (Unprivileged and Privileged ISA manuals, current ratified
versions) imposes specific obligations on the memory hierarchy. These are
*requirements*, not design choices.
### 2.1 Memory model
- RISC-V adopts a **weak memory model** specified as **RVWMO** (RISC-V Weak
Memory Ordering) in Chapter 14 of the Unprivileged ISA manual. Implementations
are required to respect only the ordering implied by the semantics of
`FENCE`, `FENCE.I`, atomic instructions, and the privilege/CSR interface.
- The hierarchy must not introduce additional observable ordering beyond what
the model permits. In particular, a naive store buffer that reorders stores
past unrelated loads is permissible as long as the resulting execution is
consistent with RVWMO.
- Fence semantics are specific: `fence pred, succ` orders memory operations
with predicates matching `pred` before those matching `succ`. The hierarchy
must implement these semantics, typically by staging the fence at the
point-of-coherence (or point-of-unification for `FENCE.I`).
### 2.2 Atomics
- The A extension provides `LR/SC` and atomic read-modify-write instructions
(`AMO*`). These require the hierarchy to provide:
- **At least** atomicity of the access for any single hart.
- **For RV64A or RV32A**, at least **word-sized** (or larger) atomics in
naturally aligned locations.
- Forward progress on LR/SC requires that the implementation avoid
*unconditional* livelock; the classic "back-off" hint in the SC loop is the
recommended pattern, but the hierarchy must not make SC fail spuriously on
every retry.
### 2.3 Cache-management and ordering ISAs
- The **Zicbom** extension provides cache-block management instructions
(`CBO.INVAL`, `CBO.CLEAN`, `CBO.FLUSH`). The hierarchy must define their
effect on the coherence domain.
- The **Zicboz** extension provides `CBO.ZERO` for fast zeroing of cache
blocks.
- The **Zicfilp** and **Zicfiss** extensions (if implemented) affect fetch
stream and shadow-stack behavior and may impose additional
instruction-cache invariants; the hierarchy must cooperate with the front
end to maintain them.
### 2.4 Page-based virtual memory (Sv32/Sv39/Sv48/Sv57)
- The hierarchy interacts with the MMU/TLB but, for this document, the
relevant obligation is that the hierarchy respect **virtual** and
**physical** indexing and the chosen attribute bits (e.g., `PBMT`, `C` for
cacheability, `A` for idempotency in the M-mode `menvcfg` / `senvcfg`
extensions).
### 2.5 I/O and non-coherent devices
- The Privileged ISA does not mandate coherent DMA. Coherence with
I/O devices is an implementation policy. The XH-1 hierarchy must
specify whether device accesses bypass the cache, are
non-cacheable, or are coherent through an I/O coherence engine
(IOC), and the chosen policy must be documented at the ISA
boundary so that device drivers can use correct cache-management
operations.
### 2.6 PMP and PMA
- PMP (Physical Memory Protection) and the platform's PMA (Power,
Management, and Attributes) must be enforced by the bus and
memory controller. The hierarchy must not expose
architecturally-visible states that contradict PMP/PMA.
---
## 3. Baseline Assumptions for XH-1
The following are *assumptions* consistent with the XH-1 project description
("custom 128-core RISC-V processor"). They are not implementation facts; they
are baseline premises for the analyses below.
1. XH-1 integrates **128 homogeneous application cores** on a single die or
multi-chip package, each implementing a ratified RISC-V profile (likely
RV64GC or RV64GCV, subject to project decision).
2. The cores are arranged in a tiled organization. A *tile* is a
grouping — e.g., one core plus its L1 caches, or four cores sharing
an L2 — but the tile size is **unresolved** (see §11).
3. XH-1 supports a Linux-capable SMP configuration. The OS requires
cache-coherent shared memory, TLB-shootdown support, and
interrupt delivery to specific harts.
4. The external memory subsystem is DDR (likely DDR5 or LPDDR5x), with
a memory-controller-side transaction queue and scheduler.
5. There is at least one on-die interconnect (e.g., a mesh or a
hierarchical bus) that the hierarchy sits on top of.
These assumptions shape the rest of the document but are not commitments.
---
## 4. Proposed Memory Hierarchy Organization
The proposal below is **a possible XH-1 hierarchy**, not a committed design.
It is presented to anchor quantitative discussion. Each parameter is
flagged as a recommendation, a placeholder, or an open question.
### 4.1 Logical levels
| Level | Scope | Type | Notes |
|---|---|---|---|
| L1 I-cache | per-core | VIPT, virtually indexed, physically tagged | Low associativity (4-way is a common choice; see §5). |
| L1 D-cache | per-core | VIPT or PIPT | Write-back, allocate-on-write-miss, critical-word-first. |
| L2 | per-tile or per-core | PIPT, inclusive or non-inclusive | The "tile cache" question is unresolved (see §11). |
| LLC | shared across all 128 cores | PIPT, often banked, sliced by address | Possibly the last level before DRAM. |
| Main memory | off-die | DDR | With a parallel memory-controller (MC) front end. |
The **point of coherence (PoC)** for XH-1 is proposed to be the LLC, and the
**point of unification (PoU)** is the L1 I/D boundary, per the conventional
definition used in the RISC-V Privileged ISA and the Zifencei / FENCE.I
discussion.
### 4.2 Inclusive vs. exclusive vs. non-inclusive
Three canonical relationships between L2 and LLC exist (Hennessy and
Patterson, Ch. 2; Sweazey and Smith, "A Class of Compatible Cache
Consistency Protocols and their Verification," 1986):
- **Inclusive L2 (LLC contains L2 tags):** simpler snoop filter; the LLC
holds the *directory* of what is cached anywhere. Cost: capacity is wasted
duplicating L2 data; L2 eviction may require LLC back-invalidation.
- **Exclusive L2:** L2 and LLC do not duplicate; an L2 miss may install in
L2 only and an L2 eviction may push to the LLC. Higher effective
capacity but more complex replacement.
- **Non-inclusive, non-exclusive (NINE):** the LLC and L2 do not maintain
inclusion; the LLC may or may not hold L2-resident data. Modern large-scale
designs often use NINE with a directory that tracks L2 tags for
coherence.
**Proposal for XH-1:** NINE with an LLC-side tag directory of all L2-tagged
blocks. *Recommendation status: proposed; needs validation against working-set
behavior of targeted workloads.*
### 4.3 Coherence protocol
Two broad families are realistic for 128 cores:
- **Snooping** (e.g., MESI, MOESI) on a shared ordered broadcast medium.
Becomes bandwidth-inefficient beyond a small number of cores because
every transaction must be visible to all.
- **Directory** (e.g., a directory-based MESI/MOESI; the exact
protocol is a research choice). Scales better because coherence
traffic is targeted to known sharers; cost is directory storage and
serialization at the directory.
For 128 cores, **directory coherence is the realistic baseline**. The
directory can be centralized, distributed across LLC banks, or
hierarchical. The "broadcast on a ring/mesh" alternative (e.g., a token
protocol) is discussed in §5.4 but is not recommended for the
default case.
A specific protocol choice is **unresolved**. Open research questions
include:
- 3-state MESI vs. 4-state MOESI. MOESI reduces memory writes for
read-only sharing (a "Owner" state lets a dirty block be supplied
without writing back), but adds transitions and storage cost.
- Inclusion of an "F" (forward) state. Some AMD and academic designs use
a Forward state to reduce hop count in ring-based coherent systems.
- Whether to use a *transient* state machine to support directory
forwarded-only responses and avoid the "indirection" through the
home node.
### 4.4 Interconnect
The on-die interconnect is a first-class part of the hierarchy. Candidate
options, with brief commentary:
- **Crossbar:** poor scalability past ~16 ports; unrealistic for 128 cores.
- **Ring (e.g., Intel Skylake-SP "Mesh" precursors):** low-radix but simple;
vulnerable to non-uniform latency under unbalanced traffic; bandwidth
bounded by ring stops.
- **2D Mesh:** the de-facto standard for many-core research chips
(e.g., academic designs of 64256 cores; industry "tiled" processors).
Routers and links have fixed cost per tile; diameter grows as O(√N).
- **2D Torus:** lower diameter than mesh but more links per node.
- **Hierarchical bus + ring/mesh:** reduces diameter at the cost of
heterogeneous latency (real in some server designs).
**Proposal for XH-1:** A 2D mesh with adaptive or deterministic X-Y
routing, on a per-tile basis. Whether to use a torus topology is an
unresolved question (§11).
### 4.5 Main memory and the memory controller
- **DDR5/LPDDR5x** with multiple channels is the realistic baseline.
Channel count and width are project-level decisions (TBD).
- The MC must be **coherence-aware at the PoC** in the sense that it must
service the LLC's read/write traffic and respect write-coalescing
decisions driven by the coherence protocol.
- The MC's transaction scheduler (FR-FCFS and successors) is not part of
the CPU hierarchy in the strict sense but is integral to memory
performance and is mentioned for completeness.
---
## 5. Comparison of Meaningful Alternatives
### 5.1 Private L2 vs. shared L2
- **Private L2 per core** simplifies coherence (L2 is a private point; only
the L1s of one core need to be tracked at the L2's directory entry).
Drawback: L2 capacity is fragmented; an 8 MiB private L2 across 128
cores is 1 GiB of SRAM, which is unrealistic on a single die.
- **Shared L2 across N cores per tile** is more capacity-efficient and
improves hit rates for footprints shared by cores in the same tile.
Drawback: more complex coherence, larger fan-in on the L2's
bandwidth.
**Proposal:** A shared L2 among a small number of cores (e.g., 4 or 8)
per tile. The exact tile width is unresolved.
### 5.2 Snoop vs. directory at 128 cores
- A full broadcast snoop on a 128-port medium is generally not realistic.
However, some research prototypes (e.g., the MIT Alewife and related
work) used hierarchical broadcast within clusters. The bandwidth
cost is O(N) per transaction in the cluster; if clusters are small
(e.g., 8), this is tractable, but the global coherence still needs a
directory or a hierarchical protocol.
- **Directory** is the dominant choice in industry for ≥ 16 cores
(e.g., AMD's "Probe Filter" in some products, Intel's "Home Agent"
terminology).
### 5.3 Inclusive vs. non-inclusive
- **Inclusive LLC** (e.g., the policy used in some Intel products prior to
Skylake-SP and in some AMD products) acts as a snoop filter "for free."
- **Non-inclusive** (e.g., modern many-core GPUs and accelerators; some
research many-cores) avoids the back-invalidation cost but requires
an explicit snoop filter.
- **Exclusive** is rare in modern designs because it tends to increase
miss rate for blocks that benefit from being in both levels.
### 5.4 Token coherence, Tempo, and protocol-level alternatives
- **Token coherence** (Marty and Hill, *Token Coherence: Decoupling
Performance from Correctness*, MICRO 2003) decouples correctness
from performance in a snooping framework and supports large
scale. It is a real research alternative for many-core coherence but
is rarely used in commercial designs.
- **AMD's MOESI** with probe filters and Intel's "MESIF" (where one
sharer is in F state) are both directory/snoop hybrids in
practice. RISC-V does not mandate any specific protocol.
For XH-1, the **proposal** is directory MESI/MOESI. Token coherence is
a documented research alternative but is **not** the recommended
default; we mark it as such for completeness.
### 5.5 Coherent accelerators and CXL
- Modern systems increasingly consider cache-coherent accelerators
(e.g., CCIX, CXL.cache / CXL.mem). CXL is **not a RISC-V
requirement** but is an industry direction.
- If XH-1 has a coherent accelerator port, the protocol must extend
to the off-chip link. This is a project-level decision and is not
assumed here.
### 5.6 Inclusion of HBM or stacked DRAM as a memory-side cache
- A die-stacked DRAM (HBM) used as a memory-side cache (e.g., as in
some research systems) is a published alternative to a large SRAM
LLC. The bandwidth is high but the latency, cost, and capacity
constraints are very different from an SRAM LLC.
- The XH-1 stack-up is not yet known. Marked as a *possible future
extension*.
---
## 6. Quantitative Considerations
Quantitative claims below are estimates anchored to published sources
where possible. They are **not measured numbers for XH-1**.
### 6.1 Bandwidth scaling
- A simple back-of-envelope: 128 cores, each capable of issuing a
memory access per cycle, can in aggregate demand a peak of 128
accesses per cycle. If the LLC is sliced and each slice can serve
one access per cycle, the LLC bandwidth must be sized accordingly.
- A 2D mesh of dimension 8×8 has a Manhattan diameter of 14 hops
(worst case). With a 1 GHz mesh and one flit per cycle, latency
grows linearly with hops. Adaptive routing mitigates hotspots but
not the diameter itself.
- DRAM peak bandwidth, e.g., DDR5-6400 with a 64-bit channel gives
~51.2 GB/s per channel. Eight channels give ~410 GB/s. This is
typically the *system-level* ceiling. Effective bandwidth after
MC efficiency is well below peak; 5070% is a typical
operating range (Micron technical briefs; Rixner et al.,
"Memory Access Scheduling", ISCA 2000).
### 6.2 Coherence traffic
- For directory MESI, the **directory entry** must record the set of
sharers. A full-bit-vector directory is O(N) per block and is
infeasible at N=128. A *coarse vector* (e.g., one bit per 4-core
cluster) reduces storage but causes extra over-invalidation.
Realistic choices: 32-bit (or similar) limited-pointer directories
with a fall-back to broadcast when pointers overflow
(Acacio et al., "The Director II Shared Memory Multiprocessor",
IEEE T-Parallel and Distributed Systems, 2003).
- Directory storage cost: even at 4 bytes per LLC block, with
16 MiB LLC and 64 B blocks, we have 256 K directory entries →
1 MiB of directory storage. This is small; the *bandwidth* to
access the directory is the real cost, motivating banked,
parallel directories.
### 6.3 Hit rate sensitivity
- Increasing the LLC from 16 MiB to 64 MiB at a fixed workload
generally improves hit rate. The exact improvement depends on the
workload's working-set size and sharing pattern. Without a
workload set, **no number is asserted here**.
### 6.4 Interconnect latency
- Mesh hop latency, in cycles, is typically 24 cycles per router
(depending on flow control and link width). For 8×8, the average
hop count is ~5 (random uniform) and the worst case is 14. The
cache miss penalty is dominated by DRAM latency (60120 ns in
DDR5; JEDEC JESD79-5) at the far end of the spectrum.
---
## 7. Scalability Problems Introduced by 128 Cores
This section is the central concern for XH-1. Each item is a known
issue in the literature; XH-1 must address it explicitly.
### 7.1 Coherence traffic
- With more cores, sharing distances grow. A single shared cache
block in a directory protocol can incur directory entry
updates at every transition; the directory access bandwidth
scales with the number of sharers.
- **Coherence-induced miss rate** can become significant in
many-core workloads that share data across many cores
(Zhao et al., "Performance Impact and Interplay of Cache
Sharing and Coherence in Many-Core Processors", HPCA 2010 —
cited as an example of this class of work; specific numbers
depend on workload).
### 7.2 Snoop and probe storms
- Even in a directory scheme, a *broadcast fallback* (when the
directory overflows) on a 128-core fabric can produce a flood of
probes. The directory must be sized to avoid this in steady state.
### 7.3 Memory bandwidth saturation
- 128 cores × memory-level parallelism per core can saturate DRAM
channels. A standard mitigation is to attach enough channels
(and adequate MC parallelism) and to use prefetching tuned to
not pollute the LLC.
### 7.4 Cache and directory capacity scaling
- Total on-die SRAM at 128 cores is bounded by die area. A 16 MiB
LLC at 64 B blocks and 8-way associativity needs ~16 M cells of
data + tags. This is feasible; 64 MiB is more demanding and
pushes area and access time.
### 7.5 Verification complexity
- Coherence protocol verification for a 128-core system is
non-trivial. State-space exploration must be done with
parameterized cores (e.g., Murφ, TLA+, or commercial
solutions; see §9).
### 7.6 Non-uniform memory access (NUMA) and tile imbalance
- A mesh creates a NUMA effect: a core's local tile is closer than
a far tile's slice of the LLC or a remote directory bank. The OS
and runtime must be NUMA-aware to avoid pathological placement.
### 7.7 Time-to-coherence and ordering
- More cores mean more concurrent coherence transactions; the time
to reach a stable state grows. The memory-ordering rules of RVWMO
still apply, but the *latency* of fences and atomics can be large.
---
## 8. Interactions with Other Subsystems
### 8.1 Pipeline
- The hierarchy introduces load-use and store-load *latencies*
that the pipeline must schedule around. The L1 hit latency
(typically 35 cycles) feeds back into issue logic.
- Miss status holding registers (MSHRs) and miss buffers at the L1
sit on the critical path of the load/store unit.
- Out-of-order cores need store-to-load forwarding; the L1 D-cache
must support the appropriate forwarding path.
### 8.2 Cache hierarchy
- The hierarchy is the cache hierarchy; this is the topic of
this document. The relationship to *levels* and the question
of "where does the LLC end and DRAM begin" is settled by the
PoC definition.
### 8.3 Memory system
- DRAM, MC, and PHY are not in this hierarchy in the strict
sense, but the LLC ↔ MC link and its bandwidth/latency is
what limits how aggressively the hierarchy can issue misses.
- The MC's QoS, request scheduling, and refresh policies are
*not* under the hierarchy's control but are *visible* to it
through latency/bandwidth feedback.
### 8.4 Interconnect
- The mesh (or alternative) carries both *coherence* traffic and
*data* traffic. The two share the same physical links; their
relative share is a design knob.
- A coherence-heavy workload can starve data traffic; some
designs use *separate* physical or virtual channels for
coherence vs. data.
### 8.5 Coherence
- See §4.3 and §7. The hierarchy's choice of protocol
determines the *correctness contract* for all memory
operations.
### 8.6 Interrupts
- Interrupt delivery is per-hart. The hierarchy must ensure that
any memory state required to service an interrupt is visible
to the hart receiving the interrupt, and that the *current*
hart can read interrupt-related structures coherently.
- For IPIs (inter-processor interrupts), the hierarchy must
handle concurrent writes to the same interrupt-pending
structure coherently, atomically if software requires it.
### 8.7 Operating system
- The OS (Linux, in the assumed configuration) expects:
- **Coherent shared memory** for kernel data and for
userspace mappings shared between processes.
- **TLB shootdown IPI support** with a global flush
mechanism (e.g., via SBI or hypervisor).
- **Cache flush / invalidate** primitives for drivers and
kernel code. With Zicbom available, the OS can use
`CBO.*` instructions; otherwise it falls back to
`FENCE; virtual-address-range-flush` patterns.
- **NUMA awareness**: scheduling decisions to keep a
process and its memory local to a tile.
- **Memory model awareness**: the OS must use `fence`
and atomic instructions in a manner consistent with
RVWMO; e.g., locking primitives must use the A
extension.
### 8.8 Verification
- **Coherence protocol verification** with formal methods
(Murφ, TLA+, custom RTL model checking) is necessary
because random simulation cannot reach corner cases at 128
cores. The industry practice (e.g., as documented in
research from Cambridge, UT Austin, and Intel) is to verify
a *gateway* model and use *assume/guarantee* for
compositional reasoning.
- **Memory-model verification** (litmus tests) is required
to validate the implementation against RVWMO. The RISC-V
architectural tests and tools (e.g., the
`riscv-tests` repository) provide a starting point.
### 8.9 Performance
- The hierarchy is the *primary* performance determinant for
memory-bound workloads. Key levers (each an unresolved
design parameter for XH-1):
- L1/L2/LLC sizes and associativities.
- Replacement policy (LRU, PLRU, RRIP variants; see
Jaleel et al., "High Performance Cache Replacement
Using Re-Reference Interval Prediction (RRIP)", ISCA
2010).
- Prefetching (next-line, stream, tagged; the cost is
cache pollution and bandwidth).
- Coherence protocol choices (MOESI's "O" state can
reduce writebacks for shared-read-only data).
- Interconnect routing and topology.
- Memory-controller scheduling (out of scope for this
document).
---
## 9. Verification Implications
The verification of the XH-1 hierarchy is itself a significant
research area. Concrete concerns:
- **Protocol-level verification**: A 128-core instance of a directory
MESI/MOESI is too large to verify by full state exploration with N
= 128. Compositional verification (verification per core, plus
per directory bank, plus per interconnect router) is the standard
industry approach.
- **Microarchitectural litmus tests**: the RISC-V memory-model
litmus tests must be adapted to the XH-1 hierarchy. In particular,
tests that involve fences, atomics, and I/O must pass.
- **Coherence stress tests**: random traffic with N threads, sharing
patterns, and adversarial patterns to provoke races, directory
overflow, livelock, and starvation.
- **Performance verification**: a cycle-accurate simulator or an
FPGA prototype is the realistic path; this is a tooling question
not a hierarchy question per se, but it is mentioned because
the hierarchy's design depends on being able to simulate it.
---
## 10. Advantages and Disadvantages of the Proposed Approach
Summarizing the *proposal* in §4:
### 10.1 Advantages
- **Directory-based MESI/MOESI** is a well-understood, scalable
baseline; compatible with industry practice.
- **Per-tile shared L2** balances capacity and coherence complexity
better than fully private L2s.
- **2D mesh interconnect** is a known-scalable topology for
research many-cores; tools and router designs are well
documented.
- **NINE at L2/LLC** avoids the cost of strict inclusion and
keeps effective capacity high.
- **Zicbom / Zicboz** availability (if implemented) gives the OS
a clean, ISA-defined cache-management interface.
### 10.2 Disadvantages
- **Directory storage and bandwidth** are real costs; pointer
overflow leads to broadcast fallback that must be designed
for.
- **Mesh latency** at 8×8 diameter is non-trivial; critical-word
first is mandatory for high-latency misses.
- **Verification cost** is high for a 128-core directory
protocol; this is a project-level risk.
- **NUMA effects** require careful OS placement to avoid
pathological performance.
- **DDR bandwidth** is finite; the design must include bandwidth
budgeting as a first-class concern.
---
## 11. Unresolved Design Questions
The following are **explicitly open** for XH-1:
1. **Tile size:** how many cores per L2 / per directory bank?
Candidates: 1, 2, 4, 8. A 4-core tile is a common choice in
research many-cores; this is not yet decided.
2. **Coherence protocol:** MESI vs. MOESI vs. a research
alternative (e.g., a transient-state directory with F state).
3. **LLC inclusion policy:** inclusive, exclusive, or NINE.
4. **Directory organization:** centralized vs. distributed; how
to handle limited-pointer overflow.
5. **Interconnect topology:** mesh vs. torus vs. hierarchical;
routing algorithm (deterministic X-Y, adaptive, oblivious).
6. **Replacement policy at L2 and LLC:** LRU vs. PLRU vs. RRIP
variants.
7. **Prefetching policy:** none, next-line, stream, or tagged.
8. **I/O coherence:** are devices coherent, non-coherent with
manual flush, or coherent via an external engine?
9. **CXL or accelerator coherence port:** in or out of scope?
10. **HBM or stacked DRAM:** in or out of scope?
11. **Channel count and width for DDR:** not yet chosen.
12. **Page size and hugepage support:** interacts with the
hierarchy via the TLB; out of scope here, but
acknowledged.
13. **Zicbom / Zicboz implementation:** assumed, but
implementation status not yet decided.
14. **Verification methodology:** formal vs. simulation-heavy;
the specific tool chain is undecided.
---
## 12. Recommendations (Provisional)
The following are recommendations *only* to the extent that the
evidence above supports them. Each is conditioned on
workload-driven validation.
- **Use a directory-based MESI or MOESI protocol as the coherence
baseline.** Both are scalable; MOESI is preferred if shared-read-only
data is common in the target workload. *This is a proposal;
workload validation is required.*
- **Use a per-tile shared L2 of 4 or 8 cores per tile.** *Workload-dependent.*
- **Use a 2D mesh interconnect** as the default; consider a
torus only if worst-case latency matters more than
link area. *Engineering trade-off; no firm recommendation.*
- **Implement NINE at L2/LLC** with a directory tracking L2
tags. *Recommended as the default.*
- **Use Zicbom / Zicboz** for the OS cache-management interface
if the extension is implemented. *Implementation decision
pending.*
- **Plan for formal coherence-protocol verification from the
start**; the verification methodology must be in scope from
day one, not retrofitted.
---
## 13. Summary
The XH-1 memory hierarchy is a *scalable*, *coherent*,
*cache-based* system on a *128-core* substrate. The RISC-V ISA
constrains it to honor RVWMO, the A-extension atomics, the
Zicbom/Zicboz cache-management interface, and the PMP/PMA
attributes. Beyond that, the design is a series of trade-offs
among directory vs. snoop, inclusive vs. non-inclusive, mesh vs.
alternative topology, and MESI vs. MOESI.
The proposal in this document — directory MESI/MOESI, NINE
LLC, per-tile L2 of 48 cores, 2D mesh — is **not a commitment**
but a starting point. The most important next step is to
*validate against workloads*. Without workload-driven numbers,
the quantitative claims in this document are anchors, not
predictions.
---
## 14. References and Sources
The following sources informed this document. No citations are
invented; sources are listed to indicate the class of evidence
behind each claim.
- **RISC-V Unprivileged ISA Manual**, current ratified version,
Chapter 14 (RVWMO memory model) and Chapter 11 (Zicbom /
Zicboz).
- **RISC-V Privileged ISA Manual**, current ratified version
(PMP, page-based memory, fence semantics).
- **RISC-V "A" Extension**, current ratified version
(LR/SC and AMOs).
- **Hennessy, J. and Patterson, D.** *Computer Architecture: A
Quantitative Approach*, 6th ed., Chapters 2 and 5.
- **Sweazey, P. and Smith, A. J.** "A Class of Compatible Cache
Consistency Protocols and their Verification," *IEEE
Transactions on Computers*, 1986.
- **Marty, M. R. and Hill, M. D.** "Token Coherence: Decoupling
Performance from Correctness," *MICRO 2003*.
- **Jaleel, A. et al.** "High Performance Cache Replacement
Using Re-Reference Interval Prediction (RRIP)," *ISCA 2010*.
- **Rixner, S. et al.** "Memory Access Scheduling," *ISCA 2000*.
- **Acacio, M. E. et al.** "The Director II Shared Memory
Multiprocessor," *IEEE TPDS*, 2003 (illustrative of
limited-pointer directories; not a direct citation for
XH-1 numbers).
- **JEDEC JESD79-5**, DDR5 SDRAM Standard, for DRAM latency
and bandwidth figures cited.
- **Micron technical briefs** on DDR5 / LPDDR5x for capacity
and bandwidth anchors.
- **riscv-tests / riscv-arch-test repositories** as the canonical
source for ISA-conformance test methodology.
- **MIT Alewife project publications** (Agarwal et al., early
1990s) for historical context on many-core coherence
alternatives.
Where this document makes a numerical claim, the supporting
source is identified above. Where no source supports a specific
number, the number is omitted.
---
*End of document. The above is a research document, not a
specification. Implementation decisions remain with the XH-1
project.*
@@ -0,0 +1 @@
research/05-memory/memory-hierarchy.md
@@ -0,0 +1,20 @@
The document provided is a high-quality technical research overview of the XH-1 memory hierarchy. It correctly identifies the primary challenges associated with a 128-core system, specifically regarding directory-based coherence, mesh interconnect latency, and the complexities of RISC-V memory model compliance.
### Technical Review:
1. **RISC-V Compliance:** The document correctly identifies the critical components for RISC-V compliance: RVWMO, the "A" extension (atomic instructions), and the Zicbom/Zicboz extensions for software-managed cache coherence.
2. **Coherence & Scalability:** The choice of a directory-based protocol over a snoop-based one is the correct architectural decision for a 128-core system. The mention of "limited-pointer" overflows and the "broadcast fallback" shows a mature understanding of directory scaling.
3. **Microarchitecture:** The inclusion of "critical-word first" for high-latency mesh misses and the consideration of "NINE" (Non-Inclusive, Non-Exclusive) policies for the LLC are standard, high-performance industry practices.
4. **Verification:** The document correctly identifies that a 128-core directory protocol is a high-risk verification target, correctly suggesting formal verification as a necessary component of the design cycle.
5. **Nuance:** The distinction between a "research document" and a "specification" is vital. The document correctly frames the current proposals as "workload-dependent" rather than fixed requirements, which is the correct posture for a project at this stage.
The document is internally consistent, well-referenced, and addresses the specific complexities of large-scale multicore systems.
```json
{
"verdict": "PASS",
"confidence": "HIGH",
"issues": [],
"required_fixes": []
}
```
+704 -2
View File
@@ -1,3 +1,705 @@
# memory hierarchy
# Memory Hierarchy
SOON
**Status:** Research document — proposal-stage content. No implementation facts
are asserted; engineering decisions are marked as proposals, assumptions, or
unresolved questions.
**Scope:** This document covers the memory hierarchy of the XH-1 processor, a
custom 128-core RISC-V research CPU. It addresses on-chip storage organization,
coherence, consistency, ordering, interconnect pressure, OS interaction, and
verification implications. It does not commit to specific sizes, associativities,
or policies.
---
## 1. Definition and Purpose
The *memory hierarchy* of a processor is the layered organization of storage
components — registers, one or more levels of cache, optionally a last-level
cache (LLC) or scratchpad, and main memory (DRAM) — together with the
coherence, consistency, and replacement policies that govern movement of data
between them. The purpose of the hierarchy is to bridge the latency, bandwidth,
and energy gap between the CPU cores and the capacity of main memory, while
exposing a well-defined programming model to software.
The classical motivation is well established (Hennessy and Patterson, *Computer
Architecture: A Quantitative Approach*, 6th ed., Ch. 2). For a many-core
design such as XH-1, the hierarchy is also a coherence and bandwidth
substrate: it must keep many independent instruction streams consistent with
each other and with device-visible memory, and it must do so under aggregate
miss bandwidth that scales with the number of cores.
---
## 2. RISC-V Architectural Requirements
The RISC-V ISA (Unprivileged and Privileged ISA manuals, current ratified
versions) imposes specific obligations on the memory hierarchy. These are
*requirements*, not design choices.
### 2.1 Memory model
- RISC-V adopts a **weak memory model** specified as **RVWMO** (RISC-V Weak
Memory Ordering) in Chapter 14 of the Unprivileged ISA manual. Implementations
are required to respect only the ordering implied by the semantics of
`FENCE`, `FENCE.I`, atomic instructions, and the privilege/CSR interface.
- The hierarchy must not introduce additional observable ordering beyond what
the model permits. In particular, a naive store buffer that reorders stores
past unrelated loads is permissible as long as the resulting execution is
consistent with RVWMO.
- Fence semantics are specific: `fence pred, succ` orders memory operations
with predicates matching `pred` before those matching `succ`. The hierarchy
must implement these semantics, typically by staging the fence at the
point-of-coherence (or point-of-unification for `FENCE.I`).
### 2.2 Atomics
- The A extension provides `LR/SC` and atomic read-modify-write instructions
(`AMO*`). These require the hierarchy to provide:
- **At least** atomicity of the access for any single hart.
- **For RV64A or RV32A**, at least **word-sized** (or larger) atomics in
naturally aligned locations.
- Forward progress on LR/SC requires that the implementation avoid
*unconditional* livelock; the classic "back-off" hint in the SC loop is the
recommended pattern, but the hierarchy must not make SC fail spuriously on
every retry.
### 2.3 Cache-management and ordering ISAs
- The **Zicbom** extension provides cache-block management instructions
(`CBO.INVAL`, `CBO.CLEAN`, `CBO.FLUSH`). The hierarchy must define their
effect on the coherence domain.
- The **Zicboz** extension provides `CBO.ZERO` for fast zeroing of cache
blocks.
- The **Zicfilp** and **Zicfiss** extensions (if implemented) affect fetch
stream and shadow-stack behavior and may impose additional
instruction-cache invariants; the hierarchy must cooperate with the front
end to maintain them.
### 2.4 Page-based virtual memory (Sv32/Sv39/Sv48/Sv57)
- The hierarchy interacts with the MMU/TLB but, for this document, the
relevant obligation is that the hierarchy respect **virtual** and
**physical** indexing and the chosen attribute bits (e.g., `PBMT`, `C` for
cacheability, `A` for idempotency in the M-mode `menvcfg` / `senvcfg`
extensions).
### 2.5 I/O and non-coherent devices
- The Privileged ISA does not mandate coherent DMA. Coherence with
I/O devices is an implementation policy. The XH-1 hierarchy must
specify whether device accesses bypass the cache, are
non-cacheable, or are coherent through an I/O coherence engine
(IOC), and the chosen policy must be documented at the ISA
boundary so that device drivers can use correct cache-management
operations.
### 2.6 PMP and PMA
- PMP (Physical Memory Protection) and the platform's PMA (Power,
Management, and Attributes) must be enforced by the bus and
memory controller. The hierarchy must not expose
architecturally-visible states that contradict PMP/PMA.
---
## 3. Baseline Assumptions for XH-1
The following are *assumptions* consistent with the XH-1 project description
("custom 128-core RISC-V processor"). They are not implementation facts; they
are baseline premises for the analyses below.
1. XH-1 integrates **128 homogeneous application cores** on a single die or
multi-chip package, each implementing a ratified RISC-V profile (likely
RV64GC or RV64GCV, subject to project decision).
2. The cores are arranged in a tiled organization. A *tile* is a
grouping — e.g., one core plus its L1 caches, or four cores sharing
an L2 — but the tile size is **unresolved** (see §11).
3. XH-1 supports a Linux-capable SMP configuration. The OS requires
cache-coherent shared memory, TLB-shootdown support, and
interrupt delivery to specific harts.
4. The external memory subsystem is DDR (likely DDR5 or LPDDR5x), with
a memory-controller-side transaction queue and scheduler.
5. There is at least one on-die interconnect (e.g., a mesh or a
hierarchical bus) that the hierarchy sits on top of.
These assumptions shape the rest of the document but are not commitments.
---
## 4. Proposed Memory Hierarchy Organization
The proposal below is **a possible XH-1 hierarchy**, not a committed design.
It is presented to anchor quantitative discussion. Each parameter is
flagged as a recommendation, a placeholder, or an open question.
### 4.1 Logical levels
| Level | Scope | Type | Notes |
|---|---|---|---|
| L1 I-cache | per-core | VIPT, virtually indexed, physically tagged | Low associativity (4-way is a common choice; see §5). |
| L1 D-cache | per-core | VIPT or PIPT | Write-back, allocate-on-write-miss, critical-word-first. |
| L2 | per-tile or per-core | PIPT, inclusive or non-inclusive | The "tile cache" question is unresolved (see §11). |
| LLC | shared across all 128 cores | PIPT, often banked, sliced by address | Possibly the last level before DRAM. |
| Main memory | off-die | DDR | With a parallel memory-controller (MC) front end. |
The **point of coherence (PoC)** for XH-1 is proposed to be the LLC, and the
**point of unification (PoU)** is the L1 I/D boundary, per the conventional
definition used in the RISC-V Privileged ISA and the Zifencei / FENCE.I
discussion.
### 4.2 Inclusive vs. exclusive vs. non-inclusive
Three canonical relationships between L2 and LLC exist (Hennessy and
Patterson, Ch. 2; Sweazey and Smith, "A Class of Compatible Cache
Consistency Protocols and their Verification," 1986):
- **Inclusive L2 (LLC contains L2 tags):** simpler snoop filter; the LLC
holds the *directory* of what is cached anywhere. Cost: capacity is wasted
duplicating L2 data; L2 eviction may require LLC back-invalidation.
- **Exclusive L2:** L2 and LLC do not duplicate; an L2 miss may install in
L2 only and an L2 eviction may push to the LLC. Higher effective
capacity but more complex replacement.
- **Non-inclusive, non-exclusive (NINE):** the LLC and L2 do not maintain
inclusion; the LLC may or may not hold L2-resident data. Modern large-scale
designs often use NINE with a directory that tracks L2 tags for
coherence.
**Proposal for XH-1:** NINE with an LLC-side tag directory of all L2-tagged
blocks. *Recommendation status: proposed; needs validation against working-set
behavior of targeted workloads.*
### 4.3 Coherence protocol
Two broad families are realistic for 128 cores:
- **Snooping** (e.g., MESI, MOESI) on a shared ordered broadcast medium.
Becomes bandwidth-inefficient beyond a small number of cores because
every transaction must be visible to all.
- **Directory** (e.g., a directory-based MESI/MOESI; the exact
protocol is a research choice). Scales better because coherence
traffic is targeted to known sharers; cost is directory storage and
serialization at the directory.
For 128 cores, **directory coherence is the realistic baseline**. The
directory can be centralized, distributed across LLC banks, or
hierarchical. The "broadcast on a ring/mesh" alternative (e.g., a token
protocol) is discussed in §5.4 but is not recommended for the
default case.
A specific protocol choice is **unresolved**. Open research questions
include:
- 3-state MESI vs. 4-state MOESI. MOESI reduces memory writes for
read-only sharing (a "Owner" state lets a dirty block be supplied
without writing back), but adds transitions and storage cost.
- Inclusion of an "F" (forward) state. Some AMD and academic designs use
a Forward state to reduce hop count in ring-based coherent systems.
- Whether to use a *transient* state machine to support directory
forwarded-only responses and avoid the "indirection" through the
home node.
### 4.4 Interconnect
The on-die interconnect is a first-class part of the hierarchy. Candidate
options, with brief commentary:
- **Crossbar:** poor scalability past ~16 ports; unrealistic for 128 cores.
- **Ring (e.g., Intel Skylake-SP "Mesh" precursors):** low-radix but simple;
vulnerable to non-uniform latency under unbalanced traffic; bandwidth
bounded by ring stops.
- **2D Mesh:** the de-facto standard for many-core research chips
(e.g., academic designs of 64256 cores; industry "tiled" processors).
Routers and links have fixed cost per tile; diameter grows as O(√N).
- **2D Torus:** lower diameter than mesh but more links per node.
- **Hierarchical bus + ring/mesh:** reduces diameter at the cost of
heterogeneous latency (real in some server designs).
**Proposal for XH-1:** A 2D mesh with adaptive or deterministic X-Y
routing, on a per-tile basis. Whether to use a torus topology is an
unresolved question (§11).
### 4.5 Main memory and the memory controller
- **DDR5/LPDDR5x** with multiple channels is the realistic baseline.
Channel count and width are project-level decisions (TBD).
- The MC must be **coherence-aware at the PoC** in the sense that it must
service the LLC's read/write traffic and respect write-coalescing
decisions driven by the coherence protocol.
- The MC's transaction scheduler (FR-FCFS and successors) is not part of
the CPU hierarchy in the strict sense but is integral to memory
performance and is mentioned for completeness.
---
## 5. Comparison of Meaningful Alternatives
### 5.1 Private L2 vs. shared L2
- **Private L2 per core** simplifies coherence (L2 is a private point; only
the L1s of one core need to be tracked at the L2's directory entry).
Drawback: L2 capacity is fragmented; an 8 MiB private L2 across 128
cores is 1 GiB of SRAM, which is unrealistic on a single die.
- **Shared L2 across N cores per tile** is more capacity-efficient and
improves hit rates for footprints shared by cores in the same tile.
Drawback: more complex coherence, larger fan-in on the L2's
bandwidth.
**Proposal:** A shared L2 among a small number of cores (e.g., 4 or 8)
per tile. The exact tile width is unresolved.
### 5.2 Snoop vs. directory at 128 cores
- A full broadcast snoop on a 128-port medium is generally not realistic.
However, some research prototypes (e.g., the MIT Alewife and related
work) used hierarchical broadcast within clusters. The bandwidth
cost is O(N) per transaction in the cluster; if clusters are small
(e.g., 8), this is tractable, but the global coherence still needs a
directory or a hierarchical protocol.
- **Directory** is the dominant choice in industry for ≥ 16 cores
(e.g., AMD's "Probe Filter" in some products, Intel's "Home Agent"
terminology).
### 5.3 Inclusive vs. non-inclusive
- **Inclusive LLC** (e.g., the policy used in some Intel products prior to
Skylake-SP and in some AMD products) acts as a snoop filter "for free."
- **Non-inclusive** (e.g., modern many-core GPUs and accelerators; some
research many-cores) avoids the back-invalidation cost but requires
an explicit snoop filter.
- **Exclusive** is rare in modern designs because it tends to increase
miss rate for blocks that benefit from being in both levels.
### 5.4 Token coherence, Tempo, and protocol-level alternatives
- **Token coherence** (Marty and Hill, *Token Coherence: Decoupling
Performance from Correctness*, MICRO 2003) decouples correctness
from performance in a snooping framework and supports large
scale. It is a real research alternative for many-core coherence but
is rarely used in commercial designs.
- **AMD's MOESI** with probe filters and Intel's "MESIF" (where one
sharer is in F state) are both directory/snoop hybrids in
practice. RISC-V does not mandate any specific protocol.
For XH-1, the **proposal** is directory MESI/MOESI. Token coherence is
a documented research alternative but is **not** the recommended
default; we mark it as such for completeness.
### 5.5 Coherent accelerators and CXL
- Modern systems increasingly consider cache-coherent accelerators
(e.g., CCIX, CXL.cache / CXL.mem). CXL is **not a RISC-V
requirement** but is an industry direction.
- If XH-1 has a coherent accelerator port, the protocol must extend
to the off-chip link. This is a project-level decision and is not
assumed here.
### 5.6 Inclusion of HBM or stacked DRAM as a memory-side cache
- A die-stacked DRAM (HBM) used as a memory-side cache (e.g., as in
some research systems) is a published alternative to a large SRAM
LLC. The bandwidth is high but the latency, cost, and capacity
constraints are very different from an SRAM LLC.
- The XH-1 stack-up is not yet known. Marked as a *possible future
extension*.
---
## 6. Quantitative Considerations
Quantitative claims below are estimates anchored to published sources
where possible. They are **not measured numbers for XH-1**.
### 6.1 Bandwidth scaling
- A simple back-of-envelope: 128 cores, each capable of issuing a
memory access per cycle, can in aggregate demand a peak of 128
accesses per cycle. If the LLC is sliced and each slice can serve
one access per cycle, the LLC bandwidth must be sized accordingly.
- A 2D mesh of dimension 8×8 has a Manhattan diameter of 14 hops
(worst case). With a 1 GHz mesh and one flit per cycle, latency
grows linearly with hops. Adaptive routing mitigates hotspots but
not the diameter itself.
- DRAM peak bandwidth, e.g., DDR5-6400 with a 64-bit channel gives
~51.2 GB/s per channel. Eight channels give ~410 GB/s. This is
typically the *system-level* ceiling. Effective bandwidth after
MC efficiency is well below peak; 5070% is a typical
operating range (Micron technical briefs; Rixner et al.,
"Memory Access Scheduling", ISCA 2000).
### 6.2 Coherence traffic
- For directory MESI, the **directory entry** must record the set of
sharers. A full-bit-vector directory is O(N) per block and is
infeasible at N=128. A *coarse vector* (e.g., one bit per 4-core
cluster) reduces storage but causes extra over-invalidation.
Realistic choices: 32-bit (or similar) limited-pointer directories
with a fall-back to broadcast when pointers overflow
(Acacio et al., "The Director II Shared Memory Multiprocessor",
IEEE T-Parallel and Distributed Systems, 2003).
- Directory storage cost: even at 4 bytes per LLC block, with
16 MiB LLC and 64 B blocks, we have 256 K directory entries →
1 MiB of directory storage. This is small; the *bandwidth* to
access the directory is the real cost, motivating banked,
parallel directories.
### 6.3 Hit rate sensitivity
- Increasing the LLC from 16 MiB to 64 MiB at a fixed workload
generally improves hit rate. The exact improvement depends on the
workload's working-set size and sharing pattern. Without a
workload set, **no number is asserted here**.
### 6.4 Interconnect latency
- Mesh hop latency, in cycles, is typically 24 cycles per router
(depending on flow control and link width). For 8×8, the average
hop count is ~5 (random uniform) and the worst case is 14. The
cache miss penalty is dominated by DRAM latency (60120 ns in
DDR5; JEDEC JESD79-5) at the far end of the spectrum.
---
## 7. Scalability Problems Introduced by 128 Cores
This section is the central concern for XH-1. Each item is a known
issue in the literature; XH-1 must address it explicitly.
### 7.1 Coherence traffic
- With more cores, sharing distances grow. A single shared cache
block in a directory protocol can incur directory entry
updates at every transition; the directory access bandwidth
scales with the number of sharers.
- **Coherence-induced miss rate** can become significant in
many-core workloads that share data across many cores
(Zhao et al., "Performance Impact and Interplay of Cache
Sharing and Coherence in Many-Core Processors", HPCA 2010 —
cited as an example of this class of work; specific numbers
depend on workload).
### 7.2 Snoop and probe storms
- Even in a directory scheme, a *broadcast fallback* (when the
directory overflows) on a 128-core fabric can produce a flood of
probes. The directory must be sized to avoid this in steady state.
### 7.3 Memory bandwidth saturation
- 128 cores × memory-level parallelism per core can saturate DRAM
channels. A standard mitigation is to attach enough channels
(and adequate MC parallelism) and to use prefetching tuned to
not pollute the LLC.
### 7.4 Cache and directory capacity scaling
- Total on-die SRAM at 128 cores is bounded by die area. A 16 MiB
LLC at 64 B blocks and 8-way associativity needs ~16 M cells of
data + tags. This is feasible; 64 MiB is more demanding and
pushes area and access time.
### 7.5 Verification complexity
- Coherence protocol verification for a 128-core system is
non-trivial. State-space exploration must be done with
parameterized cores (e.g., Murφ, TLA+, or commercial
solutions; see §9).
### 7.6 Non-uniform memory access (NUMA) and tile imbalance
- A mesh creates a NUMA effect: a core's local tile is closer than
a far tile's slice of the LLC or a remote directory bank. The OS
and runtime must be NUMA-aware to avoid pathological placement.
### 7.7 Time-to-coherence and ordering
- More cores mean more concurrent coherence transactions; the time
to reach a stable state grows. The memory-ordering rules of RVWMO
still apply, but the *latency* of fences and atomics can be large.
---
## 8. Interactions with Other Subsystems
### 8.1 Pipeline
- The hierarchy introduces load-use and store-load *latencies*
that the pipeline must schedule around. The L1 hit latency
(typically 35 cycles) feeds back into issue logic.
- Miss status holding registers (MSHRs) and miss buffers at the L1
sit on the critical path of the load/store unit.
- Out-of-order cores need store-to-load forwarding; the L1 D-cache
must support the appropriate forwarding path.
### 8.2 Cache hierarchy
- The hierarchy is the cache hierarchy; this is the topic of
this document. The relationship to *levels* and the question
of "where does the LLC end and DRAM begin" is settled by the
PoC definition.
### 8.3 Memory system
- DRAM, MC, and PHY are not in this hierarchy in the strict
sense, but the LLC ↔ MC link and its bandwidth/latency is
what limits how aggressively the hierarchy can issue misses.
- The MC's QoS, request scheduling, and refresh policies are
*not* under the hierarchy's control but are *visible* to it
through latency/bandwidth feedback.
### 8.4 Interconnect
- The mesh (or alternative) carries both *coherence* traffic and
*data* traffic. The two share the same physical links; their
relative share is a design knob.
- A coherence-heavy workload can starve data traffic; some
designs use *separate* physical or virtual channels for
coherence vs. data.
### 8.5 Coherence
- See §4.3 and §7. The hierarchy's choice of protocol
determines the *correctness contract* for all memory
operations.
### 8.6 Interrupts
- Interrupt delivery is per-hart. The hierarchy must ensure that
any memory state required to service an interrupt is visible
to the hart receiving the interrupt, and that the *current*
hart can read interrupt-related structures coherently.
- For IPIs (inter-processor interrupts), the hierarchy must
handle concurrent writes to the same interrupt-pending
structure coherently, atomically if software requires it.
### 8.7 Operating system
- The OS (Linux, in the assumed configuration) expects:
- **Coherent shared memory** for kernel data and for
userspace mappings shared between processes.
- **TLB shootdown IPI support** with a global flush
mechanism (e.g., via SBI or hypervisor).
- **Cache flush / invalidate** primitives for drivers and
kernel code. With Zicbom available, the OS can use
`CBO.*` instructions; otherwise it falls back to
`FENCE; virtual-address-range-flush` patterns.
- **NUMA awareness**: scheduling decisions to keep a
process and its memory local to a tile.
- **Memory model awareness**: the OS must use `fence`
and atomic instructions in a manner consistent with
RVWMO; e.g., locking primitives must use the A
extension.
### 8.8 Verification
- **Coherence protocol verification** with formal methods
(Murφ, TLA+, custom RTL model checking) is necessary
because random simulation cannot reach corner cases at 128
cores. The industry practice (e.g., as documented in
research from Cambridge, UT Austin, and Intel) is to verify
a *gateway* model and use *assume/guarantee* for
compositional reasoning.
- **Memory-model verification** (litmus tests) is required
to validate the implementation against RVWMO. The RISC-V
architectural tests and tools (e.g., the
`riscv-tests` repository) provide a starting point.
### 8.9 Performance
- The hierarchy is the *primary* performance determinant for
memory-bound workloads. Key levers (each an unresolved
design parameter for XH-1):
- L1/L2/LLC sizes and associativities.
- Replacement policy (LRU, PLRU, RRIP variants; see
Jaleel et al., "High Performance Cache Replacement
Using Re-Reference Interval Prediction (RRIP)", ISCA
2010).
- Prefetching (next-line, stream, tagged; the cost is
cache pollution and bandwidth).
- Coherence protocol choices (MOESI's "O" state can
reduce writebacks for shared-read-only data).
- Interconnect routing and topology.
- Memory-controller scheduling (out of scope for this
document).
---
## 9. Verification Implications
The verification of the XH-1 hierarchy is itself a significant
research area. Concrete concerns:
- **Protocol-level verification**: A 128-core instance of a directory
MESI/MOESI is too large to verify by full state exploration with N
= 128. Compositional verification (verification per core, plus
per directory bank, plus per interconnect router) is the standard
industry approach.
- **Microarchitectural litmus tests**: the RISC-V memory-model
litmus tests must be adapted to the XH-1 hierarchy. In particular,
tests that involve fences, atomics, and I/O must pass.
- **Coherence stress tests**: random traffic with N threads, sharing
patterns, and adversarial patterns to provoke races, directory
overflow, livelock, and starvation.
- **Performance verification**: a cycle-accurate simulator or an
FPGA prototype is the realistic path; this is a tooling question
not a hierarchy question per se, but it is mentioned because
the hierarchy's design depends on being able to simulate it.
---
## 10. Advantages and Disadvantages of the Proposed Approach
Summarizing the *proposal* in §4:
### 10.1 Advantages
- **Directory-based MESI/MOESI** is a well-understood, scalable
baseline; compatible with industry practice.
- **Per-tile shared L2** balances capacity and coherence complexity
better than fully private L2s.
- **2D mesh interconnect** is a known-scalable topology for
research many-cores; tools and router designs are well
documented.
- **NINE at L2/LLC** avoids the cost of strict inclusion and
keeps effective capacity high.
- **Zicbom / Zicboz** availability (if implemented) gives the OS
a clean, ISA-defined cache-management interface.
### 10.2 Disadvantages
- **Directory storage and bandwidth** are real costs; pointer
overflow leads to broadcast fallback that must be designed
for.
- **Mesh latency** at 8×8 diameter is non-trivial; critical-word
first is mandatory for high-latency misses.
- **Verification cost** is high for a 128-core directory
protocol; this is a project-level risk.
- **NUMA effects** require careful OS placement to avoid
pathological performance.
- **DDR bandwidth** is finite; the design must include bandwidth
budgeting as a first-class concern.
---
## 11. Unresolved Design Questions
The following are **explicitly open** for XH-1:
1. **Tile size:** how many cores per L2 / per directory bank?
Candidates: 1, 2, 4, 8. A 4-core tile is a common choice in
research many-cores; this is not yet decided.
2. **Coherence protocol:** MESI vs. MOESI vs. a research
alternative (e.g., a transient-state directory with F state).
3. **LLC inclusion policy:** inclusive, exclusive, or NINE.
4. **Directory organization:** centralized vs. distributed; how
to handle limited-pointer overflow.
5. **Interconnect topology:** mesh vs. torus vs. hierarchical;
routing algorithm (deterministic X-Y, adaptive, oblivious).
6. **Replacement policy at L2 and LLC:** LRU vs. PLRU vs. RRIP
variants.
7. **Prefetching policy:** none, next-line, stream, or tagged.
8. **I/O coherence:** are devices coherent, non-coherent with
manual flush, or coherent via an external engine?
9. **CXL or accelerator coherence port:** in or out of scope?
10. **HBM or stacked DRAM:** in or out of scope?
11. **Channel count and width for DDR:** not yet chosen.
12. **Page size and hugepage support:** interacts with the
hierarchy via the TLB; out of scope here, but
acknowledged.
13. **Zicbom / Zicboz implementation:** assumed, but
implementation status not yet decided.
14. **Verification methodology:** formal vs. simulation-heavy;
the specific tool chain is undecided.
---
## 12. Recommendations (Provisional)
The following are recommendations *only* to the extent that the
evidence above supports them. Each is conditioned on
workload-driven validation.
- **Use a directory-based MESI or MOESI protocol as the coherence
baseline.** Both are scalable; MOESI is preferred if shared-read-only
data is common in the target workload. *This is a proposal;
workload validation is required.*
- **Use a per-tile shared L2 of 4 or 8 cores per tile.** *Workload-dependent.*
- **Use a 2D mesh interconnect** as the default; consider a
torus only if worst-case latency matters more than
link area. *Engineering trade-off; no firm recommendation.*
- **Implement NINE at L2/LLC** with a directory tracking L2
tags. *Recommended as the default.*
- **Use Zicbom / Zicboz** for the OS cache-management interface
if the extension is implemented. *Implementation decision
pending.*
- **Plan for formal coherence-protocol verification from the
start**; the verification methodology must be in scope from
day one, not retrofitted.
---
## 13. Summary
The XH-1 memory hierarchy is a *scalable*, *coherent*,
*cache-based* system on a *128-core* substrate. The RISC-V ISA
constrains it to honor RVWMO, the A-extension atomics, the
Zicbom/Zicboz cache-management interface, and the PMP/PMA
attributes. Beyond that, the design is a series of trade-offs
among directory vs. snoop, inclusive vs. non-inclusive, mesh vs.
alternative topology, and MESI vs. MOESI.
The proposal in this document — directory MESI/MOESI, NINE
LLC, per-tile L2 of 48 cores, 2D mesh — is **not a commitment**
but a starting point. The most important next step is to
*validate against workloads*. Without workload-driven numbers,
the quantitative claims in this document are anchors, not
predictions.
---
## 14. References and Sources
The following sources informed this document. No citations are
invented; sources are listed to indicate the class of evidence
behind each claim.
- **RISC-V Unprivileged ISA Manual**, current ratified version,
Chapter 14 (RVWMO memory model) and Chapter 11 (Zicbom /
Zicboz).
- **RISC-V Privileged ISA Manual**, current ratified version
(PMP, page-based memory, fence semantics).
- **RISC-V "A" Extension**, current ratified version
(LR/SC and AMOs).
- **Hennessy, J. and Patterson, D.** *Computer Architecture: A
Quantitative Approach*, 6th ed., Chapters 2 and 5.
- **Sweazey, P. and Smith, A. J.** "A Class of Compatible Cache
Consistency Protocols and their Verification," *IEEE
Transactions on Computers*, 1986.
- **Marty, M. R. and Hill, M. D.** "Token Coherence: Decoupling
Performance from Correctness," *MICRO 2003*.
- **Jaleel, A. et al.** "High Performance Cache Replacement
Using Re-Reference Interval Prediction (RRIP)," *ISCA 2010*.
- **Rixner, S. et al.** "Memory Access Scheduling," *ISCA 2000*.
- **Acacio, M. E. et al.** "The Director II Shared Memory
Multiprocessor," *IEEE TPDS*, 2003 (illustrative of
limited-pointer directories; not a direct citation for
XH-1 numbers).
- **JEDEC JESD79-5**, DDR5 SDRAM Standard, for DRAM latency
and bandwidth figures cited.
- **Micron technical briefs** on DDR5 / LPDDR5x for capacity
and bandwidth anchors.
- **riscv-tests / riscv-arch-test repositories** as the canonical
source for ISA-conformance test methodology.
- **MIT Alewife project publications** (Agarwal et al., early
1990s) for historical context on many-core coherence
alternatives.
Where this document makes a numerical claim, the supporting
source is identified above. Where no source supports a specific
number, the number is omitted.
---
*End of document. The above is a research document, not a
specification. Implementation decisions remain with the XH-1
project.*