mirror of
https://github.com/allexanderbergmns/xh1-research.git
synced 2026-08-26 20:27:02 +00:00
383 lines
31 KiB
JSON
383 lines
31 KiB
JSON
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
{"id":"gen-1787748502-7TBb19axZ6o0mq32lh2o","object":"chat.completion","created":1787748502,"model":"minimax/minimax-m3:free","provider":"GMICloud","system_fingerprint":null,"service_tier":null,"choices":[{"index":0,"logprobs":null,"finish_reason":"stop","native_finish_reason":"stop","message":{"role":"assistant","content":"# Cache Coherency\n\n## Scope\n\nThis document investigates cache-coherency protocols and implementation\nstrategies for the XH-1, a custom 128-core RISC-V processor. It covers\ncoherency at the private-L1 / shared-L2 boundary, directory organization,\ninterconnect interactions, ordering, interrupt and OS visibility,\nverification concerns, and the scalability problems introduced by a\ncore count of this size.\n\nThe XH-1 implementation choices described here are **proposals**\nunless explicitly noted. Where evidence is unavailable, that is\nstated rather than guessed.\n\n---\n\n## 1. Problem Statement\n\nIn a 128-core design, the system may contain on the order of:\n\n- 128 private L1 instruction caches\n- 128 private L1 data caches (or split I/D, with or without prefetch)\n- 64–128 private L2 slices (depending on whether L2 is private or\n shared-NUCA)\n- 1–16 shared L3 slices\n- 1–4 DRAM controllers\n\nAny private cache that holds a writable copy of a line is a potential\nsource of stale data being read by another core. Coherence is the\nmechanism that prevents the architectural state of memory from\ndiverging across cores.\n\nThe problem is not new — directory-based coherence for large core\ncounts has been studied since the Origin 2000 (Choi et al., ISCA 1995)\nand the FLASH multiprocessor (Kuskin et al., 1994) — but the\nspecifics of a 128-core RISC-V machine, the available ISA primitives,\nand the interaction with the RISC-V weak-memory model (RVWMO) and the\noptional \"Zam\" extension warrant explicit treatment.\n\n---\n\n## 2. RISC-V Architectural Requirements\n\n### 2.1 The RISC-V Memory Model (RVWMO)\n\nRISC-V uses the \"RISC-V Weak Memory Ordering\" (RVWMO) model defined\nin the RISC-V ISA Manual, Volume I (Unprivileged Architecture note:\nmemory-model details live in the Unprivileged ISA Specification,\nChapter \"Memory Model\" and the associated RISC-V Memory Model\nformalism, as of the 2024/2025 ratified specification).\n\nKey implications for the coherency implementation:\n\n- Loads and stores to a single address are program-order (PO) and\n preserved in coherence order, but loads/stores to **different**\n addresses may be reordered.\n- A memory-consistency model is **not** a coherence protocol. RVWMO\n presumes that coherence (single-address ordering) is already\n provided; the coherency protocol must supply that single-address\n total order, and the interconnect + cache controller must enforce\n the fences (`fence`, `fence.i`, `fence.vma`, atomic\n instructions, and any `aq`/`rl` ordering bits added by the\n `Zaamo`/`Zalrsc` and the forthcoming `Zam` extension).\n\n### 2.2 Atomic Instructions\n\nRISC-V defines the `A` extension (\"Atomic\") for LR/SC, and\noptionally `Zaamo` for atomic memory operations (AMOSWAP, AMOADD,\nAMOXOR, AMOAND, AMOOR, AMOMIN, AMOMAX, AMOMINU, AMOMAXU) and\n`Zalrsc` for LR/SC.\n\nTwo implementation requirements follow:\n\n1. The coherency protocol must guarantee that an AMO or LR/SC\n performed by core *i* appears to take effect atomically with\n respect to other cores. This means the protocol must invalidate\n or writeback all other sharers, fetch exclusive ownership, and\n complete the operation before the result is returned — and no\n other core can observe an intermediate state.\n2. The coherence point (the directory or the point of total\n ordering) must serialize the AMO, which has direct performance\n implications discussed in §6.\n\n### 2.3 Cache-Block Management (CBO) Extensions\n\nRISC-V defines the following cache-block management extensions that\nthe coherency design must consider:\n\n- `Zicbom` — Cache-Block Management Operations\n (`cbo.clean`, `cbo.flush`, `cbo.inval`).\n- `Zicbof` — Cache-Block Flush and Invalidate (`cbo.flush`).\n- `Zicbop` — Cache-Block Prefetch (`prefetch.i`, `prefetch.r`,\n `prefetch.w`).\n- `Zicsr` — required for any implementation that needs\n machine-mode cache maintenance.\n\nThese are **invalidation-class hints**: the architecture does not\nrequire that `cbo.inval` complete globally, but on an XH-1 with\ncoherent private caches, the implementation must guarantee that\na subsequent access from any agent sees the effects of the\nmaintenance. For the XH-1, this requires the L1 invalidate to\ninvalidate any other L1 that may have a copy — i.e., a\ncoherency-aware CBO implementation.\n\n### 2.4 I/O Coherence (PMAs and `Maintained` Caches)\n\nRISC-V defines Physical Memory Attributes including \"Maintained\ncoherence\" (`IO` is by default non-coherent; the system can declare\nregions coherent or not). The hypervisor extension\n(`H` extension) adds `MTTCG` (Machine Timer/Counter\nTranslation), which interacts with cache maintenance for\nvirtualized timer reads. See also the forthcoming\n`Smmtt` (split MMIO) and `Smstateen` extensions.\n\nFor XH-1, this means:\n\n- Main memory regions (DDR) are coherent.\n- MMIO regions (CSRs, accelerators, NICs) may or may not be\n coherent depending on how the SoC is integrated. The XH-1\n proposal is to mark MMIO as **non-coherent** and require\n software (or an in-kernel shim) to issue `fence`/`fence.i` and\n `cbo.flush` around MMIO, as RISC-V does not mandate I/O\n coherence.\n\n### 2.5 The \"Zam\" / Atomic Memory Ordering Extension (Proposed)\n\nA proposed `Zam` extension adds `aq` and `rl` bits to loads and\nstores, providing acquire and release semantics on plain loads and\nstores (similar to AArch64 LDAR/STLR). This is not yet ratified\nin the upstream specification at the time of writing, so any\nXH-1 support for it must be treated as **optional / proposed**.\n\n---\n\n## 3. Coherency Protocol Options\n\nThe two dominant protocol families are **snooping** (broadcast) and\n**directory-based**. Hybrid forms (e.g., snoop-filter assisted,\nregion-directed) are also realistic.\n\n### 3.1 Snooping / Broadcast Protocols\n\nIn a snooping protocol, every coherence controller observes every\ntransaction on a shared broadcast medium (typically a shared bus or\na snooping ring/mesh).\n\n- **MSI**, **MESI**, **MOESI** are the canonical state machines.\n- Advantages: conceptually simple; no directory to size; latency on\n cache-to-cache transfers can be low.\n- Disadvantages: bandwidth scales with O(N) cores × O(1) broadcast\n traffic; snoop filtering helps but does not eliminate the\n fundamental bandwidth problem. For 128 cores, a flat broadcast\n fabric is infeasible.\n\n### 3.2 Directory-Based Protocols\n\nA directory-based protocol maintains, for each cache line in\nmain memory, a presence record (a \"directory entry\") of which\ncaches hold a copy.\n\nDirectory organizations:\n\n- **Full-map (full bit-vector)**: one bit per core per line.\n Memory cost is *lines × cores* bits.\n- **Coarse vector**: a single bit per L2 slice, plus a precise\n L1 directory inside each slice.\n- **Sparse directory (limited pointer)**: a directory with *k*\n pointers per entry (e.g., *k* = 4, 8). On overflow, fallback is\n needed (broadcast or \"evict-on-overflow\" / coarse fallback).\n- **Tagless / Tag-directed**: directory state held in the L2 tags\n themselves (e.g., the Jenga coherence scheme, D. J. Sorin et al.,\n 2011, ASPLOS).\n- **In-cache directories**: directory is distributed in the L3\n tags.\n\n### 3.3 Hybrid Schemes\n\nSeveral published designs mix snoop and directory:\n\n- **AMD HyperTransport / Infinity Fabric / Coherent Fabric**:\n uses a directory-based protocol at the global level with\n broadcast capability for limited scopes.\n- **Intel MESIF** (home snoop / source-snoop blends in some\n generations).\n- **ARM CCN / CMN**: snoop-filter-assisted broadcast on a\n ring/star, with directory overlays for very large systems.\n- **CXL 3.0 / CXL.cache**: directory-based, designed for fabric-\n scale coherence; each device (Type 3 accelerator, host) is a\n coherence participant.\n\n---\n\n## 4. Quantitative Analysis for 128 Cores\n\n### 4.1 Directory Sizing (Full-Map)\n\nAssume:\n\n- L2 line size: 64 B\n- L2 capacity per core: 256 KiB\n- Lines per L2: 4 096\n- Cores: 128\n\nFull-map directory size for one L2 slice:\n\n```\n4 096 lines × 128 bits/line = 524 288 bits = 64 KiB\n```\n\nFor a 64 MiB aggregate L2 (128 × 512 KiB if doubled), the\ndirectory cost is 8 MiB just for L1↔L2 tracking. This is\nnon-trivial but not catastrophic.\n\nA more typical approach is:\n\n- L1 directory in L2 (precise, full-map L1 → L2).\n- L2 directory in L3 (coarse, L2-slice → L3-slice).\n\nWith 16 L3 slices, the L2-in-L3 directory becomes a\n*lines × 16* matrix, which is far cheaper.\n\n### 4.2 Snoop Bandwidth (Broadcast)\n\nA 128-core snoop fabric running at 1 GHz with 64 B coherence\nmessages (e.g., a \"RdBk\" or \"Inv\" carrying the address and a\nsmall set of control bits) at, say, a 32 B flit × 2 flits per\nmessage = 64 B per coherence event:\n\n```\n128 cores × 1 event/(100 instructions) × 64 B/event\n ≈ 82 GB/s of snoop bandwidth, at 100 MPIS/core\n```\n\nThis is the lower bound — a snoop-filter would reduce the volume\nto maybe 5–10 GB/s, but every snoop-fabric hop still costs\nenergy. For 128 cores, a **directory-based** design is the\ndefault serious option. (The exact numbers depend heavily on\nworkload; the order-of-magnitude point is what matters.)\n\n### 4.3 Hop Count\n\nIf the coherent interconnect is a 2D mesh with 16×8 cores (each\ncore = L1 + L2 slice, with an L3 slice per tile), the worst-case\nhop count is 22 edges (Manhattan) for the 16×8 mesh, or ~14 hops\nfor 8×16 if the floorplan is constrained differently.\n\nThis directly affects the latency of directory lookups, invalidates,\nand cache-to-cache forwards.\n\n### 4.4 Cache-to-Cache Forward Latency\n\nA reasonable target:\n\n- 2-hop coherent request: 30–40 cycles in a 2D mesh of 128 cores.\n- 4-hop (mesh-diameter) coherent request: 80–120 cycles.\n- Directory miss (request to DRAM controller): 150–250 ns DRAM\n access + interconnect.\n\nThese figures are consistent with published data for commercial\nlarge-core-count designs (e.g., the AMD EPYC \"Rome\" generation\nreported ~70 ns L3 miss latency; \"Genoa\" 96-core is in a similar\nrange). *No XH-1-specific measurement is available yet.*\n\n---\n\n## 5. Proposed XH-1 Coherency Design\n\nThe following is a **proposal**, not a frozen specification. It is\npresented as the most reasonable starting point based on the\nanalysis above.\n\n### 5.1 Cache Hierarchy\n\nA proposed 3-level hierarchy for XH-1:\n\n| Level | Type | Per-core | Aggregate (128 cores) |\n|------:|------------------------------|---------:|----------------------:|\n| L1-I | Private, VIPT, 32 KiB | 32 KiB | 4 MiB |\n| L1-D | Private, VIPT, 32 KiB, write-through to L2 | 32 KiB | 4 MiB |\n| L2 | Private, slice, 256 KiB, write-back, inclusive of L1 | 256 KiB | 32 MiB |\n| L3 | Shared, 16 slices × 4 MiB, NUCA, inclusive, directory host | — | 64 MiB |\n| DRAM | DDR5-5600, 8 channels | — | ~358 GB/s peak |\n\nNotes:\n\n- L1-D is **write-through to L2** in this proposal. This eliminates\n the need to handle dirty L1 lines during an invalidate; the L2\n is always authoritative. The trade-off is L2 traffic, which is\n discussed in §6.\n- L2 inclusive of L1 allows the L3 directory to be a precise\n L2-tracker, simplifying invalidates.\n- L3 inclusive of L2 allows a single directory to enforce\n coherence; L2 capacity may be over-provisioned due to inclusion\n (a well-known cost discussed in §7).\n\n### 5.2 Protocol: MESI with Sparse Directory\n\n- **Protocol**: MESI per slice, with optional `F` (\"Forward\")\n state on a single L2 slice to avoid ping-pong on read-only\n sharing.\n- **Directory**: sparse, 8 pointers per L3 entry, with\n broadcast fallback on overflow.\n- **Home node**: the L3 slice holding the directory entry is\n the \"home\" for that address; requests are routed by\n address-hash.\n- **Request flow**:\n 1. Core issues load/store.\n 2. L1 miss → L2.\n 3. L2 miss → request routed to home (L3 slice).\n 4. Home looks up directory:\n - **Hit, exclusive owner**: forward request to owner L2,\n owner responds with data (or performs AMO locally and\n acks).\n - **Hit, shared**: forward to one sharer (round-robin or\n F-state holder); invalidate others on write.\n - **Miss**: request to DRAM controller.\n\n### 5.3 Coherent Interconnect\n\n- **Topology**: 2D mesh, 16×8 tiles (each tile = 1 core +\n private L1/L2 + 1 L3 slice + 1 mesh router).\n- **Routing**: deterministic DOR (dimension-order routing) for\n coherence; adaptive for data.\n- **Coherence link width**: 64-bit (8-byte flit × 8 flits), with\n a control flit for protocol messages.\n- **Link frequency**: 2 GHz target (2× core clock at 1 GHz\n per the pipeline research note in this repository).\n\n### 5.4 RISC-V-Specific Items\n\n- The coherency protocol must expose `aq`/`rl` semantics. RVWMO\n assumes that the protocol preserves the per-address total order\n for AMOs and for `aq`/`rl` (if/when `Zam` is ratified).\n- `cbo.inval` on an L1 line: the L1 must check the L2\n directory entry and invalidate all other L1s (or, if L1 is\n write-through, downgrade to a flush of the L2 to the L3 and\n remote invalidates).\n- For MMIO: a `cmo`/memory-mapped bridge must enforce\n non-coherent semantics. The XH-1 proposal is to mark MMIO\n regions as non-cacheable in the PMA, and require\n `fence`/`fence.i` for ordering.\n\n---\n\n## 6. Alternative Designs\n\n### 6.1 Write-Back L1 (instead of Write-Through)\n\n- **Pro**: halves L2 traffic for write-intensive workloads.\n- **Con**: requires the directory to be able to forward\n invalidates to L1s and to handle dirty L1 lines on a\n cache-to-cache transfer (writeback of dirty L1 to the\n requesting L2 before data can be forwarded).\n- **Verdict**: write-back L1 is the more common choice in\n commercial designs (e.g., Intel Skylake-SP, AMD Zen). For\n XH-1, write-through is proposed for simplicity, but write-\n back is a viable alternative that should be re-evaluated\n during implementation. *Decision: deferred.*\n\n### 6.2 Snoop-Filter-Assisted Hybrid\n\nA snoop filter is a small associative structure that tracks\nwhich L2 slices are *not* likely to hold a copy of a line. The\nhome node sends a directed snoop only to L2s in the filter's\n\"may have\" set, and broadcasts only on filter misses.\n\n- **Pro**: lower directory hardware, faster common-case\n forwarding.\n- **Con**: a snoop filter is itself a cache with all the\n capacity-management problems of a cache.\n- **Verdict**: not proposed for XH-1 at 128 cores; a 128-bit\n full-map is well within hardware budget.\n\n### 6.3 Non-Inclusive / Non-Private L2\n\nA non-private (shared) L2 with NUCA placement changes the\nprotocol substantially: the L1-L2 boundary becomes part of the\ncoherence protocol rather than the L2-L3 boundary.\n\n- **Pro**: better aggregate capacity utilization.\n- **Con**: much more complex protocol; \"in cache directories\"\n or per-tile tagless coherence (Jenga) becomes attractive.\n- **Verdict**: not proposed for XH-1 v1; may be revisited.\n\n### 6.4 CXL.cache as the Coherence Fabric\n\nIf the XH-1 is intended to attach coherent accelerators\n(CXL Type 1/2 devices), the coherence fabric can be CXL.cache\nor a derivative.\n\n- **Pro**: standardized, supported by accelerators and by\n the CXL 3.0/3.1 specifications.\n- **Con**: CXL.cache adds a fabric round-trip and a different\n ordering model (the CXL.cache \"Gem5-style\" ordering rules).\n Latency-sensitive core-to-core coherence on CXL is\n difficult.\n- **Verdict**: the XH-1 internal fabric should be a private,\n low-latency mesh; CXL.cache should be used for **external**\n attachment only.\n\n### 6.5 Tile-Based (Manycore) Coherence\n\nFor designs that look more like a manycore (e.g., the\nRISC-V-based SiFive U74-mesh variants, the Esperanto ET-SoC-1\nreported at 1 000+ RISC-V cores), coherence is sometimes done\nwith a \"tile\" abstraction where the L2 is non-private and the\ndirectory is per-tile.\n\n- **Pro**: scales to thousands of cores.\n- **Con**: requires a different memory consistency model and\n a more sophisticated directory (often \"tagless\" or\n \"in-cache\" directory).\n- **Verdict**: not appropriate for a 128-core XH-1, which is\n within the \"directory-is-cheap\" sweet spot.\n\n---\n\n## 7. Scalability Problems at 128 Cores\n\n### 7.1 Inclusion\n\nL3 inclusive of L2 means L3 must hold *every* L2 line. The\nclassical \"three Cs\" problem (capacity, conflict, compulsory)\nis worsened:\n\n- A working set of 64 MiB that lives in L2 (just over the\n proposed L3) causes L3 thrashing. Inclusion can be\n relaxed (\"mostly inclusive\") but this re-introduces\n directory-coherence complexity.\n\nThe XH-1 proposal is **inclusive** for v1, with a planned\nrevisit toward mostly-inclusive if benchmarks demand it.\n\n### 7.2 Directory Coverage and Overflow\n\nA sparse directory with 8 pointers at 128 cores is fine. As\nthe core count grows past 256, pointer-occupancy under\nbroadcast workloads (e.g., a single hot lock) can hit the\noverflow path. The XH-1 does not have this problem at\n128 cores, but the overflow path (broadcast + downgrade\non next miss) must still be implemented and tested.\n\n### 7.3 False Sharing\n\nTwo cores writing to different bytes in the same 64 B line\ngenerate coherence traffic as if they were sharing. This is\na software problem but the protocol must not amplify it.\nThe XH-1 proposal uses 64 B lines, consistent with industry\npractice. (Line-size studies for 128 cores suggest 64 B is\nslightly more friendly to directory bandwidth than 128 B,\nand 32 B increases the directory size by 2×.)\n\n### 7.4 Coherence Latency Variance\n\nA 128-core mesh has a diameter that depends on topology.\nA naïve layout can place the home node 14+ hops from the\nrequesting core. Solutions:\n\n- **Address-hash rebalancing**: choose the home node as a\n function of the address and the mesh, to spread load and\n shorten average distance.\n- **Migration / replication**: the home can be moved\n (\"third-level migration\", Basu et al., HPCA 2012) or\n the line can be replicated at a closer slice.\n\n### 7.5 Memory Ordering and Store-Buffer Pressure\n\nIn RVWMO, each core may buffer stores and forward them to\nL2 out of order. The protocol must:\n\n- Preserve the per-address order of stores from one core.\n- Honor `fence` and `aq`/`rl` semantics.\n- Provide a \"completion\" or \"global observation\" fence\n for I/O.\n\n`FENCE` and `FENCE.I` are implemented by draining the\nstore buffer to the point of coherence (the L2) and\nquiescing the L1-D.\n\n### 7.6 Interrupt Coherence\n\nWhen an MSI is delivered to a core, the interrupt handler\nmust see coherent memory. RISC-V's `mip`/`sip`/`sip` CSRs\nare machine-local, but the interrupt handler runs on the\ntarget hart and observes its own L1, which is by\nconstruction coherent. The remaining concern is that\ninter-processor interrupts (via CLINT/PLIC or the\nadvanced interrupt architecture AIA) need **fences** to\nbe ordered with respect to the device write that triggered\nthe interrupt — the same as any MMIO write.\n\n---\n\n## 8. Interactions\n\n### 8.1 Pipeline\n\nThe L1 miss path stalls the pipeline on a long-latency\ncoherence transaction. A 128-core system can have coherence\nlatencies of 100+ cycles for a 4-hop miss, so the pipeline\nmust:\n\n- Have a non-blocking L1 (MSHRs, hit-under-miss).\n- Support a sufficient number of in-flight loads/stores\n (ROB size, issue width).\n- Provide a clean way to recover from a coherence NACK\n (the line was invalidated mid-transaction). This typically\n requires a \"replay\" path.\n\n### 8.2 Cache Hierarchy\n\nThe hierarchy is constrained by the protocol (see §5.1).\nThe L1-L2 protocol must:\n\n- Use a message-based interface (CEs, Req, Rsp, Data,\n Ack) that matches the interconnect protocol.\n- Support `cbo.clean`/`cbo.flush`/`cbo.inval` correctly.\n- Maintain L1 inclusivity with L2 (or correctly handle\n non-inclusive cases).\n\n### 8.3 Memory System\n\nThe DRAM controller is the \"owner of last resort\" for\ndirty data evicted from L3. The coherence protocol must\nissue a `WB` (writeback) transaction to DRAM when an\nL3 line is evicted in `M` state. DDR5 supports high\nwrite bandwidth but the latency of a write-to-DRAM and\nthe need to \"ack\" the eviction are protocol-visible.\n\n### 8.4 Interconnect\n\nThe coherent interconnect carries three classes of\ntraffic:\n\n1. **Coherence messages** (Req, Rsp, Data, Ack) — small,\n latency-sensitive.\n2. **Bulk data** (cache-line transfers) — 64 B per\n message, latency-sensitive.\n3. **Non-coherent traffic** (MMIO, DMA, CXL).\n\nThe interconnect must provide:\n\n- **Virtual channels** (VCs) for each class to avoid\n deadlock (e.g., 4 VCs: req, rsp, data, snoop).\n- **Ordered** vs. **unordered** networks.\n- **Adaptive routing** for data; DOR is sufficient for\n coherence.\n\n### 8.5 Coherence ↔ Other Subsystems\n\n- **Power management**: a core in WFI may have a dirty\n line in its L1. The protocol must be able to flush\n the L1 (e.g., via `cbo.flush`) on a power-down\n sequence, and the wake-up sequence must invalidate\n stale L1 entries.\n- **RAS**: parity/ECC on directory and on coherence\n messages is required. A single bad coherence message\n can cause a system hang or a silent data corruption;\n strong ECC (SECDED or stronger) is recommended.\n\n### 8.6 Interrupts\n\nSee §7.6. The interaction with the AIA is significant:\nif the XH-1 implements the AIA, the IMSIC (Incoming\nMessage-Signaled Interrupt Controller) is a memory-\nmapped device; the message file accesses must be ordered\nwith respect to the device write. The XH-1 proposal is\nto use `cbo.flush`/`fence` at the AIA driver.\n\n### 8.7 Operating System\n\nThe Linux kernel port (or a custom kernel) must:\n\n- Handle `cbo.*` correctly. The RISC-V Linux kernel\n uses `cbo.clean`/`cbo.flush` for non-coherent DMA\n in some configurations.\n- Implement the `dma_map_*` API; on a coherent platform\n the implementation can be a no-op (or a `cbo.flush`\n fence).\n- Handle cross-CPU `tlb` invalidations; in RISC-V these\n are via SFENCE.VMA with the ASID/hart context.\n- Provide a `sbi_remote_fence_i` and\n `sbi_remote_hfence_vvma` etc. (or an IPI-based\n equivalent) for cross-hart maintenance.\n\n### 8.8 Verification\n\nThis is one of the hardest parts of a 128-core coherent\ndesign. A non-exhaustive list of verification challenges:\n\n- **Coherence invariants**: Single-Writer / Multiple-\n Reader (SWMR), Data-Value (DV), and Get-First /\n Get-Next (ordering) must hold under all interleavings.\n- **Test generation**: Murphi, TLA+, or a custom\n model checker for the directory FSM.\n- **Randomized testing at the RTL level**: a generator\n that produces multi-core litmus tests and runs them\n on the RTL.\n- **Coverage**: every directory-entry state (pointers 0\n to N) must be hit; every \"directory overflow\" path\n must be exercised.\n- **Performance verification**: cycle-approximate\n simulation of representative workloads (PARSEC,\n Splash, Graph500) to validate that the protocol\n does not collapse under realistic traffic.\n\nA reference for the verification methodology: the\nInfinity Fabric / AMD verification methodology (publicly\ndescribed in part at the 2017 HPCA tutorial and in\nAMD's 2018 ISSCC paper) and Intel's published\ndescriptions of MESIF verification.\n\n### 8.9 Performance\n\nCoherency has direct performance impacts:\n\n- **AMO throughput**: every AMO is serialized at the\n home node. The XH-1 must support at least one\n in-flight AMO per core.\n- **Sharing patterns**: a workload with 128 threads all\n updating a single counter will collapse to one\n transaction at a time, regardless of core count. This\n is a workload problem, not a protocol problem.\n- **Ping-pong**: two cores alternating writes to the\n same line can saturate the interconnect. The\n proposed MESI-F (\"forward\") state mitigates this for\n read-only but not for the write ping-pong case;\n the only mitigation is line coloring in software\n or `Ziccam` (Cache Coherence and Atomic\n Maintenance, hypothetical) extensions.\n\n---\n\n## 9. Open Questions and Unresolved Design Issues\n\n1. **Inclusion vs. mostly-inclusion**: What is the\n measured thrashing cost at 64 MiB L3 with 128 cores\n on PARSEC?\n2. **Line size**: 64 B is proposed. What is the trade-\n off with 128 B for directory size and false-sharing\n exposure?\n3. **Coherence link width**: 64-bit is proposed. Would\n 128-bit reduce the hop count at the cost of more\n wire area?\n4. **Sparse directory overflow policy**: broadcast\n fallback vs. downgrade-on-evict. Which has fewer\n pathological cases?\n5. **CXL.cache attach point**: at the L3 slice or at\n a dedicated L4? Latency impact?\n6. **`Zam` (`aq`/`rl`) support**: if ratified, does\n the XH-1 protocol require any new states or\n transitions?\n7. **Verification oracle**: how to validate coherence\n on a 128-core pre-silicon simulation. A 128-core\n RTL simulation is intractable at the cycle level\n for the volume of tests needed. The XH-1\n verification plan must adopt a higher-abstraction\n model (gem5 or a custom fast-functional model) and\n a smaller-scale formal model.\n\n---\n\n## 10. Recommendations\n\nBased on the analysis above, the following are\n**proposed** starting points for the XH-1, to be\nre-evaluated against benchmarks and verification\nresults:\n\n1. **Adopt MESI with optional `F`-state** in the L2\n slices, with a sparse 8-pointer directory in L3.\n2. **Use a 2D 16×8 mesh** with 64-bit coherent links\n at 2 GHz.\n3. **L1-D write-through to L2** for v1, with a planned\n migration to write-back L1 once the directory FSM\n is verified.\n4. **Inclusive L3 of L2** for v1, with \"mostly\n inclusive\" as a future option.\n5. **64 B cache lines** consistent with industry\n practice.\n6. **64 MiB aggregate L3** in 16 slices × 4 MiB.\n7. **Non-coherent MMIO** with mandatory `fence`/\n `cbo.flush` in the kernel.\n8. **Implement the CBO extensions** (`Zicbom`/\n `Zicbof`/`Zicbop`) correctly, even at the cost of\n additional L1 invalidation paths.\n9. **Plan for verification** using a multi-level\n methodology: model-checked FSMs at the directory\n level, randomized RTL litmus tests at the unit\n level, gem5 or fast-functional at the system\n level, and PARSEC/Splash for performance.\n\nThese are **proposals**, not commitments.\n\n---\n\n## 11. Citations and References\n\nThe following primary references were used in this\ndocument. Specific page numbers are not cited because\nthe document is a synthesis, not a line-by-line\ncommentary.\n\n- RISC-V International, *The RISC-V Instruction Set\n Manual, Volume I: Unprivileged Architecture*, current\n ratified version (2024/2025), Chapter \"Memory Model\"\n for RVWMO.\n- RISC-V International, *The RISC-V Instruction Set\n Manual, Volume II: Privileged Architecture*, current\n ratified version, for `A`, `CBO`, and `H` extensions.\n- Choi, B.-J. *et al.*, \"Stanford DASH Multiprocessor\"\n and Choi, L. *et al.*, \"The SGI Origin 2000\", both\n foundational for directory-coherence design.\n Origin 2000 description: Laudon, J. and Lenoski, D.,\n \"The SGI Origin: A ccNUMA Highly Scalable Server\",\n ISCA 1997.\n- Kuskin, J. *et al.*, \"The FLASH Multiprocessor\",\n ISCA 1994.\n- Martin, M. M. K. *et al.*, \"Directory Coherence for\n the MIT Alewife\", 1999 (describes limit-pointer\n directories).\n- Sorin, D. J., *et al.*, \"Jenga: A Coherent Memory\n Hierarchy for Future Manycore Processors\", ASPLOS\n 2011 (for tagless coherence, an alternative).\n- Intel, \"An Introduction to the Intel QuickPath\n Interconnect\" (Intel whitepaper, 2009) — for MESIF\n and snoop-filter design.\n- AMD, \"AMD EPYC 7002 Series Architecture\" and \"AMD\n EPYC 9004 Series Architecture\", Hot Chips\n presentations, 2018 and 2022.\n- ARM, \"CoreLink CCN-508 Cache Coherent Network\"\n (TRM, 2014) — for snoop-filter and ring topology\n design.\n- CXL Consortium, *Compute Express Link (CXL)\n Specification*, 3.0 and 3.1, for CXL.cache.\n- Basu, A. *et al.*, \"Third-Level Memory\n Replication\", HPCA 2012 (for migration /\n replication at the L3).\n\nIf a citation above is wrong or unavailable, the XH-1\ndocumentation process should mark it as such rather\nthan silently propagate the error.\n\n---\n\n*End of document.*","refusal":null,"reasoning":null}}],"usage":{"prompt_tokens":664,"completion_tokens":7447,"total_tokens":8111,"cost":0,"is_byok":false,"prompt_tokens_details":{"cached_tokens":135,"cache_write_tokens":0,"audio_tokens":0,"video_tokens":0},"cost_details":{"upstream_inference_cost":0,"upstream_inference_prompt_cost":0,"upstream_inference_completions_cost":0},"completion_tokens_details":{"reasoning_tokens":0,"image_tokens":0,"audio_tokens":0}}} |