mirror of
https://github.com/allexanderbergmns/xh1-research.git
synced 2026-08-26 23:07:01 +00:00
413 lines
34 KiB
JSON
413 lines
34 KiB
JSON
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
{"id":"gen-1787748873-3C9d1nhOhf8LfxFSsOpJ","object":"chat.completion","created":1787748873,"model":"minimax/minimax-m3:free","provider":"GMICloud","system_fingerprint":null,"service_tier":null,"choices":[{"index":0,"logprobs":null,"finish_reason":"stop","native_finish_reason":"stop","message":{"role":"assistant","content":"# Memory Hierarchy\n\n**Status:** Research document — proposal-stage content. No implementation facts\nare asserted; engineering decisions are marked as proposals, assumptions, or\nunresolved questions.\n\n**Scope:** This document covers the memory hierarchy of the XH-1 processor, a\ncustom 128-core RISC-V research CPU. It addresses on-chip storage organization,\ncoherence, consistency, ordering, interconnect pressure, OS interaction, and\nverification implications. It does not commit to specific sizes, associativities,\nor policies.\n\n---\n\n## 1. Definition and Purpose\n\nThe *memory hierarchy* of a processor is the layered organization of storage\ncomponents — registers, one or more levels of cache, optionally a last-level\ncache (LLC) or scratchpad, and main memory (DRAM) — together with the\ncoherence, consistency, and replacement policies that govern movement of data\nbetween them. The purpose of the hierarchy is to bridge the latency, bandwidth,\nand energy gap between the CPU cores and the capacity of main memory, while\nexposing a well-defined programming model to software.\n\nThe classical motivation is well established (Hennessy and Patterson, *Computer\nArchitecture: A Quantitative Approach*, 6th ed., Ch. 2). For a many-core\ndesign such as XH-1, the hierarchy is also a coherence and bandwidth\nsubstrate: it must keep many independent instruction streams consistent with\neach other and with device-visible memory, and it must do so under aggregate\nmiss bandwidth that scales with the number of cores.\n\n---\n\n## 2. RISC-V Architectural Requirements\n\nThe RISC-V ISA (Unprivileged and Privileged ISA manuals, current ratified\nversions) imposes specific obligations on the memory hierarchy. These are\n*requirements*, not design choices.\n\n### 2.1 Memory model\n\n- RISC-V adopts a **weak memory model** specified as **RVWMO** (RISC-V Weak\n Memory Ordering) in Chapter 14 of the Unprivileged ISA manual. Implementations\n are required to respect only the ordering implied by the semantics of\n `FENCE`, `FENCE.I`, atomic instructions, and the privilege/CSR interface.\n- The hierarchy must not introduce additional observable ordering beyond what\n the model permits. In particular, a naive store buffer that reorders stores\n past unrelated loads is permissible as long as the resulting execution is\n consistent with RVWMO.\n- Fence semantics are specific: `fence pred, succ` orders memory operations\n with predicates matching `pred` before those matching `succ`. The hierarchy\n must implement these semantics, typically by staging the fence at the\n point-of-coherence (or point-of-unification for `FENCE.I`).\n\n### 2.2 Atomics\n\n- The A extension provides `LR/SC` and atomic read-modify-write instructions\n (`AMO*`). These require the hierarchy to provide:\n - **At least** atomicity of the access for any single hart.\n - **For RV64A or RV32A**, at least **word-sized** (or larger) atomics in\n naturally aligned locations.\n- Forward progress on LR/SC requires that the implementation avoid\n *unconditional* livelock; the classic \"back-off\" hint in the SC loop is the\n recommended pattern, but the hierarchy must not make SC fail spuriously on\n every retry.\n\n### 2.3 Cache-management and ordering ISAs\n\n- The **Zicbom** extension provides cache-block management instructions\n (`CBO.INVAL`, `CBO.CLEAN`, `CBO.FLUSH`). The hierarchy must define their\n effect on the coherence domain.\n- The **Zicboz** extension provides `CBO.ZERO` for fast zeroing of cache\n blocks.\n- The **Zicfilp** and **Zicfiss** extensions (if implemented) affect fetch\n stream and shadow-stack behavior and may impose additional\n instruction-cache invariants; the hierarchy must cooperate with the front\n end to maintain them.\n\n### 2.4 Page-based virtual memory (Sv32/Sv39/Sv48/Sv57)\n\n- The hierarchy interacts with the MMU/TLB but, for this document, the\n relevant obligation is that the hierarchy respect **virtual** and\n **physical** indexing and the chosen attribute bits (e.g., `PBMT`, `C` for\n cacheability, `A` for idempotency in the M-mode `menvcfg` / `senvcfg`\n extensions).\n\n### 2.5 I/O and non-coherent devices\n\n- The Privileged ISA does not mandate coherent DMA. Coherence with\n I/O devices is an implementation policy. The XH-1 hierarchy must\n specify whether device accesses bypass the cache, are\n non-cacheable, or are coherent through an I/O coherence engine\n (IOC), and the chosen policy must be documented at the ISA\n boundary so that device drivers can use correct cache-management\n operations.\n\n### 2.6 PMP and PMA\n\n- PMP (Physical Memory Protection) and the platform's PMA (Power,\n Management, and Attributes) must be enforced by the bus and\n memory controller. The hierarchy must not expose\n architecturally-visible states that contradict PMP/PMA.\n\n---\n\n## 3. Baseline Assumptions for XH-1\n\nThe following are *assumptions* consistent with the XH-1 project description\n(\"custom 128-core RISC-V processor\"). They are not implementation facts; they\nare baseline premises for the analyses below.\n\n1. XH-1 integrates **128 homogeneous application cores** on a single die or\n multi-chip package, each implementing a ratified RISC-V profile (likely\n RV64GC or RV64GCV, subject to project decision).\n2. The cores are arranged in a tiled organization. A *tile* is a\n grouping — e.g., one core plus its L1 caches, or four cores sharing\n an L2 — but the tile size is **unresolved** (see §11).\n3. XH-1 supports a Linux-capable SMP configuration. The OS requires\n cache-coherent shared memory, TLB-shootdown support, and\n interrupt delivery to specific harts.\n4. The external memory subsystem is DDR (likely DDR5 or LPDDR5x), with\n a memory-controller-side transaction queue and scheduler.\n5. There is at least one on-die interconnect (e.g., a mesh or a\n hierarchical bus) that the hierarchy sits on top of.\n\nThese assumptions shape the rest of the document but are not commitments.\n\n---\n\n## 4. Proposed Memory Hierarchy Organization\n\nThe proposal below is **a possible XH-1 hierarchy**, not a committed design.\nIt is presented to anchor quantitative discussion. Each parameter is\nflagged as a recommendation, a placeholder, or an open question.\n\n### 4.1 Logical levels\n\n| Level | Scope | Type | Notes |\n|---|---|---|---|\n| L1 I-cache | per-core | VIPT, virtually indexed, physically tagged | Low associativity (4-way is a common choice; see §5). |\n| L1 D-cache | per-core | VIPT or PIPT | Write-back, allocate-on-write-miss, critical-word-first. |\n| L2 | per-tile or per-core | PIPT, inclusive or non-inclusive | The \"tile cache\" question is unresolved (see §11). |\n| LLC | shared across all 128 cores | PIPT, often banked, sliced by address | Possibly the last level before DRAM. |\n| Main memory | off-die | DDR | With a parallel memory-controller (MC) front end. |\n\nThe **point of coherence (PoC)** for XH-1 is proposed to be the LLC, and the\n**point of unification (PoU)** is the L1 I/D boundary, per the conventional\ndefinition used in the RISC-V Privileged ISA and the Zifencei / FENCE.I\ndiscussion.\n\n### 4.2 Inclusive vs. exclusive vs. non-inclusive\n\nThree canonical relationships between L2 and LLC exist (Hennessy and\nPatterson, Ch. 2; Sweazey and Smith, \"A Class of Compatible Cache\nConsistency Protocols and their Verification,\" 1986):\n\n- **Inclusive L2 (LLC contains L2 tags):** simpler snoop filter; the LLC\n holds the *directory* of what is cached anywhere. Cost: capacity is wasted\n duplicating L2 data; L2 eviction may require LLC back-invalidation.\n- **Exclusive L2:** L2 and LLC do not duplicate; an L2 miss may install in\n L2 only and an L2 eviction may push to the LLC. Higher effective\n capacity but more complex replacement.\n- **Non-inclusive, non-exclusive (NINE):** the LLC and L2 do not maintain\n inclusion; the LLC may or may not hold L2-resident data. Modern large-scale\n designs often use NINE with a directory that tracks L2 tags for\n coherence.\n\n**Proposal for XH-1:** NINE with an LLC-side tag directory of all L2-tagged\nblocks. *Recommendation status: proposed; needs validation against working-set\nbehavior of targeted workloads.*\n\n### 4.3 Coherence protocol\n\nTwo broad families are realistic for 128 cores:\n\n- **Snooping** (e.g., MESI, MOESI) on a shared ordered broadcast medium.\n Becomes bandwidth-inefficient beyond a small number of cores because\n every transaction must be visible to all.\n- **Directory** (e.g., a directory-based MESI/MOESI; the exact\n protocol is a research choice). Scales better because coherence\n traffic is targeted to known sharers; cost is directory storage and\n serialization at the directory.\n\nFor 128 cores, **directory coherence is the realistic baseline**. The\ndirectory can be centralized, distributed across LLC banks, or\nhierarchical. The \"broadcast on a ring/mesh\" alternative (e.g., a token\nprotocol) is discussed in §5.4 but is not recommended for the\ndefault case.\n\nA specific protocol choice is **unresolved**. Open research questions\ninclude:\n\n- 3-state MESI vs. 4-state MOESI. MOESI reduces memory writes for\n read-only sharing (a \"Owner\" state lets a dirty block be supplied\n without writing back), but adds transitions and storage cost.\n- Inclusion of an \"F\" (forward) state. Some AMD and academic designs use\n a Forward state to reduce hop count in ring-based coherent systems.\n- Whether to use a *transient* state machine to support directory\n forwarded-only responses and avoid the \"indirection\" through the\n home node.\n\n### 4.4 Interconnect\n\nThe on-die interconnect is a first-class part of the hierarchy. Candidate\noptions, with brief commentary:\n\n- **Crossbar:** poor scalability past ~16 ports; unrealistic for 128 cores.\n- **Ring (e.g., Intel Skylake-SP \"Mesh\" precursors):** low-radix but simple;\n vulnerable to non-uniform latency under unbalanced traffic; bandwidth\n bounded by ring stops.\n- **2D Mesh:** the de-facto standard for many-core research chips\n (e.g., academic designs of 64–256 cores; industry \"tiled\" processors).\n Routers and links have fixed cost per tile; diameter grows as O(√N).\n- **2D Torus:** lower diameter than mesh but more links per node.\n- **Hierarchical bus + ring/mesh:** reduces diameter at the cost of\n heterogeneous latency (real in some server designs).\n\n**Proposal for XH-1:** A 2D mesh with adaptive or deterministic X-Y\nrouting, on a per-tile basis. Whether to use a torus topology is an\nunresolved question (§11).\n\n### 4.5 Main memory and the memory controller\n\n- **DDR5/LPDDR5x** with multiple channels is the realistic baseline.\n Channel count and width are project-level decisions (TBD).\n- The MC must be **coherence-aware at the PoC** in the sense that it must\n service the LLC's read/write traffic and respect write-coalescing\n decisions driven by the coherence protocol.\n- The MC's transaction scheduler (FR-FCFS and successors) is not part of\n the CPU hierarchy in the strict sense but is integral to memory\n performance and is mentioned for completeness.\n\n---\n\n## 5. Comparison of Meaningful Alternatives\n\n### 5.1 Private L2 vs. shared L2\n\n- **Private L2 per core** simplifies coherence (L2 is a private point; only\n the L1s of one core need to be tracked at the L2's directory entry).\n Drawback: L2 capacity is fragmented; an 8 MiB private L2 across 128\n cores is 1 GiB of SRAM, which is unrealistic on a single die.\n- **Shared L2 across N cores per tile** is more capacity-efficient and\n improves hit rates for footprints shared by cores in the same tile.\n Drawback: more complex coherence, larger fan-in on the L2's\n bandwidth.\n\n**Proposal:** A shared L2 among a small number of cores (e.g., 4 or 8)\nper tile. The exact tile width is unresolved.\n\n### 5.2 Snoop vs. directory at 128 cores\n\n- A full broadcast snoop on a 128-port medium is generally not realistic.\n However, some research prototypes (e.g., the MIT Alewife and related\n work) used hierarchical broadcast within clusters. The bandwidth\n cost is O(N) per transaction in the cluster; if clusters are small\n (e.g., 8), this is tractable, but the global coherence still needs a\n directory or a hierarchical protocol.\n- **Directory** is the dominant choice in industry for ≥ 16 cores\n (e.g., AMD's \"Probe Filter\" in some products, Intel's \"Home Agent\"\n terminology).\n\n### 5.3 Inclusive vs. non-inclusive\n\n- **Inclusive LLC** (e.g., the policy used in some Intel products prior to\n Skylake-SP and in some AMD products) acts as a snoop filter \"for free.\"\n- **Non-inclusive** (e.g., modern many-core GPUs and accelerators; some\n research many-cores) avoids the back-invalidation cost but requires\n an explicit snoop filter.\n- **Exclusive** is rare in modern designs because it tends to increase\n miss rate for blocks that benefit from being in both levels.\n\n### 5.4 Token coherence, Tempo, and protocol-level alternatives\n\n- **Token coherence** (Marty and Hill, *Token Coherence: Decoupling\n Performance from Correctness*, MICRO 2003) decouples correctness\n from performance in a snooping framework and supports large\n scale. It is a real research alternative for many-core coherence but\n is rarely used in commercial designs.\n- **AMD's MOESI** with probe filters and Intel's \"MESIF\" (where one\n sharer is in F state) are both directory/snoop hybrids in\n practice. RISC-V does not mandate any specific protocol.\n\nFor XH-1, the **proposal** is directory MESI/MOESI. Token coherence is\na documented research alternative but is **not** the recommended\ndefault; we mark it as such for completeness.\n\n### 5.5 Coherent accelerators and CXL\n\n- Modern systems increasingly consider cache-coherent accelerators\n (e.g., CCIX, CXL.cache / CXL.mem). CXL is **not a RISC-V\n requirement** but is an industry direction.\n- If XH-1 has a coherent accelerator port, the protocol must extend\n to the off-chip link. This is a project-level decision and is not\n assumed here.\n\n### 5.6 Inclusion of HBM or stacked DRAM as a memory-side cache\n\n- A die-stacked DRAM (HBM) used as a memory-side cache (e.g., as in\n some research systems) is a published alternative to a large SRAM\n LLC. The bandwidth is high but the latency, cost, and capacity\n constraints are very different from an SRAM LLC.\n- The XH-1 stack-up is not yet known. Marked as a *possible future\n extension*.\n\n---\n\n## 6. Quantitative Considerations\n\nQuantitative claims below are estimates anchored to published sources\nwhere possible. They are **not measured numbers for XH-1**.\n\n### 6.1 Bandwidth scaling\n\n- A simple back-of-envelope: 128 cores, each capable of issuing a\n memory access per cycle, can in aggregate demand a peak of 128\n accesses per cycle. If the LLC is sliced and each slice can serve\n one access per cycle, the LLC bandwidth must be sized accordingly.\n- A 2D mesh of dimension 8×8 has a Manhattan diameter of 14 hops\n (worst case). With a 1 GHz mesh and one flit per cycle, latency\n grows linearly with hops. Adaptive routing mitigates hotspots but\n not the diameter itself.\n- DRAM peak bandwidth, e.g., DDR5-6400 with a 64-bit channel gives\n ~51.2 GB/s per channel. Eight channels give ~410 GB/s. This is\n typically the *system-level* ceiling. Effective bandwidth after\n MC efficiency is well below peak; 50–70% is a typical\n operating range (Micron technical briefs; Rixner et al.,\n \"Memory Access Scheduling\", ISCA 2000).\n\n### 6.2 Coherence traffic\n\n- For directory MESI, the **directory entry** must record the set of\n sharers. A full-bit-vector directory is O(N) per block and is\n infeasible at N=128. A *coarse vector* (e.g., one bit per 4-core\n cluster) reduces storage but causes extra over-invalidation.\n Realistic choices: 32-bit (or similar) limited-pointer directories\n with a fall-back to broadcast when pointers overflow\n (Acacio et al., \"The Director II Shared Memory Multiprocessor\",\n IEEE T-Parallel and Distributed Systems, 2003).\n- Directory storage cost: even at 4 bytes per LLC block, with\n 16 MiB LLC and 64 B blocks, we have 256 K directory entries →\n 1 MiB of directory storage. This is small; the *bandwidth* to\n access the directory is the real cost, motivating banked,\n parallel directories.\n\n### 6.3 Hit rate sensitivity\n\n- Increasing the LLC from 16 MiB to 64 MiB at a fixed workload\n generally improves hit rate. The exact improvement depends on the\n workload's working-set size and sharing pattern. Without a\n workload set, **no number is asserted here**.\n\n### 6.4 Interconnect latency\n\n- Mesh hop latency, in cycles, is typically 2–4 cycles per router\n (depending on flow control and link width). For 8×8, the average\n hop count is ~5 (random uniform) and the worst case is 14. The\n cache miss penalty is dominated by DRAM latency (60–120 ns in\n DDR5; JEDEC JESD79-5) at the far end of the spectrum.\n\n---\n\n## 7. Scalability Problems Introduced by 128 Cores\n\nThis section is the central concern for XH-1. Each item is a known\nissue in the literature; XH-1 must address it explicitly.\n\n### 7.1 Coherence traffic\n\n- With more cores, sharing distances grow. A single shared cache\n block in a directory protocol can incur directory entry\n updates at every transition; the directory access bandwidth\n scales with the number of sharers.\n- **Coherence-induced miss rate** can become significant in\n many-core workloads that share data across many cores\n (Zhao et al., \"Performance Impact and Interplay of Cache\n Sharing and Coherence in Many-Core Processors\", HPCA 2010 —\n cited as an example of this class of work; specific numbers\n depend on workload).\n\n### 7.2 Snoop and probe storms\n\n- Even in a directory scheme, a *broadcast fallback* (when the\n directory overflows) on a 128-core fabric can produce a flood of\n probes. The directory must be sized to avoid this in steady state.\n\n### 7.3 Memory bandwidth saturation\n\n- 128 cores × memory-level parallelism per core can saturate DRAM\n channels. A standard mitigation is to attach enough channels\n (and adequate MC parallelism) and to use prefetching tuned to\n not pollute the LLC.\n\n### 7.4 Cache and directory capacity scaling\n\n- Total on-die SRAM at 128 cores is bounded by die area. A 16 MiB\n LLC at 64 B blocks and 8-way associativity needs ~16 M cells of\n data + tags. This is feasible; 64 MiB is more demanding and\n pushes area and access time.\n\n### 7.5 Verification complexity\n\n- Coherence protocol verification for a 128-core system is\n non-trivial. State-space exploration must be done with\n parameterized cores (e.g., Murφ, TLA+, or commercial\n solutions; see §9).\n\n### 7.6 Non-uniform memory access (NUMA) and tile imbalance\n\n- A mesh creates a NUMA effect: a core's local tile is closer than\n a far tile's slice of the LLC or a remote directory bank. The OS\n and runtime must be NUMA-aware to avoid pathological placement.\n\n### 7.7 Time-to-coherence and ordering\n\n- More cores mean more concurrent coherence transactions; the time\n to reach a stable state grows. The memory-ordering rules of RVWMO\n still apply, but the *latency* of fences and atomics can be large.\n\n---\n\n## 8. Interactions with Other Subsystems\n\n### 8.1 Pipeline\n\n- The hierarchy introduces load-use and store-load *latencies*\n that the pipeline must schedule around. The L1 hit latency\n (typically 3–5 cycles) feeds back into issue logic.\n- Miss status holding registers (MSHRs) and miss buffers at the L1\n sit on the critical path of the load/store unit.\n- Out-of-order cores need store-to-load forwarding; the L1 D-cache\n must support the appropriate forwarding path.\n\n### 8.2 Cache hierarchy\n\n- The hierarchy is the cache hierarchy; this is the topic of\n this document. The relationship to *levels* and the question\n of \"where does the LLC end and DRAM begin\" is settled by the\n PoC definition.\n\n### 8.3 Memory system\n\n- DRAM, MC, and PHY are not in this hierarchy in the strict\n sense, but the LLC ↔ MC link and its bandwidth/latency is\n what limits how aggressively the hierarchy can issue misses.\n- The MC's QoS, request scheduling, and refresh policies are\n *not* under the hierarchy's control but are *visible* to it\n through latency/bandwidth feedback.\n\n### 8.4 Interconnect\n\n- The mesh (or alternative) carries both *coherence* traffic and\n *data* traffic. The two share the same physical links; their\n relative share is a design knob.\n- A coherence-heavy workload can starve data traffic; some\n designs use *separate* physical or virtual channels for\n coherence vs. data.\n\n### 8.5 Coherence\n\n- See §4.3 and §7. The hierarchy's choice of protocol\n determines the *correctness contract* for all memory\n operations.\n\n### 8.6 Interrupts\n\n- Interrupt delivery is per-hart. The hierarchy must ensure that\n any memory state required to service an interrupt is visible\n to the hart receiving the interrupt, and that the *current*\n hart can read interrupt-related structures coherently.\n- For IPIs (inter-processor interrupts), the hierarchy must\n handle concurrent writes to the same interrupt-pending\n structure coherently, atomically if software requires it.\n\n### 8.7 Operating system\n\n- The OS (Linux, in the assumed configuration) expects:\n - **Coherent shared memory** for kernel data and for\n userspace mappings shared between processes.\n - **TLB shootdown IPI support** with a global flush\n mechanism (e.g., via SBI or hypervisor).\n - **Cache flush / invalidate** primitives for drivers and\n kernel code. With Zicbom available, the OS can use\n `CBO.*` instructions; otherwise it falls back to\n `FENCE; virtual-address-range-flush` patterns.\n - **NUMA awareness**: scheduling decisions to keep a\n process and its memory local to a tile.\n - **Memory model awareness**: the OS must use `fence`\n and atomic instructions in a manner consistent with\n RVWMO; e.g., locking primitives must use the A\n extension.\n\n### 8.8 Verification\n\n- **Coherence protocol verification** with formal methods\n (Murφ, TLA+, custom RTL model checking) is necessary\n because random simulation cannot reach corner cases at 128\n cores. The industry practice (e.g., as documented in\n research from Cambridge, UT Austin, and Intel) is to verify\n a *gateway* model and use *assume/guarantee* for\n compositional reasoning.\n- **Memory-model verification** (litmus tests) is required\n to validate the implementation against RVWMO. The RISC-V\n architectural tests and tools (e.g., the\n `riscv-tests` repository) provide a starting point.\n\n### 8.9 Performance\n\n- The hierarchy is the *primary* performance determinant for\n memory-bound workloads. Key levers (each an unresolved\n design parameter for XH-1):\n - L1/L2/LLC sizes and associativities.\n - Replacement policy (LRU, PLRU, RRIP variants; see\n Jaleel et al., \"High Performance Cache Replacement\n Using Re-Reference Interval Prediction (RRIP)\", ISCA\n 2010).\n - Prefetching (next-line, stream, tagged; the cost is\n cache pollution and bandwidth).\n - Coherence protocol choices (MOESI's \"O\" state can\n reduce writebacks for shared-read-only data).\n - Interconnect routing and topology.\n - Memory-controller scheduling (out of scope for this\n document).\n\n---\n\n## 9. Verification Implications\n\nThe verification of the XH-1 hierarchy is itself a significant\nresearch area. Concrete concerns:\n\n- **Protocol-level verification**: A 128-core instance of a directory\n MESI/MOESI is too large to verify by full state exploration with N\n = 128. Compositional verification (verification per core, plus\n per directory bank, plus per interconnect router) is the standard\n industry approach.\n- **Microarchitectural litmus tests**: the RISC-V memory-model\n litmus tests must be adapted to the XH-1 hierarchy. In particular,\n tests that involve fences, atomics, and I/O must pass.\n- **Coherence stress tests**: random traffic with N threads, sharing\n patterns, and adversarial patterns to provoke races, directory\n overflow, livelock, and starvation.\n- **Performance verification**: a cycle-accurate simulator or an\n FPGA prototype is the realistic path; this is a tooling question\n not a hierarchy question per se, but it is mentioned because\n the hierarchy's design depends on being able to simulate it.\n\n---\n\n## 10. Advantages and Disadvantages of the Proposed Approach\n\nSummarizing the *proposal* in §4:\n\n### 10.1 Advantages\n\n- **Directory-based MESI/MOESI** is a well-understood, scalable\n baseline; compatible with industry practice.\n- **Per-tile shared L2** balances capacity and coherence complexity\n better than fully private L2s.\n- **2D mesh interconnect** is a known-scalable topology for\n research many-cores; tools and router designs are well\n documented.\n- **NINE at L2/LLC** avoids the cost of strict inclusion and\n keeps effective capacity high.\n- **Zicbom / Zicboz** availability (if implemented) gives the OS\n a clean, ISA-defined cache-management interface.\n\n### 10.2 Disadvantages\n\n- **Directory storage and bandwidth** are real costs; pointer\n overflow leads to broadcast fallback that must be designed\n for.\n- **Mesh latency** at 8×8 diameter is non-trivial; critical-word\n first is mandatory for high-latency misses.\n- **Verification cost** is high for a 128-core directory\n protocol; this is a project-level risk.\n- **NUMA effects** require careful OS placement to avoid\n pathological performance.\n- **DDR bandwidth** is finite; the design must include bandwidth\n budgeting as a first-class concern.\n\n---\n\n## 11. Unresolved Design Questions\n\nThe following are **explicitly open** for XH-1:\n\n1. **Tile size:** how many cores per L2 / per directory bank?\n Candidates: 1, 2, 4, 8. A 4-core tile is a common choice in\n research many-cores; this is not yet decided.\n2. **Coherence protocol:** MESI vs. MOESI vs. a research\n alternative (e.g., a transient-state directory with F state).\n3. **LLC inclusion policy:** inclusive, exclusive, or NINE.\n4. **Directory organization:** centralized vs. distributed; how\n to handle limited-pointer overflow.\n5. **Interconnect topology:** mesh vs. torus vs. hierarchical;\n routing algorithm (deterministic X-Y, adaptive, oblivious).\n6. **Replacement policy at L2 and LLC:** LRU vs. PLRU vs. RRIP\n variants.\n7. **Prefetching policy:** none, next-line, stream, or tagged.\n8. **I/O coherence:** are devices coherent, non-coherent with\n manual flush, or coherent via an external engine?\n9. **CXL or accelerator coherence port:** in or out of scope?\n10. **HBM or stacked DRAM:** in or out of scope?\n11. **Channel count and width for DDR:** not yet chosen.\n12. **Page size and hugepage support:** interacts with the\n hierarchy via the TLB; out of scope here, but\n acknowledged.\n13. **Zicbom / Zicboz implementation:** assumed, but\n implementation status not yet decided.\n14. **Verification methodology:** formal vs. simulation-heavy;\n the specific tool chain is undecided.\n\n---\n\n## 12. Recommendations (Provisional)\n\nThe following are recommendations *only* to the extent that the\nevidence above supports them. Each is conditioned on\nworkload-driven validation.\n\n- **Use a directory-based MESI or MOESI protocol as the coherence\n baseline.** Both are scalable; MOESI is preferred if shared-read-only\n data is common in the target workload. *This is a proposal;\n workload validation is required.*\n- **Use a per-tile shared L2 of 4 or 8 cores per tile.** *Workload-dependent.*\n- **Use a 2D mesh interconnect** as the default; consider a\n torus only if worst-case latency matters more than\n link area. *Engineering trade-off; no firm recommendation.*\n- **Implement NINE at L2/LLC** with a directory tracking L2\n tags. *Recommended as the default.*\n- **Use Zicbom / Zicboz** for the OS cache-management interface\n if the extension is implemented. *Implementation decision\n pending.*\n- **Plan for formal coherence-protocol verification from the\n start**; the verification methodology must be in scope from\n day one, not retrofitted.\n\n---\n\n## 13. Summary\n\nThe XH-1 memory hierarchy is a *scalable*, *coherent*,\n*cache-based* system on a *128-core* substrate. The RISC-V ISA\nconstrains it to honor RVWMO, the A-extension atomics, the\nZicbom/Zicboz cache-management interface, and the PMP/PMA\nattributes. Beyond that, the design is a series of trade-offs\namong directory vs. snoop, inclusive vs. non-inclusive, mesh vs.\nalternative topology, and MESI vs. MOESI.\n\nThe proposal in this document — directory MESI/MOESI, NINE\nLLC, per-tile L2 of 4–8 cores, 2D mesh — is **not a commitment**\nbut a starting point. The most important next step is to\n*validate against workloads*. Without workload-driven numbers,\nthe quantitative claims in this document are anchors, not\npredictions.\n\n---\n\n## 14. References and Sources\n\nThe following sources informed this document. No citations are\ninvented; sources are listed to indicate the class of evidence\nbehind each claim.\n\n- **RISC-V Unprivileged ISA Manual**, current ratified version,\n Chapter 14 (RVWMO memory model) and Chapter 11 (Zicbom /\n Zicboz).\n- **RISC-V Privileged ISA Manual**, current ratified version\n (PMP, page-based memory, fence semantics).\n- **RISC-V \"A\" Extension**, current ratified version\n (LR/SC and AMOs).\n- **Hennessy, J. and Patterson, D.** *Computer Architecture: A\n Quantitative Approach*, 6th ed., Chapters 2 and 5.\n- **Sweazey, P. and Smith, A. J.** \"A Class of Compatible Cache\n Consistency Protocols and their Verification,\" *IEEE\n Transactions on Computers*, 1986.\n- **Marty, M. R. and Hill, M. D.** \"Token Coherence: Decoupling\n Performance from Correctness,\" *MICRO 2003*.\n- **Jaleel, A. et al.** \"High Performance Cache Replacement\n Using Re-Reference Interval Prediction (RRIP),\" *ISCA 2010*.\n- **Rixner, S. et al.** \"Memory Access Scheduling,\" *ISCA 2000*.\n- **Acacio, M. E. et al.** \"The Director II Shared Memory\n Multiprocessor,\" *IEEE TPDS*, 2003 (illustrative of\n limited-pointer directories; not a direct citation for\n XH-1 numbers).\n- **JEDEC JESD79-5**, DDR5 SDRAM Standard, for DRAM latency\n and bandwidth figures cited.\n- **Micron technical briefs** on DDR5 / LPDDR5x for capacity\n and bandwidth anchors.\n- **riscv-tests / riscv-arch-test repositories** as the canonical\n source for ISA-conformance test methodology.\n- **MIT Alewife project publications** (Agarwal et al., early\n 1990s) for historical context on many-core coherence\n alternatives.\n\nWhere this document makes a numerical claim, the supporting\nsource is identified above. Where no source supports a specific\nnumber, the number is omitted.\n\n---\n\n*End of document. The above is a research document, not a\nspecification. Implementation decisions remain with the XH-1\nproject.*","refusal":null,"reasoning":null}}],"usage":{"prompt_tokens":663,"completion_tokens":7525,"total_tokens":8188,"cost":0,"is_byok":false,"prompt_tokens_details":{"cached_tokens":135,"cache_write_tokens":0,"audio_tokens":0,"video_tokens":0},"cost_details":{"upstream_inference_cost":0,"upstream_inference_prompt_cost":0,"upstream_inference_completions_cost":0},"completion_tokens_details":{"reasoning_tokens":0,"image_tokens":0,"audio_tokens":0}}} |