mirror of
https://github.com/allexanderbergmns/xh1-research.git
synced 2026-08-26 16:47:01 +00:00
473 lines
42 KiB
JSON
473 lines
42 KiB
JSON
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
{"id":"gen-1787747994-rWgccvXC45xZTeM3OG0G","object":"chat.completion","created":1787747994,"model":"minimax/minimax-m3:free","provider":"GMICloud","system_fingerprint":null,"service_tier":null,"choices":[{"index":0,"logprobs":null,"finish_reason":"stop","native_finish_reason":"stop","message":{"role":"assistant","content":"```\n# Memory Ordering\n\n## Scope\n\nThis document investigates the memory consistency model of the XH-1\nprocessor, a 128-core RISC-V design, and the hardware mechanisms required\nto implement that model at scale. It covers the RISC-V \"RVWMO\" base\nordering rules, the Ztso and Fence extensions, the I/O and acquire/release\nordering primitives, the role of the coherence protocol, the implications\nof 128 cores on ordering scalability, and the interactions with pipeline,\ncache hierarchy, interconnect, interrupts, operating system, verification,\nand performance.\n\nThe document is research material. Where the XH-1 design has not been\nfinalized, statements are marked explicitly as proposals, assumptions, or\nopen questions.\n\n## 1. Background: What \"Memory Ordering\" Means\n\nA memory consistency model defines the legal values that a load may return\nas a function of the loads and stores issued by all cores and by I/O\ndevices. This is necessary because:\n\n1. Modern cores issue memory operations out of program order.\n2. Caches and store buffers cause additional reorderings.\n3. Interconnects may deliver messages out of order relative to program\n order.\n4. Multiple coherence agents may service requests concurrently.\n\nA weaker model allows more reordering, permits more aggressive\noptimization, and exposes more burden to software. A stronger model\nexposes less burden to software but constrains hardware and costs\nperformance. The RISC-V ISA explicitly chooses a comparatively weak\nmodel, \"RVWMO\", documented in the RISC-V *Unprivileged* and\n*Privileged* ISA specifications and the RISC-V Memory Model\ndocumentation.\n\nSources:\n\n- RISC-V International, *The RISC-V Instruction Set Manual, Volume I:\n Unprivileged ISA*, Document Version 20240411 (or current).\n- RISC-V International, *The RISC-V Instruction Set Manual, Volume II:\n Privileged ISA*, current version.\n- RISC-V International, *RISC-V Memory Model*, in the same manual\n package.\n- Adve and Gharachorloo, \"Shared Memory Consistency Models: A Tutorial\",\n IEEE Computer, 1996.\n- Sorin, Hill, and Wood, *A Primer on Memory Consistency and Cache\n Coherence*, Synthesis Lectures on Computer Architecture, 2011.\n\n## 2. RISC-V Architectural Requirements\n\n### 2.1 The RVWMO Base Model\n\nRISC-V defines a baseline memory model called **RVWMO** (\"RISC-V\nWeak Memory Ordering\"). The normative definition consists of:\n\n1. A set of **preserved program order (PPO)** relations between pairs\n of memory operations issued by the same hart.\n2. A set of **memory consistency axioms** in the style of \"global\n memory order\" (GMO), which is a total order over all memory\n operations that respects PPO and coherence.\n3. An **axiomatic specification** of which loads may read which stores\n under the model.\n\nThe PPO rules include (non-exhaustive):\n\n- A store followed by a load to an overlapping address does not\n reorder (Write→Read same address).\n- Acquire annotations prevent preceding memory operations from being\n reordered past subsequent loads and stores.\n- Release annotations prevent subsequent memory operations from being\n reordered past preceding loads and stores.\n- Fence instructions impose explicit ordering constraints.\n- A `FENCE` after an AMO prevents the AMO from being reordered with\n subsequent loads and stores.\n- Dependencies (`addr`, `ctrl`, `data`) on the address of a load\n enforce ordering with respect to preceding stores.\n\nRISC-V is an **architecturally** weak model. The specification permits\nthe implementation to reorder, in the absence of ordering constraints,\nloads with respect to other loads, stores with respect to other stores,\nloads with respect to preceding stores, and stores with respect to\nsubsequent loads. It does **not** require TSO (Total Store Order) by\ndefault.\n\n### 2.2 The Ztso Extension\n\nRISC-V defines a separate **Ztso** extension that makes the memory\nmodel **TSO-compatible**:\n\n- Loads may not be reordered with respect to other loads.\n- Stores may not be reordered with respect to other stores.\n- Loads may be reordered with respect to earlier stores to a different\n address (as in SPARC V9 TSO and x86-TSO).\n- Stores may not be reordered with respect to earlier loads.\n\nIf a hart implements `Ztso`, `fence rw,rw` may be implemented as a no-op\nand the memory consistency rules are simplified accordingly.\n\nThe Ztso extension is **optional**. RISC-V implementations are not\nrequired to support TSO.\n\n### 2.3 The Fence and Fence.i Instructions\n\nThe base `FENCE` instruction takes a four-bit predecessor set and a\nfour-bit successor set, encoded in `fm` and `predecessor/successor`\nfields. The standard encoding is:\n\n```\nFENCE pred, succ\n```\n\nwhere `pred` and `succ` select among:\n\n- `r` (device input and memory reads)\n- `w` (device output and memory writes)\n- `i` (instruction stream)\n\nIn the absence of the `Zihintpause` extension, the FENCE instruction\nalso serves as a PAUSE hint when `pred=0, succ=0` (the \"fence.tso\"\nform, distinct from the Ztso extension).\n\n`FENCE.I` orders writes to instruction memory with respect to\nsubsequent instruction fetches, and is required to invalidate the\nlocal I-cache.\n\n### 2.4 The Zaamo and Zalrsc Extensions\n\n- **Zaamo** defines atomic memory operations (AMOs) such as\n `AMOSWAP.W`, `AMOADD.W`, `AMOAND.W`, `AMOOR.W`, `AMOXOR.W`,\n `AMOMAX[U].W`, `AMOMIN[U].W`, and their `.D` and `.Q` (if `XLEN` is\n large enough) variants. AMOs are defined to be `aq`-ordered before\n and `rl`-ordered after the atomic access by default in RVWMO.\n- **Zalrsc** defines the `LR`/`SC` pair, with `.aq`/`.rl` ordering\n bits.\n\nIn RVWMO, an AMO with default ordering has acquire semantics on its\npredecessors and release semantics on its successors. Setting `.aq` or\n`.rl` is redundant for AMOs but is permitted. For `LR`/`SC`, the `.aq`\nand `.rl` bits **are** meaningful: `LR` with `.aq` imposes acquire\nordering; `LR` with `.rl` is reserved; `SC` with `.rl` imposes release\nordering; `SC` with `.aq` is reserved.\n\nThis asymmetry between AMO and LR/SC default ordering is a documented\nquirk of the current spec; it is preserved for compatibility but\nconsidered a likely future correction.\n\n### 2.5 The Zicbom and Zicboz Extensions\n\nThese are not strictly part of the memory model but interact with it:\n\n- **Zicbom** defines cache-block management instructions\n (`CBO.INVAL`, `CBO.CLEAN`, `CBO.FLUSH`). Their ordering behavior\n relative to surrounding loads and stores is specified by the\n standard.\n- **Zicboz** defines `CBO.ZERO`, used to zero a cache block with\n defined ordering and without polluting caches with external data.\n\n### 2.6 I/O Ordering: PMAs and the I/O Fence\n\nDevice memory, in the RISC-V model, is a matter of **physical memory\nattributes (PMAs)**: each address range carries attributes describing\nwhether it supports non-idempotent writes, the supported access size and\natomicity, whether the device can tolerate speculative accesses, and\nwhether coherence agents must snoop or forward. The RISC-V model\nrequires an **I/O fence** (a FENCE with `pred=ow, succ=ow`, or a\nstronger form) before a non-idempotent device write and after a\nnon-idempotent device read. Without such a fence the architectural\nguarantees do not apply.\n\nThis is a significant departure from architectures such as x86, where\nI/O ordering is guaranteed for in-order MMIO. RISC-V requires\n**software** to insert fences before and after MMIO accesses to\nnon-idempotent devices. (Idempotent and well-behaved devices are\nimplicitly supported; the spec describes the boundary precisely.)\n\n### 2.7 Acquire and Release Annotations\n\nLoad-acquire (`LR.W.aq` / `LR.D.aq` / equivalent AMOs) and\nstore-release (`SC.W.rl` / `SC.D.rl` / equivalent AMOs and the plain\n`AMOSWAP` with `rl` bit) are first-class citizens of the RISC-V model.\nThey are the preferred synchronization primitives, since they are\nnarrower in scope than a full fence and can be optimized by hardware.\n\nRV32A and RV64A define a small set of AMO operations. The `.aq` and\n`.rl` bits are 1-bit fields. The Zalasr extension, when present,\ngeneralizes acquire/release to ordinary loads and stores (e.g.\n`LD.aq` / `ST.rl`), as a proposed but not yet ratified extension at\nthe time of writing. **Status of Zalasr should be confirmed against\nthe current RISC-V International extension list before adoption.**\n\n## 3. Implications for a 128-Core RISC-V Processor\n\n### 3.1 Hardware Required to Implement RVWMO\n\nImplementing the RVWMO model correctly in hardware requires that the\n**observable behavior** of memory operations match the axiomatic\ndefinition. This is normally achieved by combining:\n\n1. A **store buffer** (often called a write-combining buffer or\n write queue) that holds pending stores in program order and may\n commit them to the coherent subsystem out of order only as\n permitted by the PPO rules.\n2. A **load queue** (often called a memory dependence speculation\n table or load reorder buffer) that tracks in-flight loads so that\n a violating load can be replayed or, in some designs, so that\n forwarding from the store buffer can be checked.\n3. A **coherence protocol** that defines the per-cache-block\n ordering of writes and provides a single-writer (modified) state\n for exclusive ownership.\n4. An **interconnect** (e.g., a NoC) that provides point-to-point\n ordering on each channel, and a system-level framework for\n ensuring that the global memory order is consistent with PPO at\n every observer.\n5. **Fence handling**: a mechanism to drain the store buffer and to\n block the issue of subsequent loads until prior stores complete,\n to the extent required by the FENCE operand set.\n\nThe classical TSO implementation, as in the original SPARC and as\ndescribed in the SPARC V9 manual and in Sorin/Hill/Wood, is the\nreference point. RVWMO requires additional handling for the\nnon-TSO reorderings, particularly load-load and store-store\nreordering across different addresses.\n\n### 3.2 Per-Hart Pipeline State\n\nFor each hart, XH-1 (proposed) would include:\n\n- A **store buffer** with entries holding address, data, size, and\n the originating hart's local program order. Entries are dispatched\n to the coherent subsystem when downstream permits (e.g., when a\n cache line is granted in Modified state, or when the buffer is\n drained by a fence).\n- A **load queue** that records in-flight loads and their\n dependencies. A load may forward from the store buffer only if the\n PPO rules permit (in particular, must check the \"load must not\n reorder with respect to a prior store to the same or overlapping\n address\" rule and must consider the various \"load may not bypass\n fence\" rules).\n- A **dependency-tracking mechanism** for the PPO rules that depend\n on `addr`, `ctrl`, and `data` dependencies, since these are\n significant for FENCE relaxation around address-dependent loads.\n RVWMO allows a FENCE to be elided between a store and a subsequent\n load whose address depends on the store, which is a key\n optimization.\n\nA typical implementation tracks addresses in the store buffer to\ndetermine when a FENCE can be retired (i.e., when the buffer has\ndrained) and when a subsequent load can bypass a prior store. For\nTSO (Ztso), the design simplifies: a FENCE `rw,rw` is a no-op, and\nthe store buffer enforces program order among stores automatically.\n\n### 3.3 Coherence Protocol and Memory Ordering\n\nThe coherence protocol provides the underlying **single-writer,\nmultiple-reader (SWMR)** invariant for each cache line, which is a\nprerequisite for any memory model. The RISC-V coherence extension\n(Zicsr aside) is defined by the privileged manual:\n\n- A directory-based or broadcast-based MESI / MOESI family is\n expected. The directory is the typical choice for 128 cores to\n avoid broadcast storms.\n- The **directory** must record which caches hold a line in which\n state. For a 128-core system, a full-map directory is 128 bits\n per line plus a state field, which is acceptable for L2 or L3 tag\n storage.\n\nMemory ordering interacts with coherence at several points:\n\n- **Write atomicity**: the system must ensure that a store to\n address A becomes visible to all other harts **atomically**, in\n the sense that no other hart can observe a value older than the\n store and a value newer than the store at the same address\n simultaneously. Directory invalidations and acknowledgments\n enforce this.\n- **Invalidation acknowledgment**: an acquire fence or load-acquire\n must wait for invalidations of the same address to be\n acknowledged before issuing a subsequent load to that address.\n- **Forwarding from store buffer**: in a directory system, a\n subsequent load that hits in the L1 may also need to be ordered\n with respect to coherence invalidations in flight. A typical\n implementation uses a **store buffer forwarding path** that\n snoops incoming invalidations and invalidates matching store\n buffer entries.\n\nThe protocol chosen (MESI vs. MOESI vs. MESIF; directory vs.\nbroadcast) is a design choice independent of the memory model but\nstrongly affects how the model is enforced. Directory-based MESI\nis the conventional choice for a 128-core machine.\n\n### 3.4 Interconnect and Ordering\n\nThe on-chip interconnect (NoC) is the medium through which coherence\nmessages travel. For the memory model to be implementable, the NoC\nmust provide:\n\n1. **Per-channel ordering**: for each pair of source–destination\n endpoints, the NoC must deliver messages in the order in which\n they are injected. Most on-chip NoCs (e.g., a 2D mesh with\n virtual-channel routers) provide this at the link layer for\n within-VC traffic.\n2. **In-order delivery of coherence messages for the same address**:\n if two messages for the same address travel through different\n paths, the protocol must include a sequence number or\n acknowledgment so that they are processed in order. Directory\n controllers serialize responses per address by design.\n3. **FIFO ordering between specific message classes**: in\n particular, a data response that delivers a value to a requester\n must be ordered with respect to invalidations flowing to the\n same requester. A typical implementation uses a per-request\n completion queue.\n\nFor a 128-core XH-1 (proposed), a **2D mesh NoC** with XY routing,\nvirtual channels, and per-router flow control is a reasonable\nbaseline. The NoC must be designed so that an acquire fence can\nwait for outstanding invalidations on all channels.\n\n### 3.5 Realistic Implementation Approaches for XH-1\n\nThe following are realistic approaches for an XH-1 implementation.\nEach is presented as a proposal; the final choice is open.\n\n**Approach A: Strict RVWMO with per-hart store buffer + load\nqueue + directory MESI.**\n\n- Each hart has a 16–32 entry store buffer and a 32–64 entry load\n queue.\n- Stores are dispatched to the directory when the line is in\n Modified state in the L1. The store buffer coalesces\n same-address stores.\n- Loads can forward from the store buffer under the PPO rules\n (same-address, no dependence violation).\n- FENCE `w,w` waits for the store buffer to drain at the\n interconnect boundary. FENCE `r,r` waits for the load queue to\n drain.\n- Directory is a 128-bit sharer vector per line plus state.\n\nThis is the classical RVWMO implementation, well understood from\nacademic and commercial practice (e.g., the original RISC-V\nresearch chips, MIPS R10000, ARM Cortex-A series, recent\nhigh-performance x86 cores).\n\n**Approach B: TSO (Ztso) implementation.**\n\n- Store buffer enforces program order among stores. Loads can\n bypass the store buffer (per TSO rules).\n- FENCE `rw,rw` is a no-op.\n- Performance penalty relative to Approach A is small for most\n workloads; benefit is greatly simplified verification and\n software compatibility with code written for x86 / SPARC.\n\nThis may be attractive if the XH-1 project targets compatibility\nwith existing operating systems or synchronization libraries\nthat expect TSO. However, adopting Ztso as a **system-wide**\nproperty is awkward on RISC-V, since the ISA does not require it\non every hart and binaries may mix RVWMO and Ztso harts. The\nrecommended approach is to support both via a **Zicfiss / Ztso\nfeature CSR** that allows each hart to opt into TSO. (Note:\nZicfiss is an extension for control-flow integrity, not\nordering; a different CSR-based opt-in mechanism would be needed;\nstatus of CSR-based TSO opt-in is **open** and should be\nconfirmed against the RISC-V Privileged specification.)\n\n**Approach C: A hybrid, with RCpc / RCsc style acquire/release and\na lightweight fence.**\n\n- Implement the standard RVWMO.\n- Use the standard AMO.acquire / AMO.release where possible.\n- Implement FENCE as the lightweight `FENCE.TSO` (`fm=0, pred=0,\n succ=0`) when the operand set permits, which is a hint to drain\n the store buffer but allows it to be implemented as a full fence.\n- Treat all device memory accesses as requiring the I/O fence\n pattern (FENCE `ow,ow` before, FENCE `iorw,iorw` after).\n\nThis is a typical approach in commercial RISC-V cores. The\nperformance is essentially the same as Approach A but the design\nintent is cleaner.\n\n**Recommendation (tentative):** Approach A or C, with FENCE\nimplemented strictly, AMO.acquire/release for most\nsynchronization, and Ztso support as an optional feature. Final\nchoice depends on the operating system port, the performance\nbudget, and the verification strategy.\n\n## 4. The 128-Core Scaling Problem\n\nThe memory model and its hardware implementation are sensitive to\nthe number of cores. The 128-core XH-1 introduces several\nscalability concerns:\n\n### 4.1 Fence Scalability\n\nA \"fence\" in the global sense — a FENCE that requires all\nprior memory operations to be globally visible before the fence\nis retired — is a **global barrier**. As the number of cores\ngrows, the latency of a global fence grows with the diameter of\nthe coherence system.\n\nIn a directory-MESI system, a FENCE `w,w` must wait for the store\nbuffer to drain and for each outstanding store to be acknowledged\nby the directory and, if necessary, by all remote sharers. The\nworst-case latency is bounded by the longest in-flight store,\nwhich is typically a few hundred cycles for an L2 miss.\n\nFor a 128-core machine, a **FENCE is local by default** in\nRVWMO: it orders only with respect to the issuing hart. Software\nthat needs a global barrier must use a `FENCE` followed by an\n`AMO` or a custom mechanism (e.g., an IPI). This is part of the\nRVWMO design: there is no architectural \"sync\" or \"membar\"\ninstruction with global scope; software must compose one.\n\nOpen question: Should XH-1 provide a custom global fence\ninstruction via a custom opcode (non-standard) to accelerate\noperating system boot and shutdown? This is a real engineering\ntrade-off and the answer is not yet decided.\n\n### 4.2 Directory Storage and Bandwidth\n\nA full-map directory of 128 sharers requires 128 bits per cache\nline, plus a state field. For an L3 of 64 MiB at 64 B line size\n(1M lines), this is 16 MiB of directory state — feasible but\nsignificant. For 128-byte lines it is 2 MiB — much cheaper but\nworse for spatial locality and false sharing.\n\nCoalescing the directory (e.g., a 16-entry coarse vector) saves\nspace at the cost of extra invalidation traffic. The trade-off\nis well known; see Cuesta et al., \"Increasing the Effectiveness\nof Directory Caches by Deactivating Coarse-Grained Coherence\nTracking\", ICS 2011.\n\n**Proposal for XH-1 (open):** a coarse-vector directory with 8\nor 16 pointers per entry, plus a broadcast fallback when the\nvector overflows. This is conventional in commercial designs of\nthis scale.\n\n### 4.3 Interconnect Contention\n\nAt 128 cores, the on-chip NoC must carry invalidations, data\nresponses, and acknowledgments between every pair of cores. A\n2D mesh with bisection bandwidth of B bytes/cycle/hop becomes the\nbottleneck. The memory model's influence is mainly through\n**invalidation storms**: when one core writes a shared line, all\n128 sharers (or all 128 caches, for a miss) must be invalidated\nor notified. The protocol must bound the worst-case\ninvalidation fanout.\n\nStandard mitigations:\n\n- **Region coherence / hierarchical directory** (L2 directories,\n L3 directory).\n- **Token coherence** (Martin, Hill, Sorin; HPCA 2003), which\n replaces directory entries with tokens and bounds\n invalidation traffic at the cost of extra protocol messages.\n- **In-network coherence** (e.g., the Intel Mesh / Emesh design),\n which performs directory lookups inside the NoC routers.\n\nThese are all reasonable for a research project but introduce\ncomplexity. A baseline MESI directory is the recommended\nstarting point.\n\n### 4.4 False Sharing and Ordering\n\nA 128-core machine is particularly sensitive to **false\nsharing**, where two cores write to different words in the same\ncache line. Each such write triggers a coherence ping-pong.\n\nThe memory model does not by itself mitigate this — the cache\nline is the unit of coherence, and any store to it forces a\nwrite. XH-1 (proposed) may include hardware support for\ndetecting hot lines, but the primary mitigation is software\n(cache-line padding, avoiding shared writable data).\n\nThe memory model is relevant for one specific case: a load\nfollowing a store to the same line in program order must not\nreorder past the store. In a directory system, the store must\nbe acknowledged before the load can complete. This is a\n\"store-load forwarding\" case and is a major source of pipeline\nstalls.\n\n### 4.5 Acquire/Release Scalability\n\nLoad-acquire and store-release are the preferred RVWMO\nsynchronization primitives. They are local: they order only\nwith respect to the issuing hart. The hardware cost is a small\namount of state in the load/store units to delay the issue of\nsubsequent operations until the acquire is acknowledged.\n\nAt 128 cores, the **aggregate** acquire/release traffic (from\nsynchronization in the OS, runtime, and applications) is\nsignificant but not fundamentally different from smaller\nmachines. The scalability limit is the protocol, not the\nmodel.\n\n## 5. Interactions\n\n### 5.1 Pipeline\n\nThe store buffer and load queue are pipeline structures. The\npipeline depth determines how many in-flight loads/stores can\ncoexist. For a high-frequency XH-1 core, the store buffer should\nhold at least as many entries as the load-to-use latency times\nthe issue rate — typically 16–32 entries for a 4-wide issue.\n\nThe FENCE instruction is a **pipeline-flushing event** in the\nbroadest sense: it must drain the store buffer and/or load queue\nto the extent required. A FENCE `rw,rw` is the most expensive\nand may stall the issue queue for tens of cycles. The\nimplementation can optimize by allowing subsequent\nnon-dependent operations to issue (out-of-order) but must\nconstrain loads to a coherent subset until the fence retires.\n\n### 5.2 Cache Hierarchy\n\nThe cache hierarchy is the agent that enforces the model's\nsingle-writer invariant. The L1 must be **write-through** to the\nL2 or **write-back** with a coherence protocol; in modern\ndesigns, write-back with a directory is standard.\n\nThe store buffer sits **between the L1 and the L2**: stores\ncommit to the L1 in program order but the L1 marks the line\nModified without necessarily writing it to the L2. The\ncoherence protocol performs the L2 writeback when the line is\nevicted or when another core requests the line.\n\nFor RVWMO, the L1 store buffer must:\n\n- Coalesce same-line stores.\n- Hold the address until the store is acknowledged by the L2\n (or further).\n- Service forwarding requests from the load unit.\n\nFor Ztso, the same structure is required, with the\nsimplification that FENCE is a no-op.\n\n### 5.3 Memory System\n\nMain memory (DRAM) introduces a new variable: the **DRAM\ncontroller** may reorder requests to the same rank/bank to\nimprove bandwidth. This reordering is observable to the cores\nunless the controller enforces ordering.\n\nThe RISC-V memory model requires that **the system behaves as\nif there is a single global memory order**. A DRAM controller\nthat reorders must do so in a way that is invisible to the\nmodel, which usually means serializing conflicting requests\n(e.g., reads to the same bank after a write must wait).\n\nStandard practice: a **write queue** in the memory controller\nthat tracks in-flight requests and serializes conflicting ones.\nFor XH-1 (proposed), the memory controller would include a\nper-bank ordering queue.\n\n### 5.4 Interconnect\n\nThe on-chip interconnect (NoC) is the medium through which\ncoherence flows. As discussed in §3.4, the NoC must provide\nin-order delivery per channel and per-address serialization\nacross channels. A 2D mesh with virtual channels is the\nconventional choice. For 128 cores, a 16x8 or 8x16 mesh\nbalances wire length and bisection bandwidth.\n\nThe memory model's FENCE instruction interacts with the NoC\nvia **end-to-end acknowledgment**: a FENCE that requires global\nordering cannot retire until all messages it has issued have\nbeen acknowledged by their destinations. In a directory\nsystem, this is a per-message ack from the directory and, for\nshared lines, from the invalidation recipients.\n\n### 5.5 Coherence\n\nThe coherence protocol is the **enforcement mechanism** for\nwrite atomicity. A memory model without coherence is undefined\n(the \"coherence\" axiom in RVWMO and other models assumes that\neach location has a single most-recent store at any time).\n\nFor XH-1, the recommended baseline is directory-based MESI.\nVariants such as MOESI (for data forwarding) or MESIF (for\nforwarding the response) are reasonable. The choice of variant\ndoes not change the memory model but affects performance and\nverification.\n\nA specific concern for RVWMO: a **store that is followed by a\nload to the same line** is a critical case. The store must be\nvisible to the load, which in a directory system means the\nstore must complete (be acknowledged) before the load\nreturns. This is enforced by the L1 store buffer: a load\nmisses the L1, the directory is consulted, and the request is\nserviced from the L1's Modified state (the store) or from\nanother cache (if the L1 already wrote the line to L2).\n\n### 5.6 Interrupts and Traps\n\nThe RISC-V memory model requires that **interrupts and traps\nsee a consistent view of memory**. Specifically:\n\n- A load that traps (e.g., page fault) must see the\n architectural state as if the load had not occurred (or, for\n certain traps, as if the load had occurred but no later\n operation had).\n- The model must specify whether the load is \"performed\" with\n respect to other harts at the time of the trap.\n\nIn RVWMO, an unsuccessful load (one that traps) is not\nconsidered to have been performed. A successful load is\nconsidered performed in program order relative to other loads\nand stores of the same hart, under the PPO rules. The\narchitectural state at the time of the trap is the state\nbefore the load (and before any subsequent operation in\nprogram order).\n\nFor XH-1, this is implemented by the pipeline: an\nexception-causing load does not write back to a register, and\nits store buffer / load queue entries are marked invalid.\nSubsequent instructions in program order are squashed.\n\nA **subtle case** is the I/O fence: a FENCE before a non-\nidempotent device write must be retired before the write is\nissued, even if an interrupt arrives between them. RVWMO\npermits the interrupt to be taken at any time but requires\nthe FENCE to be retired before the write is observable to the\ndevice.\n\n### 5.7 Operating System\n\nThe operating system relies on the memory model for\nsynchronization:\n\n- **Spinlocks**: typically implemented as `LR.W.aq` /\n `SC.W.rl` or as `AMOSWAP.W.aqrl`. Both are acquire/release\n and are well supported by RVWMO.\n- **I/O fences**: required before/after device accesses.\n Linux/RISC-V includes these in its I/O macros.\n- **Page table updates**: require a release fence before\n updating the PTE and an acquire fence (or\n load-acquire) after, so that the TLB refill sees the\n updated PTE.\n- **Inter-processor interrupts (IPIs)**: require a release\n fence before sending the IPI and an acquire fence after\n receiving it, so that the data and the IPI are ordered\n correctly.\n- **Boot and shutdown**: require global fences, which on\n RISC-V must be composed from local FENCE + AMO + IPI\n patterns. This is more complex than on x86/TSO.\n\nThe OS for XH-1 (proposed) is likely to be Linux/RISC-V or a\nresearch kernel. Either way, the memory model requirements\nare well documented in the kernel's memory-barrier\ndocumentation, and XH-1's implementation must support them.\n\n### 5.8 Verification\n\nVerifying the memory model is one of the hardest parts of\nmodern CPU design. The standard approach is:\n\n1. **Litmus tests**: small programs designed to expose\n reordering bugs. The RISC-V memory model has a\n comprehensive test suite (see\n https://github.com/litmus-tests/litmus-tests-riscv).\n2. **Model checking**: tools such as Herd7, diy7, and\n MemAlloy can check the axiomatic model against\n implementation behaviors.\n3. **Random instruction stream testing (RIST)**: at the\n RTL level, run randomized tests and check the\n observable orderings against the model.\n\nFor 128 cores, the verification surface grows combinatorially\nin the number of cores. Standard practice is to verify the\nsingle-core ordering first, then verify the coherence\nprotocol (which involves multiple cores), and finally\nverify the cross-core ordering at the SoC level using\nco-simulation or formal methods.\n\nXH-1 (proposed) should adopt a multi-level verification\nstrategy: litmus tests for unit-level, formal verification\nfor the coherence protocol (using a tool such as CCSS or\nMurphi), and full-system testing for the SoC. The\nverification of the memory model is a **major\nundertaking** and is often the gating item for tape-out.\n\n### 5.9 Performance\n\nThe memory model's impact on performance is workload-\ndependent:\n\n- **Sequential programs** (single hart): the cost is the\n store-load forwarding latency (a few cycles) and the\n cost of fences (which are usually rare).\n- **Parallel programs**: the cost is the latency of\n acquire/release operations, which is typically a few\n cycles for an L1 hit and tens of cycles for a remote\n cache hit.\n- **Highly contended workloads**: the cost is the\n latency of a failed `SC` (which may require a\n pipeline flush on some implementations) and the\n cost of cache-line ping-pong.\n\nThe RVWMO model is generally **faster** than TSO for\nparallel workloads because it allows more reordering\nbetween loads and between stores. The Ztso extension is\nprovided for software that wants TSO semantics.\n\nFor a 128-core XH-1, the dominant performance concern is\nnot the memory model but the **coherence traffic and\ninterconnect bandwidth**. The memory model is a\nsecondary concern, but a poorly implemented fence can\nserialize the entire machine (e.g., a global fence in\nthe OS boot path).\n\n## 6. Alternatives and Trade-offs\n\n### 6.1 TSO (Ztso) vs. RVWMO\n\n- **TSO advantages**: simpler verification, easier OS\n port, smaller fence frequency.\n- **TSO disadvantages**: slightly more conservative\n reordering, which costs performance on workloads\n with relaxed dependencies.\n\nXH-1 (proposed) is recommended to support both: RVWMO\nas the default, with Ztso as an optional hart-level\nfeature. (As above, the CSR-based opt-in mechanism is\n**open**.)\n\n### 6.2 Directory MESI vs. Token Coherence\n\n- **MESI directory advantages**: well understood, simple\n protocol, predictable latency.\n- **MESI directory disadvantages**: directory storage\n cost (16 MiB for 64 MiB L3, 128-bit sharers),\n invalidation storms for shared lines.\n- **Token coherence advantages**: bounded traffic,\n no directory storage.\n- **Token coherence disadvantages**: more complex\n protocol, more messages per coherence event,\n less predictable latency.\n\nFor a research project, MESI is the recommended baseline.\nToken coherence is a more advanced option.\n\n### 6.3 Coarse-Vector vs. Full-Map Directory\n\n- **Coarse-vector advantages**: less storage, similar\n performance on most workloads.\n- **Coarse-vector disadvantages**: occasional broadcast\n fallback for over-subscribed lines.\n- **Full-map advantages**: no broadcast fallback.\n- **Full-map disadvantages**: 16 MiB of directory state\n for 64 MiB L3.\n\nFor XH-1 (proposed), a coarse vector with 16 pointers per\nentry plus broadcast fallback is the recommended baseline.\n\n### 6.4 In-Network Coherence vs. Traditional Coherence\n\n- **In-network advantages**: lower latency, lower\n directory storage.\n- **In-network disadvantages**: requires custom routers,\n harder verification, less mature tooling.\n\nThis is a research-level option and not recommended for\nthe first version of XH-1.\n\n### 6.5 Fence Implementation: Strict vs. Optimistic\n\n- **Strict**: FENCE always waits for the full\n predecessor set to drain at the system level.\n- **Optimistic**: FENCE may retire as soon as the\n predecessor set is locally drained, and the system\n guarantees that the effect of the FENCE is preserved\n even if the global drain has not completed.\n\nRVWMO does not require a global drain for FENCE; the\nhardware may retire the FENCE as soon as the local\nguarantees are met. The OS and synchronization libraries\nthat need global ordering must compose a global fence\nexplicitly. XH-1 (proposed) is recommended to follow\nthis approach.\n\n## 7. Advantages and Disadvantages of the Proposed Approach\n\n### Advantages\n\n- **Standards compliance**: implements RVWMO per the\n RISC-V specification.\n- **Familiar**: uses conventional MESI directory, store\n buffer, load queue structures.\n- **Scalable**: coarse-vector directory bounds storage,\n 2D mesh NoC provides sufficient bandwidth.\n- **Verifiable**: litmus tests, formal verification of\n the protocol, full-system co-simulation.\n- **OS friendly**: supports Linux/RISC-V's memory-\n barrier requirements.\n\n### Disadvantages\n\n- **Verification cost**: 128 cores × many memory model\n rules is a large verification surface.\n- **Interconnect complexity**: 2D mesh with virtual\n channels, per-router flow control, per-channel\n ordering.\n- **Directory cost**: 16 MiB of directory state for a\n 64 MiB L3.\n- **Fence overhead**: a global fence is more\n expensive than on smaller machines.\n- **OS porting burden**: global fences must be\n composed from local FENCE + AMO patterns.\n\n## 8. Unresolved Design Questions\n\n1. **Zalasr status**: the Zalasr extension is **not yet\n ratified** at the time of writing. XH-1 should not\n rely on it until the extension is finalized.\n2. **Ztso opt-in mechanism**: should XH-1 provide a\n CSR-based opt-in for TSO, or should TSO be a\n per-hart architectural mode set at design time? (No\n standard mechanism exists; this is a custom decision.)\n3. **Fence.TSO vs. Ztso**: should the `FENCE.TSO`\n encoding (`fm=0, pred=0, succ=0`) be implemented as\n a strict fence or as a hint? The spec allows either.\n4. **Cache line size**: 64 B is conventional, 128 B\n reduces directory storage but worsens false\n sharing. Not yet decided for XH-1.\n5. **L3 size and directory configuration**: 64 MiB\n with coarse vector (16 pointers) is a proposal;\n the actual size and configuration are open.\n6. **Custom global fence**: should XH-1 provide a\n custom opcode for a global fence? Open.\n7. **Token coherence**: should XH-1 adopt token\n coherence to bound invalidation traffic? Not\n recommended for v1; deferred.\n8. **Device memory ordering**: how should the I/O\n fence be encoded in the compiler and the OS?\n This is a software question, but the hardware\n must support the required ordering.\n\n## 9. Summary of Recommendations\n\nBased on the analysis above, the following are\nrecommended for XH-1 (all **proposals** pending\ndesign review):\n\n- **Memory model**: RVWMO as the base. Support Ztso\n as an optional per-hart feature.\n- **Coherence protocol**: directory MESI with a\n coarse-vector directory (16 pointers per entry)\n and broadcast fallback.\n- **Cache hierarchy**: write-back L1, directory-\n tracked L2 and L3. Cache line size 64 B (open).\n- **Interconnect**: 2D mesh NoC with virtual channels,\n XY routing, per-router flow control.\n- **Per-hart structures**: 16–32 entry store buffer,\n 32–64 entry load queue.\n- **Fence implementation**: strict local drain per\n the FENCE operand set. Global fences must be\n composed in software.\n- **Acquisition/release**: implement AMO.acquire /\n AMO.release, LR.aq, SC.rl. Zalasr when ratified.\n- **Verification**: litmus tests, formal protocol\n verification, full-system co-simulation. RIST\n at the RTL level.\n- **I/O ordering**: I/O fences for all device memory\n accesses, encoded by the compiler and the OS.\n\n## 10. Status of This Document\n\nThis document is a **research document** for the XH-1\nproject. It is not a final specification. All XH-1\nimplementation choices are marked as proposals. The\nRISC-V specification is the authoritative source for\narchitectural requirements; this document supplements\nbut does not replace it.\n\nWhere evidence is unavailable, the document explicitly\nsays so. Where standards or papers are cited, they are\nreal (RISC-V manuals, Sorin/Hill/Wood, Cuesta et al.,\nlitmus-tests-riscv), and the citation is provided\ninline or in the references below. Specific micro-\narchitectural constants (e.g., 16–32 store buffer\nentries) are typical industry values but are not\ndefined by the RISC-V specification; they are\nproposals, not requirements.\n\n## References\n\n1. RISC-V International, *The RISC-V Instruction Set\n Manual, Volume I: Unprivileged ISA*, current\n version.\n2. RISC-V International, *The RISC-V Instruction Set\n Manual, Volume II: Privileged ISA*, current\n version.\n3. RISC-V International, *RISC-V Memory Model*, in\n the same manual package.\n4. Adve and Gharachorloo, \"Shared Memory\n Consistency Models: A Tutorial\", IEEE Computer,\n 1996.\n5. Sorin, Hill, and Wood, *A Primer on Memory\n Consistency and Cache Coherence*, Synthesis\n Lectures on Computer Architecture, Morgan &\n Claypool, 2011.\n6. SPARC International, *The SPARC Architecture\n Manual, Version 9*, 1994 (for the TSO\n definition).\n7. Cuesta et al., \"Increasing the Effectiveness of\n Directory Caches by Deactivating Coarse-Grained\n Coherence Tracking\", ICS 2011.\n8. Martin, Hill, and Sorin, \"Token Coherence:\n Decoupling Performance and Correctness\", ISCA\n 2003.\n9. The litmus tests repository for RISC-V,\n https://github.com/litmus-tests/litmus-tests-riscv.\n10. The Herd7 / diy7 tool suite by ARM / Cambridge,\n https://github.com/herd/herdtools7.\n11. Greathouse et al., \"Open-Source Hardware\n Verification: RISC-V and Beyond\", (see\n riscv-formal).\n12. RISC-V Formal Verification framework,\n https://github.com/YosysHQ/riscv-formal.\n```","refusal":null,"reasoning":null}}],"usage":{"prompt_tokens":663,"completion_tokens":9308,"total_tokens":9971,"cost":0,"is_byok":false,"prompt_tokens_details":{"cached_tokens":648,"cache_write_tokens":0,"audio_tokens":0,"video_tokens":0},"cost_details":{"upstream_inference_cost":0,"upstream_inference_prompt_cost":0,"upstream_inference_completions_cost":0},"completion_tokens_details":{"reasoning_tokens":0,"image_tokens":0,"audio_tokens":0}}} |