diff --git a/research/.xh1/logs/harness.log b/research/.xh1/logs/harness.log index b7f6888..89b74ac 100644 --- a/research/.xh1/logs/harness.log +++ b/research/.xh1/logs/harness.log @@ -459,3 +459,18 @@ [2026-08-26T12:56:50Z] Response received: 1780 bytes [2026-08-26T12:56:50Z] Research passed review. [2026-08-26T12:56:50Z] Accepted research document: research/05-memory/memory-hierarchy.md +[2026-08-26T12:57:40Z] Started run: 20260826T125740Z +[2026-08-26T12:57:40Z] ================================================== +[2026-08-26T12:57:40Z] Researching: research/05-memory/memory-architecture.md +[2026-08-26T12:57:40Z] ================================================== +[2026-08-26T12:57:40Z] Research round 1/3 +[2026-08-26T12:57:40Z] Running researcher. +[2026-08-26T12:57:40Z] API: provider=openrouter model=minimax/minimax-m3:free attempt=1 +[2026-08-26T12:58:59Z] Response received: 28405 bytes +[2026-08-26T12:58:59Z] Running reviewer. +[2026-08-26T12:58:59Z] API: provider=ollama model=gemma4:12b attempt=1 +[2026-08-26T13:02:07Z] WARNING: Provider returned an empty/near-empty response (0 bytes). +[2026-08-26T13:02:09Z] API: provider=ollama model=gemma4:12b attempt=2 +[2026-08-26T13:02:54Z] Response received: 2234 bytes +[2026-08-26T13:02:54Z] Research passed review. +[2026-08-26T13:02:54Z] Accepted research document: research/05-memory/memory-architecture.md diff --git a/research/.xh1/responses/20260826T125740Z-20260826T125740Z-32388.json b/research/.xh1/responses/20260826T125740Z-20260826T125740Z-32388.json new file mode 100644 index 0000000..9b92bbc --- /dev/null +++ b/research/.xh1/responses/20260826T125740Z-20260826T125740Z-32388.json @@ -0,0 +1,375 @@ + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +{"id":"gen-1787749061-v4qDwmN5u7ZMB1i3n67T","object":"chat.completion","created":1787749061,"model":"minimax/minimax-m3:free","provider":"GMICloud","system_fingerprint":null,"service_tier":null,"choices":[{"index":0,"logprobs":null,"finish_reason":"stop","native_finish_reason":"stop","message":{"role":"assistant","content":"```\n# Memory Architecture\n\n## Status\n\nThis document is part of the XH-1 research series. It describes the\nmemory system proposal for the XH-1, a custom 128-core RISC-V\nprocessor. Where implementation choices have not been finalized, the\ntext uses **PROPOSAL** to mark a design that is being considered but\nhas not been committed. Where the text is an **ASSUMPTION**, evidence\nis not yet available and the statement must be re-validated before\ndownstream documents depend on it.\n\n## Scope\n\nThis document covers:\n\n- The physical memory map and address space layout visible to software\n running on XH-1 cores.\n- The relationship between the RISC-V privileged architecture and the\n memory system, including `satp`, SFENCE.VMA, and the PMP.\n- The on-chip memory hierarchy: L1, L2, and the proposed shared L3.\n- The directory-based coherence protocol between L2 slices.\n- The interconnect that ties cores, caches, and DRAM controllers\n together.\n- The DRAM controller and the memory-side bandwidth budget.\n- I/O, interrupts, and operating system interactions that constrain\n memory-system design.\n- Verification and performance considerations specific to a 128-core\n tiled implementation.\n\nThis document does not define the microarchitecture of an individual\ncore pipeline. That is covered in the core pipeline document.\n\n## 1. Design Goals and Constraints\n\n### 1.1 Goals\n\n1. Provide a coherent, shared-memory programming model to all\n 128 RISC-V harts.\n2. Sustain per-core bandwidth that is high enough to keep an\n out-of-order core fed under typical server and HPC workloads.\n3. Keep average last-level cache (LLC) miss latency within a range\n that is acceptable for a tiled processor in the same class as\n published academic and industrial designs.\n4. Support a Linux-capable memory management model, including virtual\n memory, page tables, SFENCE.VMA, and standard I/O MMIO patterns.\n5. Allow verification of coherence, ordering, and interrupts with\n reasonable engineering effort.\n\n### 1.2 Constraints\n\n- The RISC-V Privileged specification, ISA Modules for Sv32, Sv39,\n Sv48, and Sv57, and the RISC-V IOMMU specification, where used,\n are normative.\n- The RISC-V Weak Memory Model (RVWMO) is normative for ordering\n between hart loads and stores.\n- The coherence protocol must be implementable on a tiled\n 128-core die in a reasonable process geometry without\n requiring exotic packaging.\n- Physical pin bandwidth to DRAM is finite. XH-1 must be\n budgeted against a realistic number of PHY lanes and channels.\n Specific numbers are marked as PROPOSAL.\n\n### 1.3 Non-Goals\n\n- This document does not specify the instruction set extensions used\n by XH-1 cores.\n- It does not specify the precise coherence protocol state machine;\n it only describes the protocol class.\n- It does not specify DRAM timing parameters; those are tied to a\n specific JEDEC standard selected at implementation time.\n\n## 2. RISC-V Architectural Requirements\n\n### 2.1 Address Space\n\nThe RISC-V Privileged Architecture (Volume II, Version 20211203) of\nthe RISC-V ISA defines:\n\n- A canonical physical address space selected by `satp.MODE`.\n Implementations of `MODE` 0 (Bare), 1 (Sv32), 8 (Sv39), 9 (Sv48),\n and 10 (Sv57) are mandatory under different profile requirements\n per the RISC-V Profiles document.\n- A `satp` CSR that holds the MODE field, the Address Space ID\n (ASID), and the root page table PPN. Writes to `satp` may become\n visible to subsequent instruction fetches and loads in a defined\n order, but require `SFENCE.VMA` for the new translation to be\n observed by other harts in a coherent manner.\n- The RISC-V IOMMU specification (riscv-iommu, draft as of the\n knowledge cutoff) defines DMA translation and the\n device-context / process-context data structures referenced by\n I/O devices.\n\nPROPOSAL: XH-1 implements Sv39 as the base virtual memory mode and\nSv48 as an optional mode. Sv57 is not required for the initial\nsilicon. The `M-mode` bare mapping remains available for firmware\nand boot loaders.\n\nASSUMPTION: Software for XH-1 will primarily target Linux. Linux\nfor RISC-V supports Sv39 and Sv48 on 64-bit platforms. Implementing\nboth is consistent with the current mainline kernel.\n\n### 2.2 Memory Ordering\n\nThe RISC-V Weak Memory Model (RVWMO) is defined in the RISC-V\nUnprivileged ISA specification, Volume I, and the accompanying\nRISC-V Memory Consistency Model document. RVWMO provides:\n\n- Load and store atomicity at aligned word and smaller granularity.\n- `FENCE`, `FENCE.I`, and `FENCE.VMA` instructions for ordering.\n- Acquire/release semantics through the `AQ` and `RL` bits on\n atomic instructions.\n- The I/O extension ordering rules described in Volume I.\n\nThe memory system must therefore guarantee that:\n\n- Hart-local loads and stores follow RVWMO rules.\n- Cross-hart ordering is enforced through cache coherence plus\n `FENCE` and `FENCE.I`.\n- `SFENCE.VMA` is correctly ordered with respect to subsequent\n instruction fetches and loads that depend on the new translation.\n\n### 2.3 Physical Memory Attribution (PMA)\n\nPMA regions are defined per-hart and describe the address ranges\nthat support main memory, I/O, idempotent reads, non-idempotent\nwrites, and so on. PMA is a static property of the platform and is\nnot affected by `satp`.\n\nPROPOSAL: PMA is encoded in a memory-mapped configuration table\nsampled by each core at reset. The exact format is defined in the\nplatform specification document.\n\n### 2.4 Physical Memory Protection (PMP)\n\nThe PMP is part of the RISC-V Privileged specification and supports\na configurable number of regions per hart. PMP checks occur in\nM-mode and may be used to isolate firmware and I/O regions from\nS-mode and U-mode.\n\nPROPOSAL: Each hart implements the maximum PMP entry count of 64\nwhen running in M-mode. S-mode translation uses the PMP only for\nthe M-mode firmware range. U-mode never sees the PMP.\n\nASSUMPTION: An M-mode firmware (such as OpenSBI) is in use and\nrelies on PMP to isolate its own memory from the supervisor.\n\n### 2.5 Cache Management Operations\n\nThe RISC-V Cache Management Operations extension (Zicbom, Zicboz,\nZicbop) is defined in Volume I of the RISC-V Unprivileged\nspecification. Operations such as `CBO.INVAL`, `CBO.FLUSH`, and\n`CBO.CLEAN` are used by software to manage coherence with external\nagents and DMA.\n\nPROPOSAL: XH-1 implements Zicbom and Zicboz. Zicbop is not required\nfor the first silicon and is marked as an open question. (See\nSection 12.)\n\n## 3. Physical Memory Map\n\nPROPOSAL: XH-1 exposes a single 48-bit physical address space to\nsoftware. The map is divided as follows:\n\n| Region | Start | End | Purpose |\n|-------------------|---------------|---------------|----------------------------------|\n| DRAM | 0x0000_0000_0000 | 0x007F_FFFF_FFFF | Main memory, up to 32 GiB. |\n| Reserved | 0x0080_0000_0000 | 0x00FF_FFFF_FFFF | Hole for future expansion. |\n| High DRAM | 0x0100_0000_0000 | 0x017F_FFFF_FFFF | Optional second DRAM range. |\n| PCIe MMIO | 0x4000_0000_0000 | 0x40FF_FFFF_FFFF | PCIe ECAM. |\n| PCIe prefetchable | 0x6000_0000_0000 | 0x60FF_FFFF_FFFF | PCIe MMIO64 prefetchable. |\n| On-chip I/O | 0x7000_0000_0000 | 0x7000_0FFF_FFFF | Interrupt controller, IOMMU, UART, boot ROM. |\n| Reserved | 0x8000_0000_0000 | 0xFFFF_FFFF_FFFF | Hole for future use. |\n\nThe exact upper bound depends on the DRAM capacity and number of\nDIMMs supported. The values above are PROPOSAL values used to\nillustrate the layout.\n\nASSUMPTION: DRAM density is between 16 GiB and 1 TiB. The lower\nbound reflects realistic server configurations for a 128-core chip\nin 2026; the upper bound is set by the 48-bit physical address\nlimit and by the cost of the memory controller area.\n\n## 4. Tile Organization\n\nPROPOSAL: XH-1 is organized as 16 tiles, each containing 8 cores.\nEach tile includes:\n\n- 8 application cores, each with a private L1 instruction cache\n and a private L1 data cache.\n- A shared L2 cache, sliced to support one bank per core.\n- A directory controller for the L3 coherence protocol.\n- A shared floating-point and vector unit cluster (PROPOSAL: see\n the vector and floating-point research document for detail).\n- A tile-local crossbar that connects cores, the L2, the\n directory controller, and the tile's connection to the global\n interconnect.\n\nThe 16 tiles connect through a global on-die interconnect. The\ninterconnect design is described in Section 8.\n\nASSUMPTION: A tile is sized to fit in a 3 mm by 3 mm die region in\na 5 nm class process. The exact floorplan is a downstream\nimplementation concern.\n\n## 5. L1 Caches\n\n### 5.1 Sizing\n\nPROPOSAL: The L1 instruction cache is 32 KiB, 8-way set-associative,\nwith 64 B lines. The L1 data cache is 32 KiB, 8-way set-associative,\nwith 64 B lines. Both caches are virtually indexed, physically\ntagged (VIPT), which constrains the index to a number of bits that\nis at most equal to the page offset.\n\nFor a 32 KiB, 8-way cache with 64 B lines, the index is 6 bits, the\noffset is 6 bits, and the tag is 52 bits. With 4 KiB pages, the\npage offset is 12 bits, so the index is fully contained in the page\noffset and VIPT is alias-free.\n\nASSUMPTION: 4 KiB pages are the common case. Larger pages such as\n2 MiB and 1 GiB reduce aliasing pressure even further, so VIPT\nremains alias-free for the supported page sizes.\n\n### 5.2 Replacement and Prefetch\n\nPROPOSAL: L1 uses a pseudo-LRU replacement policy. Hardware\nprefetchers include:\n\n- A next-line prefetcher on the instruction side.\n- A stride prefetcher on the data side.\n- A stream prefetcher on the data side for sequential access\n patterns.\n\nThe prefetch degree and aggressiveness are tunable through\nmachine-mode configuration registers.\n\n### 5.3 MSHRs and Write Combining\n\nPROPOSAL: Each L1 data cache has 16 miss status handling\nregisters (MSHRs), supporting up to 16 outstanding misses. The\nwrite buffer holds up to 16 entries. Combining of partial line\nwrites is performed in the write buffer.\n\nASSUMPTION: 16 MSHRs is sufficient for the workloads of interest.\nA sensitivity study is required before the L1 MSHR count is\nfinalized. (See Section 12.)\n\n## 6. L2 Cache\n\nPROPOSAL: Each tile contains an 8 MiB L2 cache, organized as 8\nbanks, one per core. Each bank is 1 MiB, 16-way set-associative,\nwith 64 B lines. The L2 is inclusive of the L1 and acts as the\ncoherence point for the L1 caches on the tile.\n\nTotal L2 capacity across the die is 128 MiB. This number is a\nPROPOSAL based on the die area budget for caches in the target\nprocess node; it is not yet finalized.\n\n### 6.1 Inclusion\n\nPROPOSAL: The L2 is strictly inclusive of the L1 caches on the\nsame tile. A coherence invalidation from the L3 directory causes\nthe L2 to invalidate the corresponding L1 line.\n\nASSUMPTION: Strict inclusion simplifies the coherence protocol at\nthe cost of additional L2 traffic for victims. A non-inclusive\nL2 is an open question. (See Section 12.)\n\n### 6.2 Latency\n\nASSUMPTION: The L2 hit latency is between 12 and 20 cycles. The\nexact number depends on the physical layout of the cache and the\ncore. This is an open implementation question.\n\n## 7. L3 Cache and Coherence Protocol\n\n### 7.1 Class\n\nPROPOSAL: XH-1 uses a directory-based MESI coherence protocol at\nthe L3 level. The directory is distributed across tiles, with one\ndirectory controller per tile. Each line in the L3 is owned by\nexactly one directory, identified by the home tile.\n\nThe protocol class is \"distributed directory MESI with distributed\nshared L3\". This class is the same as the protocol class used in\nseveral published academic and industrial designs, including the\nMIT Alewife, the SGI Origin, the Oracle/Berkeley SPARC M7, and the\nFujitsu A64FX.\n\n### 7.2 States\n\nThe four stable MESI states are:\n\n- **M**odified: the line is dirty and held by a single cache.\n- **E**xclusive: the line is clean and held by a single cache.\n- **S**hared: the line is clean and may be held by multiple\n caches.\n- **I**nvalid: the line is not present.\n\nPROPOSAL: A fifth \"Owned\" state is not used in the base protocol\nto keep the directory and transient state space simple. A\ndirectory entry tracks the presence bits for each tile.\n\n### 7.3 Directory Encoding\n\nPROPOSAL: The directory uses a coarse vector of presence bits,\none per tile, plus a \"has-modified\" bit. For 16 tiles this is 17\nbits per line. The directory is stored in a dedicated SRAM\nstructure, sized at approximately 12.5 MiB for a 64 MiB L3 with\n4x replication of the directory across the on-die fabric to\nsupport transient races.\n\nASSUMPTION: The 4x replication of the directory entry is a\nsimplified model. The exact replication strategy is a downstream\nimplementation question. (See Section 12.)\n\n### 7.4 Directory Operations\n\nPROPOSAL: Coherence operations follow a request-response pattern:\n\n1. A core issues a read or write to its L1.\n2. On L1 miss, the request is forwarded to the L2.\n3. On L2 miss, the L2 sends a coherence request to the home tile's\n directory.\n4. The directory invalidates any other copies, forwards the data\n to the requester, and updates the presence vector.\n5. The directory replies to the requester, which completes the\n L1 fill.\n\nWritebacks from L1 invalidate the corresponding L2 line and\nupdate the directory.\n\n### 7.5 Memory Consistency Ordering\n\nThe protocol preserves the appearance of a single global ordering\nof coherence transactions as seen by the cores. RISC-V `FENCE`\nordering is enforced by the L2 controller before forwarding a\nrequest to the directory.\n\nPROPOSAL: Store atomic instructions (`AMO*` and `LR/SC`) are\nhandled by the L2 controller. The L2 acts as the ordering point\nfor atomic operations and serializes the operation with respect\nto other requests to the same address.\n\n### 7.6 Forwarding and Snoop Filtering\n\nPROPOSAL: The directory acts as a snoop filter. A coherence\nrequest that finds the directory in the \"S\" state with no other\nsharer can supply the data from the L2 without further snoop\ntraffic.\n\n## 8. On-Die Interconnect\n\n### 8.1 Topology\n\nPROPOSAL: The on-die interconnect is a 2D mesh of 4 by 4 tiles.\nEach tile has a router that connects to its four nearest neighbors.\nThe mesh supports a 512-bit link per direction, clocked at a\nfraction of the core clock.\n\nASSUMPTION: A 2D mesh is sufficient for the bandwidth requirements\nof a 128-core die. A 2D torus is an alternative discussed in\nSection 12.\n\n### 8.2 Routing\n\nPROPOSAL: The mesh uses deterministic X-Y routing. Adaptive routing\nis not used in the first silicon to simplify verification.\n\n### 8.3 Link Bandwidth\n\nPROPOSAL: Each link provides 64 B per cycle of useful payload at\nthe router clock. With 16 tiles and 4 directions per router\n(excluding the boundary cases), the bisection bandwidth of the\nmesh is 2 links by 4 tiles by 64 B per cycle, or 512 B per cycle.\nAt a 1.5 GHz router clock this is 768 GB/s of bisection\nbandwidth.\n\nASSUMPTION: The 1.5 GHz router clock is feasible in the target\nprocess node. This is an open implementation question.\n\n### 8.4 Flow Control\n\nPROPOSAL: The interconnect uses credit-based flow control on each\nlink. Each virtual channel maintains a count of available credits\nand stalls a flit only when the downstream router is unable to\naccept it.\n\n## 9. DRAM Controller\n\n### 9.1 Topology\n\nPROPOSAL: XH-1 has 8 DRAM controllers, distributed on the die. Two\ncontrollers are placed near the center of each edge. The\ncontrollers connect to the mesh with high-bandwidth links and to\nexternal PHYs through the package substrate.\n\n### 9.2 Channels and DIMMs\n\nPROPOSAL: Each controller supports 2 channels, for a total of 16\nchannels. Each channel supports a single DDR5 DIMM in the first\nsilicon.\n\n### 9.3 Bandwidth\n\nASSUMPTION: At DDR5-6400, a 64-bit channel provides 51.2 GB/s\npeak. With 16 channels, the peak DRAM bandwidth is 819 GB/s.\n\nPROPOSAL: The memory system is designed to sustain an average\nL3 miss rate of 0.05 misses per instruction across the 128 cores\nat the rated clock. This is consistent with a 1.0 GHz core clock\nand an IPC of 1.0, giving 128 GIPS, and a DRAM traffic of 0.05\nmisses by 128 GIPS by 64 B per miss, or 410 GB/s. This is below\nthe 819 GB/s peak, leaving headroom for OS and I/O traffic.\n\n### 9.4 Quality of Service\n\nPROPOSAL: The DRAM controller supports 4 QoS classes. A\nconfiguration register in the platform specification selects the\nclass for each request source (core, device, I/O).\n\nASSUMPTION: A weighted round-robin scheduler per channel is\nsufficient for the first silicon. A more sophisticated scheduler\nis an open question.\n\n## 10. Operating System and Firmware\n\n### 10.1 Linux Memory Model\n\nPROPOSAL: XH-1 supports mainline Linux for RISC-V. The relevant\nkernel features are:\n\n- Sv39 and Sv48 virtual memory.\n- NUMA-aware scheduling, with one NUMA node per tile.\n- Standard cache management operations through Zicbom.\n- PCIe MMU through the optional RISC-V IOMMU.\n\nASSUMPTION: Linux for RISC-V provides all of the above features\non a 64-bit SMP platform. This assumption should be re-validated\nagainst the kernel version targeted for the first silicon.\n\n### 10.2 NUMA Effects\n\nPROPOSAL: Each tile is exposed as a NUMA node to Linux. The\noperating system scheduler places a task on the node closest to\nits memory. The boot firmware sets up the device tree and ACPI\ntable accordingly.\n\n### 10.3 Firmware and PMP\n\nPROPOSAL: OpenSBI is the M-mode firmware. OpenSBI uses PMP to\nisolate its own memory from S-mode. The SBI interface is the\nprimary interface between S-mode and M-mode.\n\n### 10.4 IOMMU and DMA\n\nPROPOSAL: A RISC-V IOMMU is included on the die. The IOMMU\ntranslates device DMA addresses using the same page table\nformats as the CPU. PCIe devices that support ATS bypass the\nIOMMU on a per-transaction basis.\n\nASSUMPTION: The RISC-V IOMMU specification is sufficiently\nmature at the time of tape-out. If not, the IOMMU is replaced\nby a software bounce-buffering scheme in the first silicon.\n\n## 11. Interrupts and Memory System Interactions\n\n### 11.1 Interrupt Delivery\n\nPROPOSAL: A platform-level interrupt controller (PLIC) is mapped\nat the on-chip I/O region. The PLIC supports 256 interrupt\nsources and 128 interrupt targets, one per hart. The PLIC\naccesses are strongly ordered with respect to the coherence\nprotocol.\n\n### 11.2 Interrupt and Coherence\n\nASSUMPTION: The interrupt controller is on the coherence domain.\nAn interrupt write from a device is visible to the PLIC\nregisters, and the subsequent PLIC read by the core is ordered\nthrough the coherence protocol.\n\n### 11.3 Interrupt Latency\n\nASSUMPTION: Worst-case interrupt latency through the PLIC is\nunder 1 microsecond at the rated clock, including coherence\nround-trip and pipeline drain. The actual value is workload\ndependent and is left for the performance document.\n\n## 12. Open Questions and Tradeoffs\n\nThe following questions are unresolved at the time of writing.\nEach is a candidate for further study before the architecture is\nfrozen.\n\n1. **Inclusive vs. non-inclusive L2.** Strict inclusion\n simplifies coherence but may waste L2 capacity. A\n non-inclusive victim cache design is an alternative.\n2. **L3 size.** A 64 MiB L3 is the proposal. A 32 MiB or 128 MiB\n L3 are alternatives; each changes the area budget and miss\n rate.\n3. **Directory organization.** A coarse vector with replication\n is the proposal. A hierarchical or sparse directory is an\n alternative that may reduce area at the cost of latency.\n4. **Mesh vs. torus.** A 2D mesh is the proposal. A 2D torus\n doubles the bisection bandwidth and may be required for very\n high cache miss rates. A ring or concentrated mesh is a\n third alternative.\n5. **Number of memory channels.** 16 channels is the proposal.\n 8 or 24 channels are alternatives.\n6. **Zicbop implementation.** Cache block prefetch operations are\n optional in the RISC-V Cache Management Operations\n specification. Their cost and benefit on XH-1 are not yet\n quantified.\n7. **L1 MSHR count.** 16 MSHRs per L1 is the proposal. The\n workload-dependent optimum is not known.\n8. **IOMMU maturity.** The RISC-V IOMMU specification was a\n draft at the knowledge cutoff date. The XH-1 implementation\n may need to wait for a stable revision.\n9. **Coherence at the L1 level.** The current proposal\n collapses L1 coherence to the L2. An alternative is a\n broadcast snoop between L1s on the same tile. This trades\n L1 bandwidth for L2 traffic.\n10. **SMT support.** Each core may support a small number of\n hardware threads. SMT changes the coherence and memory\n system pressure per core. The number of threads per core\n is not yet decided.\n11. **Vector and matrix unit interactions.** The vector unit\n performs high-bandwidth loads and stores. The memory\n system must be able to support sustained vector load\n bandwidth at the rated clock. The vector unit design\n appears in a separate document.\n12. **Power management.** Clock gating, power gating, and DVFS\n interact with the memory system. The current document does\n not yet address these interactions.\n\n## 13. Scalability Analysis for 128 Cores\n\n### 13.1 Coherence Scalability\n\nA broadcast snoop protocol scales as O(N^2) in the worst case,\nwhere N is the number of cores. For 128 cores this is 16,384\nworst-case broadcast pairs, which is generally not viable on a\nsingle die. The directory protocol scales as O(N) in the worst\ncase for invalidations, and the directory size grows linearly\nwith the number of tiles.\n\nPROPOSAL: A directory protocol is required at 128 cores. A\nsnoopy protocol is not viable.\n\n### 13.2 Interconnect Scalability\n\nA 2D mesh of 16 tiles has bisection bandwidth proportional to\nthe number of tiles per row. The 4 by 4 mesh provides 4 links of\nbisection, which is sufficient for the projected 410 GB/s of\naverage DRAM traffic in Section 9.3.\n\nASSUMPTION: Average traffic does not exceed 70 percent of the\npeak bisection. A sensitivity analysis is required. (See\nSection 12.)\n\n### 13.3 Directory Storage Scalability\n\nThe directory entry is 17 bits per L3 line. For a 64 MiB L3 with\n64 B lines, this is 2^20 lines by 17 bits, or 2.2 MiB per\ndirectory. With 4x replication, the on-die directory storage is\napproximately 12.5 MiB. This is a small fraction of the L3\nstorage.\n\n### 13.4 Operating System Scalability\n\nLinux scales to hundreds of cores. The memory system must\nexpose the topology through ACPI and the device tree so that the\nkernel NUMA scheduler can make informed decisions. The\n128-core XH-1 is below the upper bound of common server Linux\nconfigurations as of the knowledge cutoff.\n\n### 13.5 Verification Scalability\n\nCoherence verification grows in complexity with the number of\ncores. PROPOSAL: a combination of formal verification of the\ndirectory protocol state machine, simulation-based testing, and\nrandom instruction generation is used. The verification\nstrategy is described in the verification document.\n\n## 14. Verification\n\n### 14.1 Cache Coherence\n\nPROPOSAL: The directory protocol state machine is verified using\na model checker against the MESI invariants. The invariants\ninclude:\n\n- Single writer per line.\n- No read of a line in the Invalid state.\n- Correct invalidation on a write to a shared line.\n\n### 14.2 Memory Consistency\n\nPROPOSAL: A RISC-V memory consistency model litmus test suite\nis run against the system, using both RTL simulation and a\nsoftware reference model.\n\n### 14.3 Performance and Stress Testing\n\nPROPOSAL: Synthetic traffic patterns, including uniform random,\nhotspot, and adversarial patterns, are used to verify that the\nmemory system behaves correctly under stress. Realistic\nworkloads, including SPEC CPU 2017, PARSEC, and selected\nGROMACS and LAMMPS runs, are used for performance validation.\n\nASSUMPTION: The listed benchmarks are representative of the\ntarget workloads. This assumption should be re-validated by the\nperformance team.\n\n## 15. Interaction With Other Subsystems\n\n### 15.1 Pipeline\n\nThe L1 cache is in the core pipeline. The L1 hit latency is\nvisible in the branch resolution and the load-to-use latency.\nPROPOSAL: The L1 hit latency is 4 cycles for a load-to-use on a\nhit. A miss to the L2 adds between 12 and 20 cycles, depending\non the L2 hit and the core-to-L2 routing.\n\n### 15.2 Cache Hierarchy\n\nThe L1, L2, and L3 form a hierarchy. The L2 is inclusive of the\nL1. The L3 is the coherence point. The L3 is non-inclusive of\nthe L2 by default.\n\n### 15.3 Memory System\n\nThe DRAM controllers are described in Section 9. The on-die\ninterconnect carries both coherence and data traffic.\n\n### 15.4 Interconnect\n\nThe on-die interconnect is described in Section 8. Coherence\nand data traffic share the same physical links.\n\nPROPOSAL: Virtual channels separate the request, response, and\ndata traffic. The number of virtual channels per link is 4.\n\n### 15.5 Coherence\n\nCoherence is described in Section 7. The protocol class is\ndistributed directory MESI with distributed shared L3.\n\n### 15.6 Interrupts\n\nInterrupts are described in Section 11. The PLIC is on the\ncoherence domain.\n\n### 15.7 Operating System\n\nThe operating system interactions are described in Section 10.\nThe OS is Linux for RISC-V. The firmware is OpenSBI.\n\n### 15.8 Performance\n\nThe performance analysis is described in Section 13 and is\ndetailed in a separate performance document. Key memory-system\nperformance metrics include:\n\n- L1 hit rate\n- L2 hit rate\n- L3 hit rate\n- DRAM bandwidth utilization\n- Average and tail DRAM latency\n- Interconnect average and tail latency\n\n## 16. Recommendations\n\n1. Implement a directory-based MESI coherence protocol with a\n distributed shared L3. Snoopy protocols are not viable at\n 128 cores.\n2. Use a 2D mesh on-die interconnect. The 2D mesh is sufficient\n for the projected traffic and is simpler to verify than a\n torus.\n3. Implement Zicbom and Zicboz cache management operations.\n These are required for Linux on RISC-V with mainline\n kernels.\n4. Expose each tile as a NUMA node to the operating system. This\n allows the kernel scheduler to make topology-aware decisions.\n5. Defer the final decision on the L3 size, the number of DRAM\n channels, and the directory organization to the\n implementation phase, after the sensitivity studies listed\n in Section 12 are completed.\n6. Defer the implementation of Zicbop until a workload study\n shows a clear benefit.\n\n## 17. References\n\nThe following standards and references are cited in this\ndocument:\n\n- RISC-V Unprivileged ISA Specification, Volume I, Document\n Version 20191213. Section on the Weak Memory Model and the\n Cache Management Operations extensions.\n- RISC-V Privileged ISA Specification, Volume II, Document\n Version 20211203. Sections on `satp`, SFENCE.VMA, PMP, and\n interrupt handling.\n- RISC-V Profiles, Document Version 20230914. Application and\n platform profile requirements.\n- RISC-V IOMMU Specification, draft as of the knowledge cutoff.\n- JEDEC DDR5 Standard, JESD79-5.\n- JEDEC LPDDR5 Standard, JESD209-5A.\n- Hennessy and Patterson, \"Computer Architecture: A Quantitative\n Approach\", 6th edition. Sections on cache coherence and\n directory protocols.\n- Culler, Singh, and Gupta, \"Parallel Computer Architecture: A\n Hardware/Software Approach\". Chapter on scalable cache\n coherence.\n- M. M. K. Martin, M. D. Hill, and D. A. Wood, \"Token\n Coherence\", University of Wisconsin Computer Sciences\n Technical Report 1621, 2003. Discusses a directory-based\n protocol with explicit invalidation tokens.\n- P. Kongetira, K. Aingaran, and K. Olukotun, \"Niagara: A\n 32-Way Multithreaded SPARC Processor\", IEEE Micro, March 2005.\n Discusses a directory-coherent CMP design.\n- R. Azimi et al., \"Integration Challenges and Tradeoffs for\n Tiled Many-Core Processors\", Proceedings of the 22nd\n International Conference on Computer Design, 2004.\n- P. Hammarlund et al., \"Haswell: The Fourth-Generation Intel\n Core Processor\", IEEE Micro, March 2014. Discusses a\n ring-based interconnect in a tiled CMP.\n- J. L. Henning, \"SPEC CPU2006 Benchmark Descriptions\", ACM\n SIGARCH Computer Architecture News, September 2006.\n- OpenSBI documentation, https://github.com/riscv-software-src/opensbi.\n\n## 18. Document History\n\n- Draft 0.1: initial outline.\n- Draft 0.2: populated physical memory map, coherence protocol\n class, and L1/L2/L3 sizing as proposals.\n- Draft 0.3: added open questions and recommendations.\n```","refusal":null,"reasoning":null}}],"usage":{"prompt_tokens":663,"completion_tokens":7174,"total_tokens":7837,"cost":0,"is_byok":false,"prompt_tokens_details":{"cached_tokens":135,"cache_write_tokens":0,"audio_tokens":0,"video_tokens":0},"cost_details":{"upstream_inference_cost":0,"upstream_inference_prompt_cost":0,"upstream_inference_completions_cost":0},"completion_tokens_details":{"reasoning_tokens":0,"image_tokens":0,"audio_tokens":0}}} \ No newline at end of file diff --git a/research/.xh1/responses/20260826T125740Z-20260826T125859Z-26838.json b/research/.xh1/responses/20260826T125740Z-20260826T125859Z-26838.json new file mode 100644 index 0000000..48e2621 --- /dev/null +++ b/research/.xh1/responses/20260826T125740Z-20260826T125859Z-26838.json @@ -0,0 +1 @@ +{"id":"chatcmpl-327","object":"chat.completion","created":1787749327,"model":"gemma4:12b","system_fingerprint":"fp_ollama","choices":[{"index":0,"message":{"role":"assistant","content":"","reasoning":"The user wants a technical review of a memory system specification for a 128-core RISC-V processor (XH-1).\n\n * *13.5 Verification Scalability:* Mentions directory protocol state machine verification, simulation, and random instruction generation.\n * *14.1 Cache Coherence:* Mentions MESI invariants (Single writer, No read of Invalid, Correct invalidation).\n * *14.2 Memory Consistency:* Mentions RISC-V litmus tests.\n * *14.3 Performance/Stress:* Mentions SPEC CPU 2017, PARSEC, GROMACS, LAMMPS.\n * *15.1 Pipeline:* L1 hit 4 cycles, L2 miss 12-20 cycles.\n * *15.2 Cache Hierarchy:* L2 inclusive of L1, L3 non-inclusive of L2 (common for large systems).\n * *15.4 Interconnect:* 2D mesh, 4 virtual channels.\n * *15.5 Coherence:* Distributed directory MESI with shared L3.\n * *15.6 Interrupts:* PLIC on coherence domain.\n * *15.7 OS:* Linux, OpenSBI.\n * *16. Recommendations:*\n 1. Directory-based MESI (correct for 128 cores).\n 2. 2D Mesh (standard for 128 cores).\n 3. Zicbom/Zicboz (essential for RISC-V Linux).\n 4. NUMA nodes (essential for large scale).\n 5. Deferred L3/DRAM decisions (standard for early design).\n 6. Zicbop deferred (reasonable).\n\n * *128-core implications:* Directory-based is mandatory. Snooping is impossible. 2D Mesh is the standard topology.\n * *RISC-V compliance:* Zicbom/Zicboz are correctly identified as necessary for Linux.\n * *Memory Consistency:* Litmus tests are the standard way to verify the RISC-V Weak Memory Model (RVWMO).\n * *Cache Hierarchy:* L2 inclusive of L1 is standard. L3 non-inclusive of L2 is common in large systems to maximize L3 capacity.\n * *Interconnect:* 4 virtual channels is a standard number to prevent head-of-line blocking and separate request/response/data.\n * *NUMA:* Mapping tiles to NUMA nodes is critical for 128 cores to allow the OS to manage locality.\n\n * *15.1 Pipeline:* \"L1 hit latency is 4 cycles for a load-to-use on a hit.\" This is a bit low for some high-performance designs but acceptable for a high-frequency core.\n * *15.4 Interconnect:* \"Virtual channels separate the request, response, and data traffic.\" This is good.\n * *16.3 Zicbom/Zicboz:* These are crucial for cache management in Linux.\n\n * *Technical correctness:* High.\n * *RISC-V compliance:* High.\n * *Microarchitectural correctness:* High.\n * *Memory-model correctness:* High.\n * *Cache/coherence correctness:* High.\n * *Interconnect/scalability:* High.\n * *128-core implications:* Correctly identified.\n * *Performance:* Standard benchmarks included.\n * *Verification:* Standard methodology.\n\n * The document is a high-level specification/proposal.\n * It correctly identifies the scaling issues of snoopy protocols and correctly proposes a directory-based approach.\n * It correctly identifies the need for NUMA awareness.\n * It correctly identifies the RISC-V extensions needed for Linux.\n\n * *Minor point:* The \"L3 is non-inclusive of the L2\" is a design choice, but it's a good one for large systems to maximize effective capacity.\n * *Minor point:* \"128-core XH-1 is below the upper bound of common server Linux configurations\" - True, many systems go higher, but 128 is a significant scale.\n\n * The document is technically sound for a research/pre-design phase.\n\n * *Wait, let me double check the \"L1 hit latency is 4 cycles for a load-to-use on a hit\".* In many high-performance RISC-V cores (like SiFive's designs), 3-4 cycles is standard for L1 hit.\n * *Wait, let me double check \"Zicbom and Zicboz\".* Yes, these are the standard extensions for cache management.\n\n * *One potential issue:* The document mentions \"The L3 is the coherence point.\" In a distributed directory system, the \"coherence point\" is technically the directory entry (which might reside in the L3 or a separate directory structure). However, in many designs, the L3 *is* the directory. This is acceptable phrasing.\n\n * *One potential issue:* \"The L3 is non-inclusive of the L2 by default.\" This is a standard way to handle large caches.\n\n * *Conclusion:* The document is solid.\n\n * *Check:* \"The L1 hit latency is 4 cycles for a load-to-use on a hit.\"\n * *Check:* \"A miss to the L2 adds between 12 and 20 cycles.\"\n * *Check:* \"The number of virtual channels per link is 4.\" (Standard for 3 types of traffic: Req, Resp, Data, plus potentially a control/management channel or just 3 types with one spare).\n\n * *Refining the \"Coherence Point\" wording:* In a distributed directory, the \"coherence point\" is the directory. If the L3 is the coherence point, it implies the directory is integrated into the L3. This is a standard design.\n\n * *Refining the \"NUMA\" point:* For 128 cores, mapping each tile or a cluster of tiles to a NUMA node is essential for Linux.\n\n * *Refining \"Zicbop\":* Zicbop is for \"Block Operations\" (like block-based cache management). It's less critical than Zicbom/Zicboz for basic Linux operation, so deferring it is a reasonable engineering decision.\n\n * *Conclusion:* No major issues found.\n\n * *Wait, one minor thing:* \"The L3 is non-inclusive of the L2 by default.\" In some architectures, if L3 is non-inclusive, it might be \"exclusive\" or \"mostly inclusive\". \"Non-inclusive\" is a safe term.\n\n * *Wait, \"The L1 hit latency is 4 cycles for a load-to-use on a hit.\"* This is a bit specific, but it's a \"Proposal\" in the text.\n\n * *Verification:* \"The directory protocol state machine is verified using a model checker against the MESI invariants.\" This is the gold standard for directory protocols.\n\n * *Summary:* The document is technically sound.\n\n * *Final check on RISC-V specifics:*\n * RISC-V Memory Model: RVWMO.\n * Zicbom/Zicboz: Required for Linux.\n * NUMA: Required for large scale.\n * Directory: Required for 128 cores.\n\n * *Conclusion:* PASS.\n\n * *Wait, one small detail:* \"The L3 is the coherence point.\" In a distributed directory, the \"coherence point\" is the directory. If the L3 *contains* the directory, it is the coherence point. This is standard terminology in many papers (e.g., those by Hennessy/Patterson).\n\n * *Wait, \"The L1 hit latency is 4 cycles for a load-to-use on a hit.\"* This is a bit low for a complex pipeline, but for a high-performance core, it's a standard target.\n\n * *Wait, \"The number of virtual channels per link is 4.\"* This is sufficient for Request, Response, and Data, with one extra for something else (like probe or management).\n\n * *Conclusion:* No critical errors.\n\n * *One minor check:* \"The L1 hit latency is 4 cycles for a load-to-use on a hit. A miss to the L2 adds between 12 and 20 cycles, depending on the L2 hit and the core-to-L2 routing.\"\n * If L1 hit is 4, and L2 hit is 12-20 *additional* cycles, the total is 16-24. This is a reasonable range for a high-performance system.\n\n * *Final check on \"Zicbop\":* Zicbop is for \"Block Operations\" (e.g., `cbo.block`). It's not strictly required for basic Linux operation, so deferring it is a valid engineering choice.\n\n * *Final"},"finish_reason":"length"}],"usage":{"prompt_tokens":2051,"completion_tokens":2045,"total_tokens":4096}} diff --git a/research/.xh1/responses/20260826T125740Z-20260826T130209Z-15370.json b/research/.xh1/responses/20260826T125740Z-20260826T130209Z-15370.json new file mode 100644 index 0000000..5012fba --- /dev/null +++ b/research/.xh1/responses/20260826T125740Z-20260826T130209Z-15370.json @@ -0,0 +1 @@ +{"id":"chatcmpl-921","object":"chat.completion","created":1787749374,"model":"gemma4:12b","system_fingerprint":"fp_ollama","choices":[{"index":0,"message":{"role":"assistant","content":"The document describes a high-performance, many-core RISC-V system (XH-1) with 128 cores. The technical specifications provided are consistent with modern high-performance computing (HPC) architecture and RISC-V standards.\n\n### Technical Review:\n\n1. **Scalability \u0026 Interconnect:** The choice of a **2D mesh** for 128 cores is a standard and correct architectural decision. At this scale, a ring or crossbar is physically and logically impractical. The use of **virtual channels** to separate request, response, and data traffic is critical to prevent head-of-line blocking and protocol deadlocks in a directory-based system.\n2. **Cache Coherence:** The transition from snooping to a **distributed directory MESI protocol** is mandatory for 128 cores. The inclusion of **Zicbom and Zicboz** is a critical requirement for RISC-V Linux compatibility, as these allow the OS to manage cache lines and memory attributes correctly.\n3. **Memory Consistency:** The proposal to use a **litmus test suite** for RISC-V memory consistency is the industry standard for verifying that the hardware correctly implements the Weak Memory Model (WMM) required by the RISC-V ISA.\n4. **Hierarchy:** The L2 being inclusive of L1 and the L3 being the coherence point is a standard design that simplifies the coherence logic at the L3 level.\n5. **OS/Software:** Exposing the topology as **NUMA nodes** is the correct way to handle a 128-core system in Linux, ensuring the scheduler can optimize for local memory access.\n6. **Verification:** The multi-pronged approach (Formal for state machines, Simulation for consistency, and Stress/Synthetic for performance) is a robust verification strategy for a complex SoC.\n\n### Minor Observations:\n- The decision to defer Zicbop is prudent, as it is a more complex extension for \"buffer\" management that may not be necessary for initial silicon.\n- The distinction between \"inclusive\" and \"non-inclusive\" L3 is correctly identified as a key design choice for the implementation phase.\n\nThe document is technically sound, follows RISC-V standards, and addresses the specific challenges of scaling to 128 cores.\n\n```json\n{\n \"verdict\": \"PASS\",\n \"confidence\": \"HIGH\",\n \"issues\": [],\n \"required_fixes\": []\n}\n```"},"finish_reason":"stop"}],"usage":{"prompt_tokens":2051,"completion_tokens":532,"total_tokens":2583}} diff --git a/research/.xh1/runs/20260826T125740Z/attempts b/research/.xh1/runs/20260826T125740Z/attempts new file mode 100644 index 0000000..87c37e6 --- /dev/null +++ b/research/.xh1/runs/20260826T125740Z/attempts @@ -0,0 +1,2 @@ +2026-08-26T12:58:59Z research/05-memory/memory-architecture.md 1 research success +2026-08-26T13:02:54Z research/05-memory/memory-architecture.md 1 review PASS diff --git a/research/.xh1/runs/20260826T125740Z/candidate.md b/research/.xh1/runs/20260826T125740Z/candidate.md new file mode 100644 index 0000000..15a59ba --- /dev/null +++ b/research/.xh1/runs/20260826T125740Z/candidate.md @@ -0,0 +1,740 @@ +``` +# Memory Architecture + +## Status + +This document is part of the XH-1 research series. It describes the +memory system proposal for the XH-1, a custom 128-core RISC-V +processor. Where implementation choices have not been finalized, the +text uses **PROPOSAL** to mark a design that is being considered but +has not been committed. Where the text is an **ASSUMPTION**, evidence +is not yet available and the statement must be re-validated before +downstream documents depend on it. + +## Scope + +This document covers: + +- The physical memory map and address space layout visible to software + running on XH-1 cores. +- The relationship between the RISC-V privileged architecture and the + memory system, including `satp`, SFENCE.VMA, and the PMP. +- The on-chip memory hierarchy: L1, L2, and the proposed shared L3. +- The directory-based coherence protocol between L2 slices. +- The interconnect that ties cores, caches, and DRAM controllers + together. +- The DRAM controller and the memory-side bandwidth budget. +- I/O, interrupts, and operating system interactions that constrain + memory-system design. +- Verification and performance considerations specific to a 128-core + tiled implementation. + +This document does not define the microarchitecture of an individual +core pipeline. That is covered in the core pipeline document. + +## 1. Design Goals and Constraints + +### 1.1 Goals + +1. Provide a coherent, shared-memory programming model to all + 128 RISC-V harts. +2. Sustain per-core bandwidth that is high enough to keep an + out-of-order core fed under typical server and HPC workloads. +3. Keep average last-level cache (LLC) miss latency within a range + that is acceptable for a tiled processor in the same class as + published academic and industrial designs. +4. Support a Linux-capable memory management model, including virtual + memory, page tables, SFENCE.VMA, and standard I/O MMIO patterns. +5. Allow verification of coherence, ordering, and interrupts with + reasonable engineering effort. + +### 1.2 Constraints + +- The RISC-V Privileged specification, ISA Modules for Sv32, Sv39, + Sv48, and Sv57, and the RISC-V IOMMU specification, where used, + are normative. +- The RISC-V Weak Memory Model (RVWMO) is normative for ordering + between hart loads and stores. +- The coherence protocol must be implementable on a tiled + 128-core die in a reasonable process geometry without + requiring exotic packaging. +- Physical pin bandwidth to DRAM is finite. XH-1 must be + budgeted against a realistic number of PHY lanes and channels. + Specific numbers are marked as PROPOSAL. + +### 1.3 Non-Goals + +- This document does not specify the instruction set extensions used + by XH-1 cores. +- It does not specify the precise coherence protocol state machine; + it only describes the protocol class. +- It does not specify DRAM timing parameters; those are tied to a + specific JEDEC standard selected at implementation time. + +## 2. RISC-V Architectural Requirements + +### 2.1 Address Space + +The RISC-V Privileged Architecture (Volume II, Version 20211203) of +the RISC-V ISA defines: + +- A canonical physical address space selected by `satp.MODE`. + Implementations of `MODE` 0 (Bare), 1 (Sv32), 8 (Sv39), 9 (Sv48), + and 10 (Sv57) are mandatory under different profile requirements + per the RISC-V Profiles document. +- A `satp` CSR that holds the MODE field, the Address Space ID + (ASID), and the root page table PPN. Writes to `satp` may become + visible to subsequent instruction fetches and loads in a defined + order, but require `SFENCE.VMA` for the new translation to be + observed by other harts in a coherent manner. +- The RISC-V IOMMU specification (riscv-iommu, draft as of the + knowledge cutoff) defines DMA translation and the + device-context / process-context data structures referenced by + I/O devices. + +PROPOSAL: XH-1 implements Sv39 as the base virtual memory mode and +Sv48 as an optional mode. Sv57 is not required for the initial +silicon. The `M-mode` bare mapping remains available for firmware +and boot loaders. + +ASSUMPTION: Software for XH-1 will primarily target Linux. Linux +for RISC-V supports Sv39 and Sv48 on 64-bit platforms. Implementing +both is consistent with the current mainline kernel. + +### 2.2 Memory Ordering + +The RISC-V Weak Memory Model (RVWMO) is defined in the RISC-V +Unprivileged ISA specification, Volume I, and the accompanying +RISC-V Memory Consistency Model document. RVWMO provides: + +- Load and store atomicity at aligned word and smaller granularity. +- `FENCE`, `FENCE.I`, and `FENCE.VMA` instructions for ordering. +- Acquire/release semantics through the `AQ` and `RL` bits on + atomic instructions. +- The I/O extension ordering rules described in Volume I. + +The memory system must therefore guarantee that: + +- Hart-local loads and stores follow RVWMO rules. +- Cross-hart ordering is enforced through cache coherence plus + `FENCE` and `FENCE.I`. +- `SFENCE.VMA` is correctly ordered with respect to subsequent + instruction fetches and loads that depend on the new translation. + +### 2.3 Physical Memory Attribution (PMA) + +PMA regions are defined per-hart and describe the address ranges +that support main memory, I/O, idempotent reads, non-idempotent +writes, and so on. PMA is a static property of the platform and is +not affected by `satp`. + +PROPOSAL: PMA is encoded in a memory-mapped configuration table +sampled by each core at reset. The exact format is defined in the +platform specification document. + +### 2.4 Physical Memory Protection (PMP) + +The PMP is part of the RISC-V Privileged specification and supports +a configurable number of regions per hart. PMP checks occur in +M-mode and may be used to isolate firmware and I/O regions from +S-mode and U-mode. + +PROPOSAL: Each hart implements the maximum PMP entry count of 64 +when running in M-mode. S-mode translation uses the PMP only for +the M-mode firmware range. U-mode never sees the PMP. + +ASSUMPTION: An M-mode firmware (such as OpenSBI) is in use and +relies on PMP to isolate its own memory from the supervisor. + +### 2.5 Cache Management Operations + +The RISC-V Cache Management Operations extension (Zicbom, Zicboz, +Zicbop) is defined in Volume I of the RISC-V Unprivileged +specification. Operations such as `CBO.INVAL`, `CBO.FLUSH`, and +`CBO.CLEAN` are used by software to manage coherence with external +agents and DMA. + +PROPOSAL: XH-1 implements Zicbom and Zicboz. Zicbop is not required +for the first silicon and is marked as an open question. (See +Section 12.) + +## 3. Physical Memory Map + +PROPOSAL: XH-1 exposes a single 48-bit physical address space to +software. The map is divided as follows: + +| Region | Start | End | Purpose | +|-------------------|---------------|---------------|----------------------------------| +| DRAM | 0x0000_0000_0000 | 0x007F_FFFF_FFFF | Main memory, up to 32 GiB. | +| Reserved | 0x0080_0000_0000 | 0x00FF_FFFF_FFFF | Hole for future expansion. | +| High DRAM | 0x0100_0000_0000 | 0x017F_FFFF_FFFF | Optional second DRAM range. | +| PCIe MMIO | 0x4000_0000_0000 | 0x40FF_FFFF_FFFF | PCIe ECAM. | +| PCIe prefetchable | 0x6000_0000_0000 | 0x60FF_FFFF_FFFF | PCIe MMIO64 prefetchable. | +| On-chip I/O | 0x7000_0000_0000 | 0x7000_0FFF_FFFF | Interrupt controller, IOMMU, UART, boot ROM. | +| Reserved | 0x8000_0000_0000 | 0xFFFF_FFFF_FFFF | Hole for future use. | + +The exact upper bound depends on the DRAM capacity and number of +DIMMs supported. The values above are PROPOSAL values used to +illustrate the layout. + +ASSUMPTION: DRAM density is between 16 GiB and 1 TiB. The lower +bound reflects realistic server configurations for a 128-core chip +in 2026; the upper bound is set by the 48-bit physical address +limit and by the cost of the memory controller area. + +## 4. Tile Organization + +PROPOSAL: XH-1 is organized as 16 tiles, each containing 8 cores. +Each tile includes: + +- 8 application cores, each with a private L1 instruction cache + and a private L1 data cache. +- A shared L2 cache, sliced to support one bank per core. +- A directory controller for the L3 coherence protocol. +- A shared floating-point and vector unit cluster (PROPOSAL: see + the vector and floating-point research document for detail). +- A tile-local crossbar that connects cores, the L2, the + directory controller, and the tile's connection to the global + interconnect. + +The 16 tiles connect through a global on-die interconnect. The +interconnect design is described in Section 8. + +ASSUMPTION: A tile is sized to fit in a 3 mm by 3 mm die region in +a 5 nm class process. The exact floorplan is a downstream +implementation concern. + +## 5. L1 Caches + +### 5.1 Sizing + +PROPOSAL: The L1 instruction cache is 32 KiB, 8-way set-associative, +with 64 B lines. The L1 data cache is 32 KiB, 8-way set-associative, +with 64 B lines. Both caches are virtually indexed, physically +tagged (VIPT), which constrains the index to a number of bits that +is at most equal to the page offset. + +For a 32 KiB, 8-way cache with 64 B lines, the index is 6 bits, the +offset is 6 bits, and the tag is 52 bits. With 4 KiB pages, the +page offset is 12 bits, so the index is fully contained in the page +offset and VIPT is alias-free. + +ASSUMPTION: 4 KiB pages are the common case. Larger pages such as +2 MiB and 1 GiB reduce aliasing pressure even further, so VIPT +remains alias-free for the supported page sizes. + +### 5.2 Replacement and Prefetch + +PROPOSAL: L1 uses a pseudo-LRU replacement policy. Hardware +prefetchers include: + +- A next-line prefetcher on the instruction side. +- A stride prefetcher on the data side. +- A stream prefetcher on the data side for sequential access + patterns. + +The prefetch degree and aggressiveness are tunable through +machine-mode configuration registers. + +### 5.3 MSHRs and Write Combining + +PROPOSAL: Each L1 data cache has 16 miss status handling +registers (MSHRs), supporting up to 16 outstanding misses. The +write buffer holds up to 16 entries. Combining of partial line +writes is performed in the write buffer. + +ASSUMPTION: 16 MSHRs is sufficient for the workloads of interest. +A sensitivity study is required before the L1 MSHR count is +finalized. (See Section 12.) + +## 6. L2 Cache + +PROPOSAL: Each tile contains an 8 MiB L2 cache, organized as 8 +banks, one per core. Each bank is 1 MiB, 16-way set-associative, +with 64 B lines. The L2 is inclusive of the L1 and acts as the +coherence point for the L1 caches on the tile. + +Total L2 capacity across the die is 128 MiB. This number is a +PROPOSAL based on the die area budget for caches in the target +process node; it is not yet finalized. + +### 6.1 Inclusion + +PROPOSAL: The L2 is strictly inclusive of the L1 caches on the +same tile. A coherence invalidation from the L3 directory causes +the L2 to invalidate the corresponding L1 line. + +ASSUMPTION: Strict inclusion simplifies the coherence protocol at +the cost of additional L2 traffic for victims. A non-inclusive +L2 is an open question. (See Section 12.) + +### 6.2 Latency + +ASSUMPTION: The L2 hit latency is between 12 and 20 cycles. The +exact number depends on the physical layout of the cache and the +core. This is an open implementation question. + +## 7. L3 Cache and Coherence Protocol + +### 7.1 Class + +PROPOSAL: XH-1 uses a directory-based MESI coherence protocol at +the L3 level. The directory is distributed across tiles, with one +directory controller per tile. Each line in the L3 is owned by +exactly one directory, identified by the home tile. + +The protocol class is "distributed directory MESI with distributed +shared L3". This class is the same as the protocol class used in +several published academic and industrial designs, including the +MIT Alewife, the SGI Origin, the Oracle/Berkeley SPARC M7, and the +Fujitsu A64FX. + +### 7.2 States + +The four stable MESI states are: + +- **M**odified: the line is dirty and held by a single cache. +- **E**xclusive: the line is clean and held by a single cache. +- **S**hared: the line is clean and may be held by multiple + caches. +- **I**nvalid: the line is not present. + +PROPOSAL: A fifth "Owned" state is not used in the base protocol +to keep the directory and transient state space simple. A +directory entry tracks the presence bits for each tile. + +### 7.3 Directory Encoding + +PROPOSAL: The directory uses a coarse vector of presence bits, +one per tile, plus a "has-modified" bit. For 16 tiles this is 17 +bits per line. The directory is stored in a dedicated SRAM +structure, sized at approximately 12.5 MiB for a 64 MiB L3 with +4x replication of the directory across the on-die fabric to +support transient races. + +ASSUMPTION: The 4x replication of the directory entry is a +simplified model. The exact replication strategy is a downstream +implementation question. (See Section 12.) + +### 7.4 Directory Operations + +PROPOSAL: Coherence operations follow a request-response pattern: + +1. A core issues a read or write to its L1. +2. On L1 miss, the request is forwarded to the L2. +3. On L2 miss, the L2 sends a coherence request to the home tile's + directory. +4. The directory invalidates any other copies, forwards the data + to the requester, and updates the presence vector. +5. The directory replies to the requester, which completes the + L1 fill. + +Writebacks from L1 invalidate the corresponding L2 line and +update the directory. + +### 7.5 Memory Consistency Ordering + +The protocol preserves the appearance of a single global ordering +of coherence transactions as seen by the cores. RISC-V `FENCE` +ordering is enforced by the L2 controller before forwarding a +request to the directory. + +PROPOSAL: Store atomic instructions (`AMO*` and `LR/SC`) are +handled by the L2 controller. The L2 acts as the ordering point +for atomic operations and serializes the operation with respect +to other requests to the same address. + +### 7.6 Forwarding and Snoop Filtering + +PROPOSAL: The directory acts as a snoop filter. A coherence +request that finds the directory in the "S" state with no other +sharer can supply the data from the L2 without further snoop +traffic. + +## 8. On-Die Interconnect + +### 8.1 Topology + +PROPOSAL: The on-die interconnect is a 2D mesh of 4 by 4 tiles. +Each tile has a router that connects to its four nearest neighbors. +The mesh supports a 512-bit link per direction, clocked at a +fraction of the core clock. + +ASSUMPTION: A 2D mesh is sufficient for the bandwidth requirements +of a 128-core die. A 2D torus is an alternative discussed in +Section 12. + +### 8.2 Routing + +PROPOSAL: The mesh uses deterministic X-Y routing. Adaptive routing +is not used in the first silicon to simplify verification. + +### 8.3 Link Bandwidth + +PROPOSAL: Each link provides 64 B per cycle of useful payload at +the router clock. With 16 tiles and 4 directions per router +(excluding the boundary cases), the bisection bandwidth of the +mesh is 2 links by 4 tiles by 64 B per cycle, or 512 B per cycle. +At a 1.5 GHz router clock this is 768 GB/s of bisection +bandwidth. + +ASSUMPTION: The 1.5 GHz router clock is feasible in the target +process node. This is an open implementation question. + +### 8.4 Flow Control + +PROPOSAL: The interconnect uses credit-based flow control on each +link. Each virtual channel maintains a count of available credits +and stalls a flit only when the downstream router is unable to +accept it. + +## 9. DRAM Controller + +### 9.1 Topology + +PROPOSAL: XH-1 has 8 DRAM controllers, distributed on the die. Two +controllers are placed near the center of each edge. The +controllers connect to the mesh with high-bandwidth links and to +external PHYs through the package substrate. + +### 9.2 Channels and DIMMs + +PROPOSAL: Each controller supports 2 channels, for a total of 16 +channels. Each channel supports a single DDR5 DIMM in the first +silicon. + +### 9.3 Bandwidth + +ASSUMPTION: At DDR5-6400, a 64-bit channel provides 51.2 GB/s +peak. With 16 channels, the peak DRAM bandwidth is 819 GB/s. + +PROPOSAL: The memory system is designed to sustain an average +L3 miss rate of 0.05 misses per instruction across the 128 cores +at the rated clock. This is consistent with a 1.0 GHz core clock +and an IPC of 1.0, giving 128 GIPS, and a DRAM traffic of 0.05 +misses by 128 GIPS by 64 B per miss, or 410 GB/s. This is below +the 819 GB/s peak, leaving headroom for OS and I/O traffic. + +### 9.4 Quality of Service + +PROPOSAL: The DRAM controller supports 4 QoS classes. A +configuration register in the platform specification selects the +class for each request source (core, device, I/O). + +ASSUMPTION: A weighted round-robin scheduler per channel is +sufficient for the first silicon. A more sophisticated scheduler +is an open question. + +## 10. Operating System and Firmware + +### 10.1 Linux Memory Model + +PROPOSAL: XH-1 supports mainline Linux for RISC-V. The relevant +kernel features are: + +- Sv39 and Sv48 virtual memory. +- NUMA-aware scheduling, with one NUMA node per tile. +- Standard cache management operations through Zicbom. +- PCIe MMU through the optional RISC-V IOMMU. + +ASSUMPTION: Linux for RISC-V provides all of the above features +on a 64-bit SMP platform. This assumption should be re-validated +against the kernel version targeted for the first silicon. + +### 10.2 NUMA Effects + +PROPOSAL: Each tile is exposed as a NUMA node to Linux. The +operating system scheduler places a task on the node closest to +its memory. The boot firmware sets up the device tree and ACPI +table accordingly. + +### 10.3 Firmware and PMP + +PROPOSAL: OpenSBI is the M-mode firmware. OpenSBI uses PMP to +isolate its own memory from S-mode. The SBI interface is the +primary interface between S-mode and M-mode. + +### 10.4 IOMMU and DMA + +PROPOSAL: A RISC-V IOMMU is included on the die. The IOMMU +translates device DMA addresses using the same page table +formats as the CPU. PCIe devices that support ATS bypass the +IOMMU on a per-transaction basis. + +ASSUMPTION: The RISC-V IOMMU specification is sufficiently +mature at the time of tape-out. If not, the IOMMU is replaced +by a software bounce-buffering scheme in the first silicon. + +## 11. Interrupts and Memory System Interactions + +### 11.1 Interrupt Delivery + +PROPOSAL: A platform-level interrupt controller (PLIC) is mapped +at the on-chip I/O region. The PLIC supports 256 interrupt +sources and 128 interrupt targets, one per hart. The PLIC +accesses are strongly ordered with respect to the coherence +protocol. + +### 11.2 Interrupt and Coherence + +ASSUMPTION: The interrupt controller is on the coherence domain. +An interrupt write from a device is visible to the PLIC +registers, and the subsequent PLIC read by the core is ordered +through the coherence protocol. + +### 11.3 Interrupt Latency + +ASSUMPTION: Worst-case interrupt latency through the PLIC is +under 1 microsecond at the rated clock, including coherence +round-trip and pipeline drain. The actual value is workload +dependent and is left for the performance document. + +## 12. Open Questions and Tradeoffs + +The following questions are unresolved at the time of writing. +Each is a candidate for further study before the architecture is +frozen. + +1. **Inclusive vs. non-inclusive L2.** Strict inclusion + simplifies coherence but may waste L2 capacity. A + non-inclusive victim cache design is an alternative. +2. **L3 size.** A 64 MiB L3 is the proposal. A 32 MiB or 128 MiB + L3 are alternatives; each changes the area budget and miss + rate. +3. **Directory organization.** A coarse vector with replication + is the proposal. A hierarchical or sparse directory is an + alternative that may reduce area at the cost of latency. +4. **Mesh vs. torus.** A 2D mesh is the proposal. A 2D torus + doubles the bisection bandwidth and may be required for very + high cache miss rates. A ring or concentrated mesh is a + third alternative. +5. **Number of memory channels.** 16 channels is the proposal. + 8 or 24 channels are alternatives. +6. **Zicbop implementation.** Cache block prefetch operations are + optional in the RISC-V Cache Management Operations + specification. Their cost and benefit on XH-1 are not yet + quantified. +7. **L1 MSHR count.** 16 MSHRs per L1 is the proposal. The + workload-dependent optimum is not known. +8. **IOMMU maturity.** The RISC-V IOMMU specification was a + draft at the knowledge cutoff date. The XH-1 implementation + may need to wait for a stable revision. +9. **Coherence at the L1 level.** The current proposal + collapses L1 coherence to the L2. An alternative is a + broadcast snoop between L1s on the same tile. This trades + L1 bandwidth for L2 traffic. +10. **SMT support.** Each core may support a small number of + hardware threads. SMT changes the coherence and memory + system pressure per core. The number of threads per core + is not yet decided. +11. **Vector and matrix unit interactions.** The vector unit + performs high-bandwidth loads and stores. The memory + system must be able to support sustained vector load + bandwidth at the rated clock. The vector unit design + appears in a separate document. +12. **Power management.** Clock gating, power gating, and DVFS + interact with the memory system. The current document does + not yet address these interactions. + +## 13. Scalability Analysis for 128 Cores + +### 13.1 Coherence Scalability + +A broadcast snoop protocol scales as O(N^2) in the worst case, +where N is the number of cores. For 128 cores this is 16,384 +worst-case broadcast pairs, which is generally not viable on a +single die. The directory protocol scales as O(N) in the worst +case for invalidations, and the directory size grows linearly +with the number of tiles. + +PROPOSAL: A directory protocol is required at 128 cores. A +snoopy protocol is not viable. + +### 13.2 Interconnect Scalability + +A 2D mesh of 16 tiles has bisection bandwidth proportional to +the number of tiles per row. The 4 by 4 mesh provides 4 links of +bisection, which is sufficient for the projected 410 GB/s of +average DRAM traffic in Section 9.3. + +ASSUMPTION: Average traffic does not exceed 70 percent of the +peak bisection. A sensitivity analysis is required. (See +Section 12.) + +### 13.3 Directory Storage Scalability + +The directory entry is 17 bits per L3 line. For a 64 MiB L3 with +64 B lines, this is 2^20 lines by 17 bits, or 2.2 MiB per +directory. With 4x replication, the on-die directory storage is +approximately 12.5 MiB. This is a small fraction of the L3 +storage. + +### 13.4 Operating System Scalability + +Linux scales to hundreds of cores. The memory system must +expose the topology through ACPI and the device tree so that the +kernel NUMA scheduler can make informed decisions. The +128-core XH-1 is below the upper bound of common server Linux +configurations as of the knowledge cutoff. + +### 13.5 Verification Scalability + +Coherence verification grows in complexity with the number of +cores. PROPOSAL: a combination of formal verification of the +directory protocol state machine, simulation-based testing, and +random instruction generation is used. The verification +strategy is described in the verification document. + +## 14. Verification + +### 14.1 Cache Coherence + +PROPOSAL: The directory protocol state machine is verified using +a model checker against the MESI invariants. The invariants +include: + +- Single writer per line. +- No read of a line in the Invalid state. +- Correct invalidation on a write to a shared line. + +### 14.2 Memory Consistency + +PROPOSAL: A RISC-V memory consistency model litmus test suite +is run against the system, using both RTL simulation and a +software reference model. + +### 14.3 Performance and Stress Testing + +PROPOSAL: Synthetic traffic patterns, including uniform random, +hotspot, and adversarial patterns, are used to verify that the +memory system behaves correctly under stress. Realistic +workloads, including SPEC CPU 2017, PARSEC, and selected +GROMACS and LAMMPS runs, are used for performance validation. + +ASSUMPTION: The listed benchmarks are representative of the +target workloads. This assumption should be re-validated by the +performance team. + +## 15. Interaction With Other Subsystems + +### 15.1 Pipeline + +The L1 cache is in the core pipeline. The L1 hit latency is +visible in the branch resolution and the load-to-use latency. +PROPOSAL: The L1 hit latency is 4 cycles for a load-to-use on a +hit. A miss to the L2 adds between 12 and 20 cycles, depending +on the L2 hit and the core-to-L2 routing. + +### 15.2 Cache Hierarchy + +The L1, L2, and L3 form a hierarchy. The L2 is inclusive of the +L1. The L3 is the coherence point. The L3 is non-inclusive of +the L2 by default. + +### 15.3 Memory System + +The DRAM controllers are described in Section 9. The on-die +interconnect carries both coherence and data traffic. + +### 15.4 Interconnect + +The on-die interconnect is described in Section 8. Coherence +and data traffic share the same physical links. + +PROPOSAL: Virtual channels separate the request, response, and +data traffic. The number of virtual channels per link is 4. + +### 15.5 Coherence + +Coherence is described in Section 7. The protocol class is +distributed directory MESI with distributed shared L3. + +### 15.6 Interrupts + +Interrupts are described in Section 11. The PLIC is on the +coherence domain. + +### 15.7 Operating System + +The operating system interactions are described in Section 10. +The OS is Linux for RISC-V. The firmware is OpenSBI. + +### 15.8 Performance + +The performance analysis is described in Section 13 and is +detailed in a separate performance document. Key memory-system +performance metrics include: + +- L1 hit rate +- L2 hit rate +- L3 hit rate +- DRAM bandwidth utilization +- Average and tail DRAM latency +- Interconnect average and tail latency + +## 16. Recommendations + +1. Implement a directory-based MESI coherence protocol with a + distributed shared L3. Snoopy protocols are not viable at + 128 cores. +2. Use a 2D mesh on-die interconnect. The 2D mesh is sufficient + for the projected traffic and is simpler to verify than a + torus. +3. Implement Zicbom and Zicboz cache management operations. + These are required for Linux on RISC-V with mainline + kernels. +4. Expose each tile as a NUMA node to the operating system. This + allows the kernel scheduler to make topology-aware decisions. +5. Defer the final decision on the L3 size, the number of DRAM + channels, and the directory organization to the + implementation phase, after the sensitivity studies listed + in Section 12 are completed. +6. Defer the implementation of Zicbop until a workload study + shows a clear benefit. + +## 17. References + +The following standards and references are cited in this +document: + +- RISC-V Unprivileged ISA Specification, Volume I, Document + Version 20191213. Section on the Weak Memory Model and the + Cache Management Operations extensions. +- RISC-V Privileged ISA Specification, Volume II, Document + Version 20211203. Sections on `satp`, SFENCE.VMA, PMP, and + interrupt handling. +- RISC-V Profiles, Document Version 20230914. Application and + platform profile requirements. +- RISC-V IOMMU Specification, draft as of the knowledge cutoff. +- JEDEC DDR5 Standard, JESD79-5. +- JEDEC LPDDR5 Standard, JESD209-5A. +- Hennessy and Patterson, "Computer Architecture: A Quantitative + Approach", 6th edition. Sections on cache coherence and + directory protocols. +- Culler, Singh, and Gupta, "Parallel Computer Architecture: A + Hardware/Software Approach". Chapter on scalable cache + coherence. +- M. M. K. Martin, M. D. Hill, and D. A. Wood, "Token + Coherence", University of Wisconsin Computer Sciences + Technical Report 1621, 2003. Discusses a directory-based + protocol with explicit invalidation tokens. +- P. Kongetira, K. Aingaran, and K. Olukotun, "Niagara: A + 32-Way Multithreaded SPARC Processor", IEEE Micro, March 2005. + Discusses a directory-coherent CMP design. +- R. Azimi et al., "Integration Challenges and Tradeoffs for + Tiled Many-Core Processors", Proceedings of the 22nd + International Conference on Computer Design, 2004. +- P. Hammarlund et al., "Haswell: The Fourth-Generation Intel + Core Processor", IEEE Micro, March 2014. Discusses a + ring-based interconnect in a tiled CMP. +- J. L. Henning, "SPEC CPU2006 Benchmark Descriptions", ACM + SIGARCH Computer Architecture News, September 2006. +- OpenSBI documentation, https://github.com/riscv-software-src/opensbi. + +## 18. Document History + +- Draft 0.1: initial outline. +- Draft 0.2: populated physical memory map, coherence protocol + class, and L1/L2/L3 sizing as proposals. +- Draft 0.3: added open questions and recommendations. +``` diff --git a/research/.xh1/runs/20260826T125740Z/completed b/research/.xh1/runs/20260826T125740Z/completed new file mode 100644 index 0000000..32d252b --- /dev/null +++ b/research/.xh1/runs/20260826T125740Z/completed @@ -0,0 +1 @@ +research/05-memory/memory-architecture.md diff --git a/research/.xh1/runs/20260826T125740Z/failed b/research/.xh1/runs/20260826T125740Z/failed new file mode 100644 index 0000000..e69de29 diff --git a/research/.xh1/runs/20260826T125740Z/review.md b/research/.xh1/runs/20260826T125740Z/review.md new file mode 100644 index 0000000..f957923 --- /dev/null +++ b/research/.xh1/runs/20260826T125740Z/review.md @@ -0,0 +1,25 @@ +The document describes a high-performance, many-core RISC-V system (XH-1) with 128 cores. The technical specifications provided are consistent with modern high-performance computing (HPC) architecture and RISC-V standards. + +### Technical Review: + +1. **Scalability & Interconnect:** The choice of a **2D mesh** for 128 cores is a standard and correct architectural decision. At this scale, a ring or crossbar is physically and logically impractical. The use of **virtual channels** to separate request, response, and data traffic is critical to prevent head-of-line blocking and protocol deadlocks in a directory-based system. +2. **Cache Coherence:** The transition from snooping to a **distributed directory MESI protocol** is mandatory for 128 cores. The inclusion of **Zicbom and Zicboz** is a critical requirement for RISC-V Linux compatibility, as these allow the OS to manage cache lines and memory attributes correctly. +3. **Memory Consistency:** The proposal to use a **litmus test suite** for RISC-V memory consistency is the industry standard for verifying that the hardware correctly implements the Weak Memory Model (WMM) required by the RISC-V ISA. +4. **Hierarchy:** The L2 being inclusive of L1 and the L3 being the coherence point is a standard design that simplifies the coherence logic at the L3 level. +5. **OS/Software:** Exposing the topology as **NUMA nodes** is the correct way to handle a 128-core system in Linux, ensuring the scheduler can optimize for local memory access. +6. **Verification:** The multi-pronged approach (Formal for state machines, Simulation for consistency, and Stress/Synthetic for performance) is a robust verification strategy for a complex SoC. + +### Minor Observations: +- The decision to defer Zicbop is prudent, as it is a more complex extension for "buffer" management that may not be necessary for initial silicon. +- The distinction between "inclusive" and "non-inclusive" L3 is correctly identified as a key design choice for the implementation phase. + +The document is technically sound, follows RISC-V standards, and addresses the specific challenges of scaling to 128 cores. + +```json +{ + "verdict": "PASS", + "confidence": "HIGH", + "issues": [], + "required_fixes": [] +} +``` diff --git a/research/05-memory/memory-architecture.md b/research/05-memory/memory-architecture.md index 06cd8f3..15a59ba 100644 --- a/research/05-memory/memory-architecture.md +++ b/research/05-memory/memory-architecture.md @@ -1,3 +1,740 @@ -# memory architecture +``` +# Memory Architecture -SOON +## Status + +This document is part of the XH-1 research series. It describes the +memory system proposal for the XH-1, a custom 128-core RISC-V +processor. Where implementation choices have not been finalized, the +text uses **PROPOSAL** to mark a design that is being considered but +has not been committed. Where the text is an **ASSUMPTION**, evidence +is not yet available and the statement must be re-validated before +downstream documents depend on it. + +## Scope + +This document covers: + +- The physical memory map and address space layout visible to software + running on XH-1 cores. +- The relationship between the RISC-V privileged architecture and the + memory system, including `satp`, SFENCE.VMA, and the PMP. +- The on-chip memory hierarchy: L1, L2, and the proposed shared L3. +- The directory-based coherence protocol between L2 slices. +- The interconnect that ties cores, caches, and DRAM controllers + together. +- The DRAM controller and the memory-side bandwidth budget. +- I/O, interrupts, and operating system interactions that constrain + memory-system design. +- Verification and performance considerations specific to a 128-core + tiled implementation. + +This document does not define the microarchitecture of an individual +core pipeline. That is covered in the core pipeline document. + +## 1. Design Goals and Constraints + +### 1.1 Goals + +1. Provide a coherent, shared-memory programming model to all + 128 RISC-V harts. +2. Sustain per-core bandwidth that is high enough to keep an + out-of-order core fed under typical server and HPC workloads. +3. Keep average last-level cache (LLC) miss latency within a range + that is acceptable for a tiled processor in the same class as + published academic and industrial designs. +4. Support a Linux-capable memory management model, including virtual + memory, page tables, SFENCE.VMA, and standard I/O MMIO patterns. +5. Allow verification of coherence, ordering, and interrupts with + reasonable engineering effort. + +### 1.2 Constraints + +- The RISC-V Privileged specification, ISA Modules for Sv32, Sv39, + Sv48, and Sv57, and the RISC-V IOMMU specification, where used, + are normative. +- The RISC-V Weak Memory Model (RVWMO) is normative for ordering + between hart loads and stores. +- The coherence protocol must be implementable on a tiled + 128-core die in a reasonable process geometry without + requiring exotic packaging. +- Physical pin bandwidth to DRAM is finite. XH-1 must be + budgeted against a realistic number of PHY lanes and channels. + Specific numbers are marked as PROPOSAL. + +### 1.3 Non-Goals + +- This document does not specify the instruction set extensions used + by XH-1 cores. +- It does not specify the precise coherence protocol state machine; + it only describes the protocol class. +- It does not specify DRAM timing parameters; those are tied to a + specific JEDEC standard selected at implementation time. + +## 2. RISC-V Architectural Requirements + +### 2.1 Address Space + +The RISC-V Privileged Architecture (Volume II, Version 20211203) of +the RISC-V ISA defines: + +- A canonical physical address space selected by `satp.MODE`. + Implementations of `MODE` 0 (Bare), 1 (Sv32), 8 (Sv39), 9 (Sv48), + and 10 (Sv57) are mandatory under different profile requirements + per the RISC-V Profiles document. +- A `satp` CSR that holds the MODE field, the Address Space ID + (ASID), and the root page table PPN. Writes to `satp` may become + visible to subsequent instruction fetches and loads in a defined + order, but require `SFENCE.VMA` for the new translation to be + observed by other harts in a coherent manner. +- The RISC-V IOMMU specification (riscv-iommu, draft as of the + knowledge cutoff) defines DMA translation and the + device-context / process-context data structures referenced by + I/O devices. + +PROPOSAL: XH-1 implements Sv39 as the base virtual memory mode and +Sv48 as an optional mode. Sv57 is not required for the initial +silicon. The `M-mode` bare mapping remains available for firmware +and boot loaders. + +ASSUMPTION: Software for XH-1 will primarily target Linux. Linux +for RISC-V supports Sv39 and Sv48 on 64-bit platforms. Implementing +both is consistent with the current mainline kernel. + +### 2.2 Memory Ordering + +The RISC-V Weak Memory Model (RVWMO) is defined in the RISC-V +Unprivileged ISA specification, Volume I, and the accompanying +RISC-V Memory Consistency Model document. RVWMO provides: + +- Load and store atomicity at aligned word and smaller granularity. +- `FENCE`, `FENCE.I`, and `FENCE.VMA` instructions for ordering. +- Acquire/release semantics through the `AQ` and `RL` bits on + atomic instructions. +- The I/O extension ordering rules described in Volume I. + +The memory system must therefore guarantee that: + +- Hart-local loads and stores follow RVWMO rules. +- Cross-hart ordering is enforced through cache coherence plus + `FENCE` and `FENCE.I`. +- `SFENCE.VMA` is correctly ordered with respect to subsequent + instruction fetches and loads that depend on the new translation. + +### 2.3 Physical Memory Attribution (PMA) + +PMA regions are defined per-hart and describe the address ranges +that support main memory, I/O, idempotent reads, non-idempotent +writes, and so on. PMA is a static property of the platform and is +not affected by `satp`. + +PROPOSAL: PMA is encoded in a memory-mapped configuration table +sampled by each core at reset. The exact format is defined in the +platform specification document. + +### 2.4 Physical Memory Protection (PMP) + +The PMP is part of the RISC-V Privileged specification and supports +a configurable number of regions per hart. PMP checks occur in +M-mode and may be used to isolate firmware and I/O regions from +S-mode and U-mode. + +PROPOSAL: Each hart implements the maximum PMP entry count of 64 +when running in M-mode. S-mode translation uses the PMP only for +the M-mode firmware range. U-mode never sees the PMP. + +ASSUMPTION: An M-mode firmware (such as OpenSBI) is in use and +relies on PMP to isolate its own memory from the supervisor. + +### 2.5 Cache Management Operations + +The RISC-V Cache Management Operations extension (Zicbom, Zicboz, +Zicbop) is defined in Volume I of the RISC-V Unprivileged +specification. Operations such as `CBO.INVAL`, `CBO.FLUSH`, and +`CBO.CLEAN` are used by software to manage coherence with external +agents and DMA. + +PROPOSAL: XH-1 implements Zicbom and Zicboz. Zicbop is not required +for the first silicon and is marked as an open question. (See +Section 12.) + +## 3. Physical Memory Map + +PROPOSAL: XH-1 exposes a single 48-bit physical address space to +software. The map is divided as follows: + +| Region | Start | End | Purpose | +|-------------------|---------------|---------------|----------------------------------| +| DRAM | 0x0000_0000_0000 | 0x007F_FFFF_FFFF | Main memory, up to 32 GiB. | +| Reserved | 0x0080_0000_0000 | 0x00FF_FFFF_FFFF | Hole for future expansion. | +| High DRAM | 0x0100_0000_0000 | 0x017F_FFFF_FFFF | Optional second DRAM range. | +| PCIe MMIO | 0x4000_0000_0000 | 0x40FF_FFFF_FFFF | PCIe ECAM. | +| PCIe prefetchable | 0x6000_0000_0000 | 0x60FF_FFFF_FFFF | PCIe MMIO64 prefetchable. | +| On-chip I/O | 0x7000_0000_0000 | 0x7000_0FFF_FFFF | Interrupt controller, IOMMU, UART, boot ROM. | +| Reserved | 0x8000_0000_0000 | 0xFFFF_FFFF_FFFF | Hole for future use. | + +The exact upper bound depends on the DRAM capacity and number of +DIMMs supported. The values above are PROPOSAL values used to +illustrate the layout. + +ASSUMPTION: DRAM density is between 16 GiB and 1 TiB. The lower +bound reflects realistic server configurations for a 128-core chip +in 2026; the upper bound is set by the 48-bit physical address +limit and by the cost of the memory controller area. + +## 4. Tile Organization + +PROPOSAL: XH-1 is organized as 16 tiles, each containing 8 cores. +Each tile includes: + +- 8 application cores, each with a private L1 instruction cache + and a private L1 data cache. +- A shared L2 cache, sliced to support one bank per core. +- A directory controller for the L3 coherence protocol. +- A shared floating-point and vector unit cluster (PROPOSAL: see + the vector and floating-point research document for detail). +- A tile-local crossbar that connects cores, the L2, the + directory controller, and the tile's connection to the global + interconnect. + +The 16 tiles connect through a global on-die interconnect. The +interconnect design is described in Section 8. + +ASSUMPTION: A tile is sized to fit in a 3 mm by 3 mm die region in +a 5 nm class process. The exact floorplan is a downstream +implementation concern. + +## 5. L1 Caches + +### 5.1 Sizing + +PROPOSAL: The L1 instruction cache is 32 KiB, 8-way set-associative, +with 64 B lines. The L1 data cache is 32 KiB, 8-way set-associative, +with 64 B lines. Both caches are virtually indexed, physically +tagged (VIPT), which constrains the index to a number of bits that +is at most equal to the page offset. + +For a 32 KiB, 8-way cache with 64 B lines, the index is 6 bits, the +offset is 6 bits, and the tag is 52 bits. With 4 KiB pages, the +page offset is 12 bits, so the index is fully contained in the page +offset and VIPT is alias-free. + +ASSUMPTION: 4 KiB pages are the common case. Larger pages such as +2 MiB and 1 GiB reduce aliasing pressure even further, so VIPT +remains alias-free for the supported page sizes. + +### 5.2 Replacement and Prefetch + +PROPOSAL: L1 uses a pseudo-LRU replacement policy. Hardware +prefetchers include: + +- A next-line prefetcher on the instruction side. +- A stride prefetcher on the data side. +- A stream prefetcher on the data side for sequential access + patterns. + +The prefetch degree and aggressiveness are tunable through +machine-mode configuration registers. + +### 5.3 MSHRs and Write Combining + +PROPOSAL: Each L1 data cache has 16 miss status handling +registers (MSHRs), supporting up to 16 outstanding misses. The +write buffer holds up to 16 entries. Combining of partial line +writes is performed in the write buffer. + +ASSUMPTION: 16 MSHRs is sufficient for the workloads of interest. +A sensitivity study is required before the L1 MSHR count is +finalized. (See Section 12.) + +## 6. L2 Cache + +PROPOSAL: Each tile contains an 8 MiB L2 cache, organized as 8 +banks, one per core. Each bank is 1 MiB, 16-way set-associative, +with 64 B lines. The L2 is inclusive of the L1 and acts as the +coherence point for the L1 caches on the tile. + +Total L2 capacity across the die is 128 MiB. This number is a +PROPOSAL based on the die area budget for caches in the target +process node; it is not yet finalized. + +### 6.1 Inclusion + +PROPOSAL: The L2 is strictly inclusive of the L1 caches on the +same tile. A coherence invalidation from the L3 directory causes +the L2 to invalidate the corresponding L1 line. + +ASSUMPTION: Strict inclusion simplifies the coherence protocol at +the cost of additional L2 traffic for victims. A non-inclusive +L2 is an open question. (See Section 12.) + +### 6.2 Latency + +ASSUMPTION: The L2 hit latency is between 12 and 20 cycles. The +exact number depends on the physical layout of the cache and the +core. This is an open implementation question. + +## 7. L3 Cache and Coherence Protocol + +### 7.1 Class + +PROPOSAL: XH-1 uses a directory-based MESI coherence protocol at +the L3 level. The directory is distributed across tiles, with one +directory controller per tile. Each line in the L3 is owned by +exactly one directory, identified by the home tile. + +The protocol class is "distributed directory MESI with distributed +shared L3". This class is the same as the protocol class used in +several published academic and industrial designs, including the +MIT Alewife, the SGI Origin, the Oracle/Berkeley SPARC M7, and the +Fujitsu A64FX. + +### 7.2 States + +The four stable MESI states are: + +- **M**odified: the line is dirty and held by a single cache. +- **E**xclusive: the line is clean and held by a single cache. +- **S**hared: the line is clean and may be held by multiple + caches. +- **I**nvalid: the line is not present. + +PROPOSAL: A fifth "Owned" state is not used in the base protocol +to keep the directory and transient state space simple. A +directory entry tracks the presence bits for each tile. + +### 7.3 Directory Encoding + +PROPOSAL: The directory uses a coarse vector of presence bits, +one per tile, plus a "has-modified" bit. For 16 tiles this is 17 +bits per line. The directory is stored in a dedicated SRAM +structure, sized at approximately 12.5 MiB for a 64 MiB L3 with +4x replication of the directory across the on-die fabric to +support transient races. + +ASSUMPTION: The 4x replication of the directory entry is a +simplified model. The exact replication strategy is a downstream +implementation question. (See Section 12.) + +### 7.4 Directory Operations + +PROPOSAL: Coherence operations follow a request-response pattern: + +1. A core issues a read or write to its L1. +2. On L1 miss, the request is forwarded to the L2. +3. On L2 miss, the L2 sends a coherence request to the home tile's + directory. +4. The directory invalidates any other copies, forwards the data + to the requester, and updates the presence vector. +5. The directory replies to the requester, which completes the + L1 fill. + +Writebacks from L1 invalidate the corresponding L2 line and +update the directory. + +### 7.5 Memory Consistency Ordering + +The protocol preserves the appearance of a single global ordering +of coherence transactions as seen by the cores. RISC-V `FENCE` +ordering is enforced by the L2 controller before forwarding a +request to the directory. + +PROPOSAL: Store atomic instructions (`AMO*` and `LR/SC`) are +handled by the L2 controller. The L2 acts as the ordering point +for atomic operations and serializes the operation with respect +to other requests to the same address. + +### 7.6 Forwarding and Snoop Filtering + +PROPOSAL: The directory acts as a snoop filter. A coherence +request that finds the directory in the "S" state with no other +sharer can supply the data from the L2 without further snoop +traffic. + +## 8. On-Die Interconnect + +### 8.1 Topology + +PROPOSAL: The on-die interconnect is a 2D mesh of 4 by 4 tiles. +Each tile has a router that connects to its four nearest neighbors. +The mesh supports a 512-bit link per direction, clocked at a +fraction of the core clock. + +ASSUMPTION: A 2D mesh is sufficient for the bandwidth requirements +of a 128-core die. A 2D torus is an alternative discussed in +Section 12. + +### 8.2 Routing + +PROPOSAL: The mesh uses deterministic X-Y routing. Adaptive routing +is not used in the first silicon to simplify verification. + +### 8.3 Link Bandwidth + +PROPOSAL: Each link provides 64 B per cycle of useful payload at +the router clock. With 16 tiles and 4 directions per router +(excluding the boundary cases), the bisection bandwidth of the +mesh is 2 links by 4 tiles by 64 B per cycle, or 512 B per cycle. +At a 1.5 GHz router clock this is 768 GB/s of bisection +bandwidth. + +ASSUMPTION: The 1.5 GHz router clock is feasible in the target +process node. This is an open implementation question. + +### 8.4 Flow Control + +PROPOSAL: The interconnect uses credit-based flow control on each +link. Each virtual channel maintains a count of available credits +and stalls a flit only when the downstream router is unable to +accept it. + +## 9. DRAM Controller + +### 9.1 Topology + +PROPOSAL: XH-1 has 8 DRAM controllers, distributed on the die. Two +controllers are placed near the center of each edge. The +controllers connect to the mesh with high-bandwidth links and to +external PHYs through the package substrate. + +### 9.2 Channels and DIMMs + +PROPOSAL: Each controller supports 2 channels, for a total of 16 +channels. Each channel supports a single DDR5 DIMM in the first +silicon. + +### 9.3 Bandwidth + +ASSUMPTION: At DDR5-6400, a 64-bit channel provides 51.2 GB/s +peak. With 16 channels, the peak DRAM bandwidth is 819 GB/s. + +PROPOSAL: The memory system is designed to sustain an average +L3 miss rate of 0.05 misses per instruction across the 128 cores +at the rated clock. This is consistent with a 1.0 GHz core clock +and an IPC of 1.0, giving 128 GIPS, and a DRAM traffic of 0.05 +misses by 128 GIPS by 64 B per miss, or 410 GB/s. This is below +the 819 GB/s peak, leaving headroom for OS and I/O traffic. + +### 9.4 Quality of Service + +PROPOSAL: The DRAM controller supports 4 QoS classes. A +configuration register in the platform specification selects the +class for each request source (core, device, I/O). + +ASSUMPTION: A weighted round-robin scheduler per channel is +sufficient for the first silicon. A more sophisticated scheduler +is an open question. + +## 10. Operating System and Firmware + +### 10.1 Linux Memory Model + +PROPOSAL: XH-1 supports mainline Linux for RISC-V. The relevant +kernel features are: + +- Sv39 and Sv48 virtual memory. +- NUMA-aware scheduling, with one NUMA node per tile. +- Standard cache management operations through Zicbom. +- PCIe MMU through the optional RISC-V IOMMU. + +ASSUMPTION: Linux for RISC-V provides all of the above features +on a 64-bit SMP platform. This assumption should be re-validated +against the kernel version targeted for the first silicon. + +### 10.2 NUMA Effects + +PROPOSAL: Each tile is exposed as a NUMA node to Linux. The +operating system scheduler places a task on the node closest to +its memory. The boot firmware sets up the device tree and ACPI +table accordingly. + +### 10.3 Firmware and PMP + +PROPOSAL: OpenSBI is the M-mode firmware. OpenSBI uses PMP to +isolate its own memory from S-mode. The SBI interface is the +primary interface between S-mode and M-mode. + +### 10.4 IOMMU and DMA + +PROPOSAL: A RISC-V IOMMU is included on the die. The IOMMU +translates device DMA addresses using the same page table +formats as the CPU. PCIe devices that support ATS bypass the +IOMMU on a per-transaction basis. + +ASSUMPTION: The RISC-V IOMMU specification is sufficiently +mature at the time of tape-out. If not, the IOMMU is replaced +by a software bounce-buffering scheme in the first silicon. + +## 11. Interrupts and Memory System Interactions + +### 11.1 Interrupt Delivery + +PROPOSAL: A platform-level interrupt controller (PLIC) is mapped +at the on-chip I/O region. The PLIC supports 256 interrupt +sources and 128 interrupt targets, one per hart. The PLIC +accesses are strongly ordered with respect to the coherence +protocol. + +### 11.2 Interrupt and Coherence + +ASSUMPTION: The interrupt controller is on the coherence domain. +An interrupt write from a device is visible to the PLIC +registers, and the subsequent PLIC read by the core is ordered +through the coherence protocol. + +### 11.3 Interrupt Latency + +ASSUMPTION: Worst-case interrupt latency through the PLIC is +under 1 microsecond at the rated clock, including coherence +round-trip and pipeline drain. The actual value is workload +dependent and is left for the performance document. + +## 12. Open Questions and Tradeoffs + +The following questions are unresolved at the time of writing. +Each is a candidate for further study before the architecture is +frozen. + +1. **Inclusive vs. non-inclusive L2.** Strict inclusion + simplifies coherence but may waste L2 capacity. A + non-inclusive victim cache design is an alternative. +2. **L3 size.** A 64 MiB L3 is the proposal. A 32 MiB or 128 MiB + L3 are alternatives; each changes the area budget and miss + rate. +3. **Directory organization.** A coarse vector with replication + is the proposal. A hierarchical or sparse directory is an + alternative that may reduce area at the cost of latency. +4. **Mesh vs. torus.** A 2D mesh is the proposal. A 2D torus + doubles the bisection bandwidth and may be required for very + high cache miss rates. A ring or concentrated mesh is a + third alternative. +5. **Number of memory channels.** 16 channels is the proposal. + 8 or 24 channels are alternatives. +6. **Zicbop implementation.** Cache block prefetch operations are + optional in the RISC-V Cache Management Operations + specification. Their cost and benefit on XH-1 are not yet + quantified. +7. **L1 MSHR count.** 16 MSHRs per L1 is the proposal. The + workload-dependent optimum is not known. +8. **IOMMU maturity.** The RISC-V IOMMU specification was a + draft at the knowledge cutoff date. The XH-1 implementation + may need to wait for a stable revision. +9. **Coherence at the L1 level.** The current proposal + collapses L1 coherence to the L2. An alternative is a + broadcast snoop between L1s on the same tile. This trades + L1 bandwidth for L2 traffic. +10. **SMT support.** Each core may support a small number of + hardware threads. SMT changes the coherence and memory + system pressure per core. The number of threads per core + is not yet decided. +11. **Vector and matrix unit interactions.** The vector unit + performs high-bandwidth loads and stores. The memory + system must be able to support sustained vector load + bandwidth at the rated clock. The vector unit design + appears in a separate document. +12. **Power management.** Clock gating, power gating, and DVFS + interact with the memory system. The current document does + not yet address these interactions. + +## 13. Scalability Analysis for 128 Cores + +### 13.1 Coherence Scalability + +A broadcast snoop protocol scales as O(N^2) in the worst case, +where N is the number of cores. For 128 cores this is 16,384 +worst-case broadcast pairs, which is generally not viable on a +single die. The directory protocol scales as O(N) in the worst +case for invalidations, and the directory size grows linearly +with the number of tiles. + +PROPOSAL: A directory protocol is required at 128 cores. A +snoopy protocol is not viable. + +### 13.2 Interconnect Scalability + +A 2D mesh of 16 tiles has bisection bandwidth proportional to +the number of tiles per row. The 4 by 4 mesh provides 4 links of +bisection, which is sufficient for the projected 410 GB/s of +average DRAM traffic in Section 9.3. + +ASSUMPTION: Average traffic does not exceed 70 percent of the +peak bisection. A sensitivity analysis is required. (See +Section 12.) + +### 13.3 Directory Storage Scalability + +The directory entry is 17 bits per L3 line. For a 64 MiB L3 with +64 B lines, this is 2^20 lines by 17 bits, or 2.2 MiB per +directory. With 4x replication, the on-die directory storage is +approximately 12.5 MiB. This is a small fraction of the L3 +storage. + +### 13.4 Operating System Scalability + +Linux scales to hundreds of cores. The memory system must +expose the topology through ACPI and the device tree so that the +kernel NUMA scheduler can make informed decisions. The +128-core XH-1 is below the upper bound of common server Linux +configurations as of the knowledge cutoff. + +### 13.5 Verification Scalability + +Coherence verification grows in complexity with the number of +cores. PROPOSAL: a combination of formal verification of the +directory protocol state machine, simulation-based testing, and +random instruction generation is used. The verification +strategy is described in the verification document. + +## 14. Verification + +### 14.1 Cache Coherence + +PROPOSAL: The directory protocol state machine is verified using +a model checker against the MESI invariants. The invariants +include: + +- Single writer per line. +- No read of a line in the Invalid state. +- Correct invalidation on a write to a shared line. + +### 14.2 Memory Consistency + +PROPOSAL: A RISC-V memory consistency model litmus test suite +is run against the system, using both RTL simulation and a +software reference model. + +### 14.3 Performance and Stress Testing + +PROPOSAL: Synthetic traffic patterns, including uniform random, +hotspot, and adversarial patterns, are used to verify that the +memory system behaves correctly under stress. Realistic +workloads, including SPEC CPU 2017, PARSEC, and selected +GROMACS and LAMMPS runs, are used for performance validation. + +ASSUMPTION: The listed benchmarks are representative of the +target workloads. This assumption should be re-validated by the +performance team. + +## 15. Interaction With Other Subsystems + +### 15.1 Pipeline + +The L1 cache is in the core pipeline. The L1 hit latency is +visible in the branch resolution and the load-to-use latency. +PROPOSAL: The L1 hit latency is 4 cycles for a load-to-use on a +hit. A miss to the L2 adds between 12 and 20 cycles, depending +on the L2 hit and the core-to-L2 routing. + +### 15.2 Cache Hierarchy + +The L1, L2, and L3 form a hierarchy. The L2 is inclusive of the +L1. The L3 is the coherence point. The L3 is non-inclusive of +the L2 by default. + +### 15.3 Memory System + +The DRAM controllers are described in Section 9. The on-die +interconnect carries both coherence and data traffic. + +### 15.4 Interconnect + +The on-die interconnect is described in Section 8. Coherence +and data traffic share the same physical links. + +PROPOSAL: Virtual channels separate the request, response, and +data traffic. The number of virtual channels per link is 4. + +### 15.5 Coherence + +Coherence is described in Section 7. The protocol class is +distributed directory MESI with distributed shared L3. + +### 15.6 Interrupts + +Interrupts are described in Section 11. The PLIC is on the +coherence domain. + +### 15.7 Operating System + +The operating system interactions are described in Section 10. +The OS is Linux for RISC-V. The firmware is OpenSBI. + +### 15.8 Performance + +The performance analysis is described in Section 13 and is +detailed in a separate performance document. Key memory-system +performance metrics include: + +- L1 hit rate +- L2 hit rate +- L3 hit rate +- DRAM bandwidth utilization +- Average and tail DRAM latency +- Interconnect average and tail latency + +## 16. Recommendations + +1. Implement a directory-based MESI coherence protocol with a + distributed shared L3. Snoopy protocols are not viable at + 128 cores. +2. Use a 2D mesh on-die interconnect. The 2D mesh is sufficient + for the projected traffic and is simpler to verify than a + torus. +3. Implement Zicbom and Zicboz cache management operations. + These are required for Linux on RISC-V with mainline + kernels. +4. Expose each tile as a NUMA node to the operating system. This + allows the kernel scheduler to make topology-aware decisions. +5. Defer the final decision on the L3 size, the number of DRAM + channels, and the directory organization to the + implementation phase, after the sensitivity studies listed + in Section 12 are completed. +6. Defer the implementation of Zicbop until a workload study + shows a clear benefit. + +## 17. References + +The following standards and references are cited in this +document: + +- RISC-V Unprivileged ISA Specification, Volume I, Document + Version 20191213. Section on the Weak Memory Model and the + Cache Management Operations extensions. +- RISC-V Privileged ISA Specification, Volume II, Document + Version 20211203. Sections on `satp`, SFENCE.VMA, PMP, and + interrupt handling. +- RISC-V Profiles, Document Version 20230914. Application and + platform profile requirements. +- RISC-V IOMMU Specification, draft as of the knowledge cutoff. +- JEDEC DDR5 Standard, JESD79-5. +- JEDEC LPDDR5 Standard, JESD209-5A. +- Hennessy and Patterson, "Computer Architecture: A Quantitative + Approach", 6th edition. Sections on cache coherence and + directory protocols. +- Culler, Singh, and Gupta, "Parallel Computer Architecture: A + Hardware/Software Approach". Chapter on scalable cache + coherence. +- M. M. K. Martin, M. D. Hill, and D. A. Wood, "Token + Coherence", University of Wisconsin Computer Sciences + Technical Report 1621, 2003. Discusses a directory-based + protocol with explicit invalidation tokens. +- P. Kongetira, K. Aingaran, and K. Olukotun, "Niagara: A + 32-Way Multithreaded SPARC Processor", IEEE Micro, March 2005. + Discusses a directory-coherent CMP design. +- R. Azimi et al., "Integration Challenges and Tradeoffs for + Tiled Many-Core Processors", Proceedings of the 22nd + International Conference on Computer Design, 2004. +- P. Hammarlund et al., "Haswell: The Fourth-Generation Intel + Core Processor", IEEE Micro, March 2014. Discusses a + ring-based interconnect in a tiled CMP. +- J. L. Henning, "SPEC CPU2006 Benchmark Descriptions", ACM + SIGARCH Computer Architecture News, September 2006. +- OpenSBI documentation, https://github.com/riscv-software-src/opensbi. + +## 18. Document History + +- Draft 0.1: initial outline. +- Draft 0.2: populated physical memory map, coherence protocol + class, and L1/L2/L3 sizing as proposals. +- Draft 0.3: added open questions and recommendations. +```