mirror of
https://github.com/allexanderbergmns/xh1-research.git
synced 2026-08-26 18:47:01 +00:00
375 lines
31 KiB
JSON
375 lines
31 KiB
JSON
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
{"id":"gen-1787749061-v4qDwmN5u7ZMB1i3n67T","object":"chat.completion","created":1787749061,"model":"minimax/minimax-m3:free","provider":"GMICloud","system_fingerprint":null,"service_tier":null,"choices":[{"index":0,"logprobs":null,"finish_reason":"stop","native_finish_reason":"stop","message":{"role":"assistant","content":"```\n# Memory Architecture\n\n## Status\n\nThis document is part of the XH-1 research series. It describes the\nmemory system proposal for the XH-1, a custom 128-core RISC-V\nprocessor. Where implementation choices have not been finalized, the\ntext uses **PROPOSAL** to mark a design that is being considered but\nhas not been committed. Where the text is an **ASSUMPTION**, evidence\nis not yet available and the statement must be re-validated before\ndownstream documents depend on it.\n\n## Scope\n\nThis document covers:\n\n- The physical memory map and address space layout visible to software\n running on XH-1 cores.\n- The relationship between the RISC-V privileged architecture and the\n memory system, including `satp`, SFENCE.VMA, and the PMP.\n- The on-chip memory hierarchy: L1, L2, and the proposed shared L3.\n- The directory-based coherence protocol between L2 slices.\n- The interconnect that ties cores, caches, and DRAM controllers\n together.\n- The DRAM controller and the memory-side bandwidth budget.\n- I/O, interrupts, and operating system interactions that constrain\n memory-system design.\n- Verification and performance considerations specific to a 128-core\n tiled implementation.\n\nThis document does not define the microarchitecture of an individual\ncore pipeline. That is covered in the core pipeline document.\n\n## 1. Design Goals and Constraints\n\n### 1.1 Goals\n\n1. Provide a coherent, shared-memory programming model to all\n 128 RISC-V harts.\n2. Sustain per-core bandwidth that is high enough to keep an\n out-of-order core fed under typical server and HPC workloads.\n3. Keep average last-level cache (LLC) miss latency within a range\n that is acceptable for a tiled processor in the same class as\n published academic and industrial designs.\n4. Support a Linux-capable memory management model, including virtual\n memory, page tables, SFENCE.VMA, and standard I/O MMIO patterns.\n5. Allow verification of coherence, ordering, and interrupts with\n reasonable engineering effort.\n\n### 1.2 Constraints\n\n- The RISC-V Privileged specification, ISA Modules for Sv32, Sv39,\n Sv48, and Sv57, and the RISC-V IOMMU specification, where used,\n are normative.\n- The RISC-V Weak Memory Model (RVWMO) is normative for ordering\n between hart loads and stores.\n- The coherence protocol must be implementable on a tiled\n 128-core die in a reasonable process geometry without\n requiring exotic packaging.\n- Physical pin bandwidth to DRAM is finite. XH-1 must be\n budgeted against a realistic number of PHY lanes and channels.\n Specific numbers are marked as PROPOSAL.\n\n### 1.3 Non-Goals\n\n- This document does not specify the instruction set extensions used\n by XH-1 cores.\n- It does not specify the precise coherence protocol state machine;\n it only describes the protocol class.\n- It does not specify DRAM timing parameters; those are tied to a\n specific JEDEC standard selected at implementation time.\n\n## 2. RISC-V Architectural Requirements\n\n### 2.1 Address Space\n\nThe RISC-V Privileged Architecture (Volume II, Version 20211203) of\nthe RISC-V ISA defines:\n\n- A canonical physical address space selected by `satp.MODE`.\n Implementations of `MODE` 0 (Bare), 1 (Sv32), 8 (Sv39), 9 (Sv48),\n and 10 (Sv57) are mandatory under different profile requirements\n per the RISC-V Profiles document.\n- A `satp` CSR that holds the MODE field, the Address Space ID\n (ASID), and the root page table PPN. Writes to `satp` may become\n visible to subsequent instruction fetches and loads in a defined\n order, but require `SFENCE.VMA` for the new translation to be\n observed by other harts in a coherent manner.\n- The RISC-V IOMMU specification (riscv-iommu, draft as of the\n knowledge cutoff) defines DMA translation and the\n device-context / process-context data structures referenced by\n I/O devices.\n\nPROPOSAL: XH-1 implements Sv39 as the base virtual memory mode and\nSv48 as an optional mode. Sv57 is not required for the initial\nsilicon. The `M-mode` bare mapping remains available for firmware\nand boot loaders.\n\nASSUMPTION: Software for XH-1 will primarily target Linux. Linux\nfor RISC-V supports Sv39 and Sv48 on 64-bit platforms. Implementing\nboth is consistent with the current mainline kernel.\n\n### 2.2 Memory Ordering\n\nThe RISC-V Weak Memory Model (RVWMO) is defined in the RISC-V\nUnprivileged ISA specification, Volume I, and the accompanying\nRISC-V Memory Consistency Model document. RVWMO provides:\n\n- Load and store atomicity at aligned word and smaller granularity.\n- `FENCE`, `FENCE.I`, and `FENCE.VMA` instructions for ordering.\n- Acquire/release semantics through the `AQ` and `RL` bits on\n atomic instructions.\n- The I/O extension ordering rules described in Volume I.\n\nThe memory system must therefore guarantee that:\n\n- Hart-local loads and stores follow RVWMO rules.\n- Cross-hart ordering is enforced through cache coherence plus\n `FENCE` and `FENCE.I`.\n- `SFENCE.VMA` is correctly ordered with respect to subsequent\n instruction fetches and loads that depend on the new translation.\n\n### 2.3 Physical Memory Attribution (PMA)\n\nPMA regions are defined per-hart and describe the address ranges\nthat support main memory, I/O, idempotent reads, non-idempotent\nwrites, and so on. PMA is a static property of the platform and is\nnot affected by `satp`.\n\nPROPOSAL: PMA is encoded in a memory-mapped configuration table\nsampled by each core at reset. The exact format is defined in the\nplatform specification document.\n\n### 2.4 Physical Memory Protection (PMP)\n\nThe PMP is part of the RISC-V Privileged specification and supports\na configurable number of regions per hart. PMP checks occur in\nM-mode and may be used to isolate firmware and I/O regions from\nS-mode and U-mode.\n\nPROPOSAL: Each hart implements the maximum PMP entry count of 64\nwhen running in M-mode. S-mode translation uses the PMP only for\nthe M-mode firmware range. U-mode never sees the PMP.\n\nASSUMPTION: An M-mode firmware (such as OpenSBI) is in use and\nrelies on PMP to isolate its own memory from the supervisor.\n\n### 2.5 Cache Management Operations\n\nThe RISC-V Cache Management Operations extension (Zicbom, Zicboz,\nZicbop) is defined in Volume I of the RISC-V Unprivileged\nspecification. Operations such as `CBO.INVAL`, `CBO.FLUSH`, and\n`CBO.CLEAN` are used by software to manage coherence with external\nagents and DMA.\n\nPROPOSAL: XH-1 implements Zicbom and Zicboz. Zicbop is not required\nfor the first silicon and is marked as an open question. (See\nSection 12.)\n\n## 3. Physical Memory Map\n\nPROPOSAL: XH-1 exposes a single 48-bit physical address space to\nsoftware. The map is divided as follows:\n\n| Region | Start | End | Purpose |\n|-------------------|---------------|---------------|----------------------------------|\n| DRAM | 0x0000_0000_0000 | 0x007F_FFFF_FFFF | Main memory, up to 32 GiB. |\n| Reserved | 0x0080_0000_0000 | 0x00FF_FFFF_FFFF | Hole for future expansion. |\n| High DRAM | 0x0100_0000_0000 | 0x017F_FFFF_FFFF | Optional second DRAM range. |\n| PCIe MMIO | 0x4000_0000_0000 | 0x40FF_FFFF_FFFF | PCIe ECAM. |\n| PCIe prefetchable | 0x6000_0000_0000 | 0x60FF_FFFF_FFFF | PCIe MMIO64 prefetchable. |\n| On-chip I/O | 0x7000_0000_0000 | 0x7000_0FFF_FFFF | Interrupt controller, IOMMU, UART, boot ROM. |\n| Reserved | 0x8000_0000_0000 | 0xFFFF_FFFF_FFFF | Hole for future use. |\n\nThe exact upper bound depends on the DRAM capacity and number of\nDIMMs supported. The values above are PROPOSAL values used to\nillustrate the layout.\n\nASSUMPTION: DRAM density is between 16 GiB and 1 TiB. The lower\nbound reflects realistic server configurations for a 128-core chip\nin 2026; the upper bound is set by the 48-bit physical address\nlimit and by the cost of the memory controller area.\n\n## 4. Tile Organization\n\nPROPOSAL: XH-1 is organized as 16 tiles, each containing 8 cores.\nEach tile includes:\n\n- 8 application cores, each with a private L1 instruction cache\n and a private L1 data cache.\n- A shared L2 cache, sliced to support one bank per core.\n- A directory controller for the L3 coherence protocol.\n- A shared floating-point and vector unit cluster (PROPOSAL: see\n the vector and floating-point research document for detail).\n- A tile-local crossbar that connects cores, the L2, the\n directory controller, and the tile's connection to the global\n interconnect.\n\nThe 16 tiles connect through a global on-die interconnect. The\ninterconnect design is described in Section 8.\n\nASSUMPTION: A tile is sized to fit in a 3 mm by 3 mm die region in\na 5 nm class process. The exact floorplan is a downstream\nimplementation concern.\n\n## 5. L1 Caches\n\n### 5.1 Sizing\n\nPROPOSAL: The L1 instruction cache is 32 KiB, 8-way set-associative,\nwith 64 B lines. The L1 data cache is 32 KiB, 8-way set-associative,\nwith 64 B lines. Both caches are virtually indexed, physically\ntagged (VIPT), which constrains the index to a number of bits that\nis at most equal to the page offset.\n\nFor a 32 KiB, 8-way cache with 64 B lines, the index is 6 bits, the\noffset is 6 bits, and the tag is 52 bits. With 4 KiB pages, the\npage offset is 12 bits, so the index is fully contained in the page\noffset and VIPT is alias-free.\n\nASSUMPTION: 4 KiB pages are the common case. Larger pages such as\n2 MiB and 1 GiB reduce aliasing pressure even further, so VIPT\nremains alias-free for the supported page sizes.\n\n### 5.2 Replacement and Prefetch\n\nPROPOSAL: L1 uses a pseudo-LRU replacement policy. Hardware\nprefetchers include:\n\n- A next-line prefetcher on the instruction side.\n- A stride prefetcher on the data side.\n- A stream prefetcher on the data side for sequential access\n patterns.\n\nThe prefetch degree and aggressiveness are tunable through\nmachine-mode configuration registers.\n\n### 5.3 MSHRs and Write Combining\n\nPROPOSAL: Each L1 data cache has 16 miss status handling\nregisters (MSHRs), supporting up to 16 outstanding misses. The\nwrite buffer holds up to 16 entries. Combining of partial line\nwrites is performed in the write buffer.\n\nASSUMPTION: 16 MSHRs is sufficient for the workloads of interest.\nA sensitivity study is required before the L1 MSHR count is\nfinalized. (See Section 12.)\n\n## 6. L2 Cache\n\nPROPOSAL: Each tile contains an 8 MiB L2 cache, organized as 8\nbanks, one per core. Each bank is 1 MiB, 16-way set-associative,\nwith 64 B lines. The L2 is inclusive of the L1 and acts as the\ncoherence point for the L1 caches on the tile.\n\nTotal L2 capacity across the die is 128 MiB. This number is a\nPROPOSAL based on the die area budget for caches in the target\nprocess node; it is not yet finalized.\n\n### 6.1 Inclusion\n\nPROPOSAL: The L2 is strictly inclusive of the L1 caches on the\nsame tile. A coherence invalidation from the L3 directory causes\nthe L2 to invalidate the corresponding L1 line.\n\nASSUMPTION: Strict inclusion simplifies the coherence protocol at\nthe cost of additional L2 traffic for victims. A non-inclusive\nL2 is an open question. (See Section 12.)\n\n### 6.2 Latency\n\nASSUMPTION: The L2 hit latency is between 12 and 20 cycles. The\nexact number depends on the physical layout of the cache and the\ncore. This is an open implementation question.\n\n## 7. L3 Cache and Coherence Protocol\n\n### 7.1 Class\n\nPROPOSAL: XH-1 uses a directory-based MESI coherence protocol at\nthe L3 level. The directory is distributed across tiles, with one\ndirectory controller per tile. Each line in the L3 is owned by\nexactly one directory, identified by the home tile.\n\nThe protocol class is \"distributed directory MESI with distributed\nshared L3\". This class is the same as the protocol class used in\nseveral published academic and industrial designs, including the\nMIT Alewife, the SGI Origin, the Oracle/Berkeley SPARC M7, and the\nFujitsu A64FX.\n\n### 7.2 States\n\nThe four stable MESI states are:\n\n- **M**odified: the line is dirty and held by a single cache.\n- **E**xclusive: the line is clean and held by a single cache.\n- **S**hared: the line is clean and may be held by multiple\n caches.\n- **I**nvalid: the line is not present.\n\nPROPOSAL: A fifth \"Owned\" state is not used in the base protocol\nto keep the directory and transient state space simple. A\ndirectory entry tracks the presence bits for each tile.\n\n### 7.3 Directory Encoding\n\nPROPOSAL: The directory uses a coarse vector of presence bits,\none per tile, plus a \"has-modified\" bit. For 16 tiles this is 17\nbits per line. The directory is stored in a dedicated SRAM\nstructure, sized at approximately 12.5 MiB for a 64 MiB L3 with\n4x replication of the directory across the on-die fabric to\nsupport transient races.\n\nASSUMPTION: The 4x replication of the directory entry is a\nsimplified model. The exact replication strategy is a downstream\nimplementation question. (See Section 12.)\n\n### 7.4 Directory Operations\n\nPROPOSAL: Coherence operations follow a request-response pattern:\n\n1. A core issues a read or write to its L1.\n2. On L1 miss, the request is forwarded to the L2.\n3. On L2 miss, the L2 sends a coherence request to the home tile's\n directory.\n4. The directory invalidates any other copies, forwards the data\n to the requester, and updates the presence vector.\n5. The directory replies to the requester, which completes the\n L1 fill.\n\nWritebacks from L1 invalidate the corresponding L2 line and\nupdate the directory.\n\n### 7.5 Memory Consistency Ordering\n\nThe protocol preserves the appearance of a single global ordering\nof coherence transactions as seen by the cores. RISC-V `FENCE`\nordering is enforced by the L2 controller before forwarding a\nrequest to the directory.\n\nPROPOSAL: Store atomic instructions (`AMO*` and `LR/SC`) are\nhandled by the L2 controller. The L2 acts as the ordering point\nfor atomic operations and serializes the operation with respect\nto other requests to the same address.\n\n### 7.6 Forwarding and Snoop Filtering\n\nPROPOSAL: The directory acts as a snoop filter. A coherence\nrequest that finds the directory in the \"S\" state with no other\nsharer can supply the data from the L2 without further snoop\ntraffic.\n\n## 8. On-Die Interconnect\n\n### 8.1 Topology\n\nPROPOSAL: The on-die interconnect is a 2D mesh of 4 by 4 tiles.\nEach tile has a router that connects to its four nearest neighbors.\nThe mesh supports a 512-bit link per direction, clocked at a\nfraction of the core clock.\n\nASSUMPTION: A 2D mesh is sufficient for the bandwidth requirements\nof a 128-core die. A 2D torus is an alternative discussed in\nSection 12.\n\n### 8.2 Routing\n\nPROPOSAL: The mesh uses deterministic X-Y routing. Adaptive routing\nis not used in the first silicon to simplify verification.\n\n### 8.3 Link Bandwidth\n\nPROPOSAL: Each link provides 64 B per cycle of useful payload at\nthe router clock. With 16 tiles and 4 directions per router\n(excluding the boundary cases), the bisection bandwidth of the\nmesh is 2 links by 4 tiles by 64 B per cycle, or 512 B per cycle.\nAt a 1.5 GHz router clock this is 768 GB/s of bisection\nbandwidth.\n\nASSUMPTION: The 1.5 GHz router clock is feasible in the target\nprocess node. This is an open implementation question.\n\n### 8.4 Flow Control\n\nPROPOSAL: The interconnect uses credit-based flow control on each\nlink. Each virtual channel maintains a count of available credits\nand stalls a flit only when the downstream router is unable to\naccept it.\n\n## 9. DRAM Controller\n\n### 9.1 Topology\n\nPROPOSAL: XH-1 has 8 DRAM controllers, distributed on the die. Two\ncontrollers are placed near the center of each edge. The\ncontrollers connect to the mesh with high-bandwidth links and to\nexternal PHYs through the package substrate.\n\n### 9.2 Channels and DIMMs\n\nPROPOSAL: Each controller supports 2 channels, for a total of 16\nchannels. Each channel supports a single DDR5 DIMM in the first\nsilicon.\n\n### 9.3 Bandwidth\n\nASSUMPTION: At DDR5-6400, a 64-bit channel provides 51.2 GB/s\npeak. With 16 channels, the peak DRAM bandwidth is 819 GB/s.\n\nPROPOSAL: The memory system is designed to sustain an average\nL3 miss rate of 0.05 misses per instruction across the 128 cores\nat the rated clock. This is consistent with a 1.0 GHz core clock\nand an IPC of 1.0, giving 128 GIPS, and a DRAM traffic of 0.05\nmisses by 128 GIPS by 64 B per miss, or 410 GB/s. This is below\nthe 819 GB/s peak, leaving headroom for OS and I/O traffic.\n\n### 9.4 Quality of Service\n\nPROPOSAL: The DRAM controller supports 4 QoS classes. A\nconfiguration register in the platform specification selects the\nclass for each request source (core, device, I/O).\n\nASSUMPTION: A weighted round-robin scheduler per channel is\nsufficient for the first silicon. A more sophisticated scheduler\nis an open question.\n\n## 10. Operating System and Firmware\n\n### 10.1 Linux Memory Model\n\nPROPOSAL: XH-1 supports mainline Linux for RISC-V. The relevant\nkernel features are:\n\n- Sv39 and Sv48 virtual memory.\n- NUMA-aware scheduling, with one NUMA node per tile.\n- Standard cache management operations through Zicbom.\n- PCIe MMU through the optional RISC-V IOMMU.\n\nASSUMPTION: Linux for RISC-V provides all of the above features\non a 64-bit SMP platform. This assumption should be re-validated\nagainst the kernel version targeted for the first silicon.\n\n### 10.2 NUMA Effects\n\nPROPOSAL: Each tile is exposed as a NUMA node to Linux. The\noperating system scheduler places a task on the node closest to\nits memory. The boot firmware sets up the device tree and ACPI\ntable accordingly.\n\n### 10.3 Firmware and PMP\n\nPROPOSAL: OpenSBI is the M-mode firmware. OpenSBI uses PMP to\nisolate its own memory from S-mode. The SBI interface is the\nprimary interface between S-mode and M-mode.\n\n### 10.4 IOMMU and DMA\n\nPROPOSAL: A RISC-V IOMMU is included on the die. The IOMMU\ntranslates device DMA addresses using the same page table\nformats as the CPU. PCIe devices that support ATS bypass the\nIOMMU on a per-transaction basis.\n\nASSUMPTION: The RISC-V IOMMU specification is sufficiently\nmature at the time of tape-out. If not, the IOMMU is replaced\nby a software bounce-buffering scheme in the first silicon.\n\n## 11. Interrupts and Memory System Interactions\n\n### 11.1 Interrupt Delivery\n\nPROPOSAL: A platform-level interrupt controller (PLIC) is mapped\nat the on-chip I/O region. The PLIC supports 256 interrupt\nsources and 128 interrupt targets, one per hart. The PLIC\naccesses are strongly ordered with respect to the coherence\nprotocol.\n\n### 11.2 Interrupt and Coherence\n\nASSUMPTION: The interrupt controller is on the coherence domain.\nAn interrupt write from a device is visible to the PLIC\nregisters, and the subsequent PLIC read by the core is ordered\nthrough the coherence protocol.\n\n### 11.3 Interrupt Latency\n\nASSUMPTION: Worst-case interrupt latency through the PLIC is\nunder 1 microsecond at the rated clock, including coherence\nround-trip and pipeline drain. The actual value is workload\ndependent and is left for the performance document.\n\n## 12. Open Questions and Tradeoffs\n\nThe following questions are unresolved at the time of writing.\nEach is a candidate for further study before the architecture is\nfrozen.\n\n1. **Inclusive vs. non-inclusive L2.** Strict inclusion\n simplifies coherence but may waste L2 capacity. A\n non-inclusive victim cache design is an alternative.\n2. **L3 size.** A 64 MiB L3 is the proposal. A 32 MiB or 128 MiB\n L3 are alternatives; each changes the area budget and miss\n rate.\n3. **Directory organization.** A coarse vector with replication\n is the proposal. A hierarchical or sparse directory is an\n alternative that may reduce area at the cost of latency.\n4. **Mesh vs. torus.** A 2D mesh is the proposal. A 2D torus\n doubles the bisection bandwidth and may be required for very\n high cache miss rates. A ring or concentrated mesh is a\n third alternative.\n5. **Number of memory channels.** 16 channels is the proposal.\n 8 or 24 channels are alternatives.\n6. **Zicbop implementation.** Cache block prefetch operations are\n optional in the RISC-V Cache Management Operations\n specification. Their cost and benefit on XH-1 are not yet\n quantified.\n7. **L1 MSHR count.** 16 MSHRs per L1 is the proposal. The\n workload-dependent optimum is not known.\n8. **IOMMU maturity.** The RISC-V IOMMU specification was a\n draft at the knowledge cutoff date. The XH-1 implementation\n may need to wait for a stable revision.\n9. **Coherence at the L1 level.** The current proposal\n collapses L1 coherence to the L2. An alternative is a\n broadcast snoop between L1s on the same tile. This trades\n L1 bandwidth for L2 traffic.\n10. **SMT support.** Each core may support a small number of\n hardware threads. SMT changes the coherence and memory\n system pressure per core. The number of threads per core\n is not yet decided.\n11. **Vector and matrix unit interactions.** The vector unit\n performs high-bandwidth loads and stores. The memory\n system must be able to support sustained vector load\n bandwidth at the rated clock. The vector unit design\n appears in a separate document.\n12. **Power management.** Clock gating, power gating, and DVFS\n interact with the memory system. The current document does\n not yet address these interactions.\n\n## 13. Scalability Analysis for 128 Cores\n\n### 13.1 Coherence Scalability\n\nA broadcast snoop protocol scales as O(N^2) in the worst case,\nwhere N is the number of cores. For 128 cores this is 16,384\nworst-case broadcast pairs, which is generally not viable on a\nsingle die. The directory protocol scales as O(N) in the worst\ncase for invalidations, and the directory size grows linearly\nwith the number of tiles.\n\nPROPOSAL: A directory protocol is required at 128 cores. A\nsnoopy protocol is not viable.\n\n### 13.2 Interconnect Scalability\n\nA 2D mesh of 16 tiles has bisection bandwidth proportional to\nthe number of tiles per row. The 4 by 4 mesh provides 4 links of\nbisection, which is sufficient for the projected 410 GB/s of\naverage DRAM traffic in Section 9.3.\n\nASSUMPTION: Average traffic does not exceed 70 percent of the\npeak bisection. A sensitivity analysis is required. (See\nSection 12.)\n\n### 13.3 Directory Storage Scalability\n\nThe directory entry is 17 bits per L3 line. For a 64 MiB L3 with\n64 B lines, this is 2^20 lines by 17 bits, or 2.2 MiB per\ndirectory. With 4x replication, the on-die directory storage is\napproximately 12.5 MiB. This is a small fraction of the L3\nstorage.\n\n### 13.4 Operating System Scalability\n\nLinux scales to hundreds of cores. The memory system must\nexpose the topology through ACPI and the device tree so that the\nkernel NUMA scheduler can make informed decisions. The\n128-core XH-1 is below the upper bound of common server Linux\nconfigurations as of the knowledge cutoff.\n\n### 13.5 Verification Scalability\n\nCoherence verification grows in complexity with the number of\ncores. PROPOSAL: a combination of formal verification of the\ndirectory protocol state machine, simulation-based testing, and\nrandom instruction generation is used. The verification\nstrategy is described in the verification document.\n\n## 14. Verification\n\n### 14.1 Cache Coherence\n\nPROPOSAL: The directory protocol state machine is verified using\na model checker against the MESI invariants. The invariants\ninclude:\n\n- Single writer per line.\n- No read of a line in the Invalid state.\n- Correct invalidation on a write to a shared line.\n\n### 14.2 Memory Consistency\n\nPROPOSAL: A RISC-V memory consistency model litmus test suite\nis run against the system, using both RTL simulation and a\nsoftware reference model.\n\n### 14.3 Performance and Stress Testing\n\nPROPOSAL: Synthetic traffic patterns, including uniform random,\nhotspot, and adversarial patterns, are used to verify that the\nmemory system behaves correctly under stress. Realistic\nworkloads, including SPEC CPU 2017, PARSEC, and selected\nGROMACS and LAMMPS runs, are used for performance validation.\n\nASSUMPTION: The listed benchmarks are representative of the\ntarget workloads. This assumption should be re-validated by the\nperformance team.\n\n## 15. Interaction With Other Subsystems\n\n### 15.1 Pipeline\n\nThe L1 cache is in the core pipeline. The L1 hit latency is\nvisible in the branch resolution and the load-to-use latency.\nPROPOSAL: The L1 hit latency is 4 cycles for a load-to-use on a\nhit. A miss to the L2 adds between 12 and 20 cycles, depending\non the L2 hit and the core-to-L2 routing.\n\n### 15.2 Cache Hierarchy\n\nThe L1, L2, and L3 form a hierarchy. The L2 is inclusive of the\nL1. The L3 is the coherence point. The L3 is non-inclusive of\nthe L2 by default.\n\n### 15.3 Memory System\n\nThe DRAM controllers are described in Section 9. The on-die\ninterconnect carries both coherence and data traffic.\n\n### 15.4 Interconnect\n\nThe on-die interconnect is described in Section 8. Coherence\nand data traffic share the same physical links.\n\nPROPOSAL: Virtual channels separate the request, response, and\ndata traffic. The number of virtual channels per link is 4.\n\n### 15.5 Coherence\n\nCoherence is described in Section 7. The protocol class is\ndistributed directory MESI with distributed shared L3.\n\n### 15.6 Interrupts\n\nInterrupts are described in Section 11. The PLIC is on the\ncoherence domain.\n\n### 15.7 Operating System\n\nThe operating system interactions are described in Section 10.\nThe OS is Linux for RISC-V. The firmware is OpenSBI.\n\n### 15.8 Performance\n\nThe performance analysis is described in Section 13 and is\ndetailed in a separate performance document. Key memory-system\nperformance metrics include:\n\n- L1 hit rate\n- L2 hit rate\n- L3 hit rate\n- DRAM bandwidth utilization\n- Average and tail DRAM latency\n- Interconnect average and tail latency\n\n## 16. Recommendations\n\n1. Implement a directory-based MESI coherence protocol with a\n distributed shared L3. Snoopy protocols are not viable at\n 128 cores.\n2. Use a 2D mesh on-die interconnect. The 2D mesh is sufficient\n for the projected traffic and is simpler to verify than a\n torus.\n3. Implement Zicbom and Zicboz cache management operations.\n These are required for Linux on RISC-V with mainline\n kernels.\n4. Expose each tile as a NUMA node to the operating system. This\n allows the kernel scheduler to make topology-aware decisions.\n5. Defer the final decision on the L3 size, the number of DRAM\n channels, and the directory organization to the\n implementation phase, after the sensitivity studies listed\n in Section 12 are completed.\n6. Defer the implementation of Zicbop until a workload study\n shows a clear benefit.\n\n## 17. References\n\nThe following standards and references are cited in this\ndocument:\n\n- RISC-V Unprivileged ISA Specification, Volume I, Document\n Version 20191213. Section on the Weak Memory Model and the\n Cache Management Operations extensions.\n- RISC-V Privileged ISA Specification, Volume II, Document\n Version 20211203. Sections on `satp`, SFENCE.VMA, PMP, and\n interrupt handling.\n- RISC-V Profiles, Document Version 20230914. Application and\n platform profile requirements.\n- RISC-V IOMMU Specification, draft as of the knowledge cutoff.\n- JEDEC DDR5 Standard, JESD79-5.\n- JEDEC LPDDR5 Standard, JESD209-5A.\n- Hennessy and Patterson, \"Computer Architecture: A Quantitative\n Approach\", 6th edition. Sections on cache coherence and\n directory protocols.\n- Culler, Singh, and Gupta, \"Parallel Computer Architecture: A\n Hardware/Software Approach\". Chapter on scalable cache\n coherence.\n- M. M. K. Martin, M. D. Hill, and D. A. Wood, \"Token\n Coherence\", University of Wisconsin Computer Sciences\n Technical Report 1621, 2003. Discusses a directory-based\n protocol with explicit invalidation tokens.\n- P. Kongetira, K. Aingaran, and K. Olukotun, \"Niagara: A\n 32-Way Multithreaded SPARC Processor\", IEEE Micro, March 2005.\n Discusses a directory-coherent CMP design.\n- R. Azimi et al., \"Integration Challenges and Tradeoffs for\n Tiled Many-Core Processors\", Proceedings of the 22nd\n International Conference on Computer Design, 2004.\n- P. Hammarlund et al., \"Haswell: The Fourth-Generation Intel\n Core Processor\", IEEE Micro, March 2014. Discusses a\n ring-based interconnect in a tiled CMP.\n- J. L. Henning, \"SPEC CPU2006 Benchmark Descriptions\", ACM\n SIGARCH Computer Architecture News, September 2006.\n- OpenSBI documentation, https://github.com/riscv-software-src/opensbi.\n\n## 18. Document History\n\n- Draft 0.1: initial outline.\n- Draft 0.2: populated physical memory map, coherence protocol\n class, and L1/L2/L3 sizing as proposals.\n- Draft 0.3: added open questions and recommendations.\n```","refusal":null,"reasoning":null}}],"usage":{"prompt_tokens":663,"completion_tokens":7174,"total_tokens":7837,"cost":0,"is_byok":false,"prompt_tokens_details":{"cached_tokens":135,"cache_write_tokens":0,"audio_tokens":0,"video_tokens":0},"cost_details":{"upstream_inference_cost":0,"upstream_inference_prompt_cost":0,"upstream_inference_completions_cost":0},"completion_tokens_details":{"reasoning_tokens":0,"image_tokens":0,"audio_tokens":0}}} |