mirror of
https://github.com/allexanderbergmns/xh1-research.git
synced 2026-08-26 19:47:01 +00:00
239 lines
22 KiB
JSON
239 lines
22 KiB
JSON
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
{"id":"gen-1787750367-KCLH3O5H7MT2UExQcqUI","object":"chat.completion","created":1787750367,"model":"minimax/minimax-m3:free","provider":"GMICloud","system_fingerprint":null,"service_tier":null,"choices":[{"index":0,"logprobs":null,"finish_reason":"stop","native_finish_reason":"stop","message":{"role":"assistant","content":"# Memory Bandwidth\n\n## Scope\n\nThis document investigates memory bandwidth requirements, limits, and\nimplementation strategies for the XH-1 processor, a proposed 128-core\nRISC-V design. The discussion covers off-chip DRAM bandwidth,\non-chip interconnect bandwidth, cache hierarchy bandwidth, and the\ninteractions between bandwidth, coherence traffic, and the memory\nsystem.\n\nThis document is part of the XH-1 research series (area 05-memory).\nStatements about XH-1 implementation are proposals unless explicitly\nmarked as decided.\n\n---\n\n## 1. Problem Statement\n\nA 128-core processor generates substantially more memory traffic than\nsingle-core or small-count-core designs. Even if each core issues a\nmodest number of outstanding memory requests, the aggregate demand can\nquickly exceed the bandwidth available from a realistic DRAM subsystem.\n\nThe XH-1 design must therefore answer several related questions:\n\n- How much DRAM bandwidth is physically available on the target package?\n- How much of that bandwidth is consumed by coherence traffic?\n- How much remains for useful application work?\n- How is that remaining bandwidth shared between cores without\n starvation or severe unfairness?\n- Where do bandwidth bottlenecks arise in the on-chip interconnect,\n cache hierarchy, and memory controllers?\n\nThe remainder of this document analyses each of these questions.\n\n---\n\n## 2. Background\n\n### 2.1 DRAM Bandwidth Fundamentals\n\nModern DDR systems provide bandwidth through wide channels and high\ntransfer rates:\n\n- DDR5 devices use 32-bit sub-channels per module rank, with\n transfer rates up to 8400 MT/s in JEDEC-published modules.\n- A single DDR5 channel (64 bits) at 6400 MT/s delivers approximately\n 51.2 GB/s peak; a dual-channel configuration roughly doubles this.\n- HBM3 stacks deliver substantially higher bandwidth per package\n (around 819 GB/s for a 6-high stack per HBM3 specification).\n- HBM3E extends this further, with published per-stack figures in the\n 1 TB/s range for selected devices.\n\nThe XH-1 design has not fixed a memory technology. Both DDR5 and\nHBM3-class options are considered below as proposals.\n\n### 2.2 RISC-V Memory Model\n\nRISC-V defines a weak memory ordering model, the RISC-V Weak Memory\nOrdering (RVWMO), in the architecture specification. RVWMO permits a\nwide range of reorderings and is similar in spirit to other modern\nweakly ordered architectures.\n\nKey consequences for bandwidth analysis:\n\n- Loads may bypass earlier stores under default rules when addresses\n differ.\n- Coherence is required only at a single total order per address\n (the \"coherence order\" axiom).\n- Fence instructions (`fence`, `fence.i`, `fence.vma`) are required\n to constrain ordering where software demands it.\n\nA weakly ordered model does not directly change bandwidth demand but\ninfluences how much buffering, speculation, and replay logic is\nrequired to expose bandwidth to software.\n\n### 2.3 Bandwidth vs. Latency\n\nBandwidth and latency are distinct metrics. A system can have ample\naggregate DRAM bandwidth but very high per-request latency, in which\ncase small working sets run well but large working sets suffer.\nA 128-core system is especially sensitive to the latency component\nbecause aggregate concurrency exposes the latency of every shared\nresource, including DRAM banks, the on-chip interconnect, and the\ndirectory or snoop filter.\n\n---\n\n## 3. Bandwidth Demand Estimation\n\n### 3.1 Per-Core Demand\n\nA modern out-of-order core can issue multiple outstanding cache misses.\nThe exact number depends on the pipeline width, miss status holding\nregister (MSHR) count, and prefetcher effectiveness, but published\nindustrial designs commonly support 10-32 outstanding L1 misses per\ncore.\n\nIf each outstanding request ultimately consumes 64 bytes from DRAM,\nand a core sustains 10 outstanding misses, the per-core demand is:\n\n 10 misses * 64 B / core / (memory round-trip latency)\n\nFor a 200 ns round-trip DRAM latency, this is roughly:\n\n 10 * 64 B / 200 ns = 3.2 GB/s per core\n\nIf all 128 cores simultaneously sustain this rate, the aggregate is\napproximately 410 GB/s, which already exceeds what a single DDR5\nchannel can deliver.\n\nThis calculation is intentionally rough. It is intended to motivate\nthe rest of the document, not to set a precise requirement.\n\n### 3.2 Realistic Demand\n\nIn practice, applications rarely sustain peak miss rates on every core\nsimultaneously. Published studies of throughput-oriented server\nworkloads report average L2 MPKI (misses per thousand instructions)\nin the range of 5-30 for memory-bound kernels and 1-5 for cache-friendly\nworkloads. These values vary significantly across workloads.\n\nMultiplying through, the worst-case aggregate bandwidth demand of a\n128-core system can plausibly reach hundreds of GB/s. The XH-1 design\nmust therefore plan for a memory subsystem substantially more capable\nthan a single DDR5 channel.\n\n### 3.3 Coherence Traffic Overhead\n\nCache coherence protocols add traffic on top of application reads and\nwrites. The magnitude depends on the protocol (see Section 6) and the\nsharing pattern of the workload. A directory-based protocol typically\ngenerates at least one directory lookup per coherence transaction and\npotentially one or more forward/intervention messages.\n\nA conservative estimate is that coherence and protocol overhead add\n10-30% to the raw application demand, but this can rise significantly\nfor workloads with intensive cross-core sharing.\n\n---\n\n## 4. On-Chip Bandwidth\n\nDRAM bandwidth is only one part of the system. The on-chip\ninterconnect and cache hierarchy must also carry:\n\n- Coherence messages between cores and directories.\n- Cache line fills from memory controllers to L2/L3 slices.\n- Write-backs from L1 to L2.\n- Snoop responses and invalidation acknowledgements.\n- Interrupt and configuration traffic (typically negligible).\n\nFor a 128-core mesh-class interconnect, published industrial designs\nindicate bisection bandwidth targets in the multi-TB/s range. Whether\nXH-1 reaches this range depends on link width, clock frequency, and\ntopology choices that are not yet decided.\n\n---\n\n## 5. Interconnect Topology and Bandwidth\n\nThe interconnect determines how effectively aggregate on-chip\nbandwidth is delivered to the DRAM controllers. Three common\ntopologies for many-core designs are:\n\n1. Crossbar\n2. Ring\n3. Mesh\n4. Concentrated mesh (cmesh)\n\n| Topology | Typical use | Bandwidth scaling | Notes |\n|----------|-------------|-------------------|-------|\n| Crossbar | small core counts | linear in core count | poor scalability beyond ~16 cores |\n| Ring | mid counts | limited by ring hops | simple but latency grows |\n| Mesh | mid-to-high counts | roughly linear with links | common in tiled designs |\n| Cmesh | high counts | similar to mesh, fewer routers | reduces router count |\n\nA crossbar at 128 cores is generally impractical due to area and\nwire count. A ring at 128 cores suffers from long hop counts and\nlimited bisection bandwidth. A mesh or concentrated mesh is the\nmost commonly chosen topology in published academic and industrial\n128-core class designs.\n\nThe XH-1 topology is not yet decided; both a 2D mesh and a 2D\nconcentrated mesh are candidate proposals.\n\n---\n\n## 6. Coherence Protocol Bandwidth\n\n### 6.1 Protocol Options\n\nTwo main families of coherence protocol exist:\n\n- Snooping protocols (broadcast-based).\n- Directory protocols (tracking-based).\n\nAt 128 cores, broadcast-based snooping is generally considered\nimpractical because every coherence transaction must reach every\ncore, consuming O(N) bandwidth per request. A directory protocol\nincurs O(1) directory lookups per request in the common case but\nrequires directory storage and handles sharing patterns differently.\n\nA directory is therefore the expected XH-1 proposal, consistent with\npublished 64-128 core research prototypes.\n\n### 6.2 Directory Organization\n\nDirectories can be implemented in several ways:\n\n- Full map directory: tracks sharers exactly. Storage cost is O(N)\n per line; prohibitive at 128 cores without compression or\n hierarchical schemes.\n- Limited pointer directory: tracks a fixed number of sharers per\n line (for example 4-6 sharers). When the count is exceeded, a\n fallback representation (often a \"broadcast\" or \"coarse vector\"\n marker) is used.\n- Coarse vector directory: groups cores into regions (e.g. 8-core\n tiles) and tracks presence at region granularity. Lower storage\n cost but less precise; can produce additional coherence traffic.\n- Sparse directory: similar to coarse vector but with different\n trade-offs.\n\nA limited pointer directory is the most commonly used scheme in\npublished industrial designs in this core-count range. The XH-1\ndesign is expected to use a limited pointer directory, possibly\naugmented with a coarse region fallback, but this is a proposal\nrather than a decision.\n\n### 6.3 Coherence Bandwidth Cost\n\nFor a limited pointer directory with P pointers, the common case\nis one directory lookup per coherence transaction. Bandwidth cost\nscales with the number of cores only when the directory overflows\ninto broadcast mode.\n\nPublished research indicates that overflow frequency is workload\ndependent. For read-mostly shared data, pointer counts typically\nstay small. For write-shared data, a limited pointer directory\ncan overflow frequently, increasing coherence traffic.\n\n---\n\n## 7. Cache Hierarchy Bandwidth\n\nThe cache hierarchy must sustain:\n\n- L1 to L2 traffic (fills, write-backs, invalidations).\n- L2 to L3 traffic.\n- L3 to DRAM traffic.\n- Snoop and probe traffic at every level.\n\nEach cache level adds banked storage and multiple read/write ports to\ndeliver bandwidth. The area and energy cost of these ports grows\nnonlinearly with port count.\n\nFor XH-1, a three-level hierarchy is the expected proposal:\n\n- Private L1 per core.\n- Private L2 per core (or per core pair) with a coherence controller.\n- Shared L3 distributed across tiles.\n\nThe exact partition of L2 vs. L3 capacity, associativity, and\nbandwidth is a design question left open in this document.\n\n---\n\n## 8. Bandwidth Allocation and QoS\n\nWhen 128 cores compete for limited DRAM bandwidth, some form of\nallocation policy is required. Options include:\n\n- Strict priority ordering.\n- Weighted fair queuing.\n- Per-core bandwidth caps.\n- Per-class (e.g. memory-mapped I/O vs. CPU) priorities.\n\nWithout allocation policy, a single core running a streaming workload\ncan saturate the memory subsystem and starve other cores. This is a\ndocumented failure mode in many-core research literature.\n\nRISC-V itself does not define memory bandwidth allocation semantics;\nthis is an XH-1 implementation choice. It is expected that the XH-1\nmemory controllers expose some form of bandwidth partitioning, but\nthis is a proposal rather than a decision.\n\n---\n\n## 9. Memory Controller Architecture\n\nThe memory controller is the point at which on-chip bandwidth meets\noff-chip DRAM. Key design choices include:\n\n- Number of independent channels.\n- Channel-to-core mapping.\n- Scheduling policy (FR-FCFS, parallelism-aware batch scheduling,\n etc.).\n- Address mapping policy.\n- Refresh and power management handling.\n\nFor DDR5-class memory, a 4-8 channel configuration is a reasonable\nstarting point for a 128-core design, but this is workload dependent\nand is presented as a proposal.\n\nFor HBM3-class memory, the channel count is fixed by the stack\nconfiguration; bandwidth is delivered through many independent\npseudo-channels per stack.\n\n---\n\n## 10. Quantitative Summary\n\nThe following table summarises representative bandwidth figures from\npublished sources. These are not XH-1 targets; they are external\nreference points.\n\n| Technology / Configuration | Peak Bandwidth | Source basis |\n|-----------------------------|----------------|--------------|\n| DDR5 single channel, 6400 MT/s | ~51.2 GB/s | JEDEC DDR5 spec |\n| DDR5 dual channel, 6400 MT/s | ~102.4 GB/s | JEDEC DDR5 spec |\n| HBM3 6-high stack | ~819 GB/s | JEDEC HBM3 spec |\n| HBM3E 8-high stack | ~1 TB/s class | vendor published figures |\n| 128-core aggregate theoretical (10 outstanding misses/core) | ~410 GB/s | rough calculation, Section 3.1 |\n\n---\n\n## 11. Advantages and Disadvantages of Candidate Memory Systems\n\n| Option | Advantages | Disadvantages |\n|--------|------------|---------------|\n| DDR5 multi-channel | Standard, widely available, lower cost per GB | Lower peak bandwidth per channel, more channels increase board complexity |\n| HBM3 / HBM3E | Very high bandwidth, high channel count per stack | Higher cost, 2.5D/3D packaging required, lower capacity per stack |\n| Hybrid (HBM + DDR) | Combine high bandwidth with high capacity | Complex memory hierarchy, two coherence domains |\n\nThe XH-1 memory technology choice is not yet decided.\n\n---\n\n## 12. Scalability Problems at 128 Cores\n\nThe following scalability problems are well documented in the\nmany-core literature and are directly relevant to XH-1:\n\n- **DRAM bandwidth saturation**: aggregate demand can exceed\n realistic DRAM bandwidth.\n- **Directory overflow**: limited pointer directories overflow\n more often as core count rises, increasing coherence traffic.\n- **Coherence traffic growth**: as more cores share data,\n coherence messages grow superlinearly in some sharing patterns.\n- **Interconnect bisection bottlenecks**: insufficient bisection\n bandwidth creates hotspots at memory controllers.\n- **Quality-of-service degradation**: without allocation policy,\n fairness collapses as core count rises.\n- **Verification cost**: coherence state space grows roughly\n combinatorially with core count.\n\nThese are not unique to XH-1 but apply to any 128-core coherent\ndesign.\n\n---\n\n## 13. Interactions With Other Subsystems\n\n### 13.1 Pipeline\n\nA deeper pipeline can support higher clock frequencies but increases\nthe penalty of cache misses. Bandwidth-bound workloads become\nsensitive to the round-trip latency the pipeline exposes. The\npipeline must include enough MSHRs to expose memory-level parallelism.\n\n### 13.2 Cache Hierarchy\n\nCache capacity and associativity directly affect miss rate and hence\nDRAM bandwidth demand. Larger caches reduce demand at the cost of\narea and access time. Bandwidth and capacity are coupled; both\nmust be planned together.\n\n### 13.3 Memory System\n\nThe memory controller scheduling policy, address mapping, and\nrefresh handling all affect effective bandwidth. Closed-page vs.\nopen-page policies, bank parallelism, and request batching\nmaterially change delivered bandwidth.\n\n### 13.4 Interconnect\n\nThe interconnect must deliver at least as much aggregate bandwidth\nas the DRAM subsystem, plus coherence overhead. Otherwise the\ninterconnect becomes the bottleneck and DRAM is underutilised.\n\n### 13.5 Coherence\n\nCoherence traffic is a first-class component of bandwidth demand.\nProtocol choice and directory organisation directly determine the\nshare of bandwidth consumed by coherence.\n\n### 13.6 Interrupts\n\nInterrupts are typically not bandwidth-limited, but interrupt\ndelivery traffic and any interrupt-cause queues add minor traffic\nto the system fabric.\n\n### 13.7 Operating System\n\nThe OS schedules processes across cores, controls page placement,\nand may implement memory allocation policies that affect locality.\nThe OS does not directly set DRAM bandwidth but its page colour\nand NUMA-style placement policies interact with the memory\ncontrollers.\n\n### 13.8 Verification\n\nCoherence protocols at 128 cores have a large state space.\nVerification complexity grows roughly with the product of\nper-core state and the number of cores. Formal methods,\nrandomised testing, and traffic pattern fuzzing are all relevant.\n\n### 13.9 Performance\n\nBandwidth is one of the dominant performance constraints for\nmemory-bound workloads. Performance counters for DRAM channel\nutilisation, interconnect link utilisation, and coherence message\ncounts are essential for tuning.\n\n---\n\n## 14. RISC-V Architectural Requirements\n\nThe following RISC-V architectural elements are relevant to bandwidth\nanalysis:\n\n- The RISC-V ISA specifies `fence`, `fence.i`, and `fence.vma`\n instructions that order memory operations.\n- The privileged architecture defines physical memory attributes\n (PMA) and, in designs with translation, page-based attributes\n (PBMT, NAPOT, etc.) that can affect memory behaviour.\n- The RISC-V coherency management extensions (for example, the\n proposed H-extension in the vector and hypervisor contexts, and\n the cache-management operations in the RISC-V Profiles and\n IOMMU-related specifications) provide instructions such as\n `cbo.clean`, `cbo.flush`, and `cbo.inval` that interact with\n cache bandwidth.\n\nRISC-V itself does not mandate a specific coherence protocol or\nmemory bandwidth management policy. These are XH-1 implementation\nchoices.\n\n---\n\n## 15. Open Questions and Unresolved Design Decisions\n\nThe following items are not yet decided and require further research\nor design effort:\n\n- Choice of DRAM technology (DDR5 multi-channel vs. HBM3-class).\n- Number and placement of memory controllers.\n- Interconnect topology (mesh vs. concentrated mesh, dimensional\n order, link width).\n- Coherence protocol variant (e.g. MESI vs. MOESI and the\n directory structure).\n- Bandwidth allocation policy and per-core fairness guarantees.\n- L2 vs. L3 capacity split and per-tile bandwidth targets.\n- Quality-of-service hooks exposed to the operating system.\n- Verification strategy for the coherence and interconnect at\n 128 cores.\n\nThese are recorded as open questions; no recommendation is made\nwhere the available evidence does not strongly favour a particular\nchoice.\n\n---\n\n## 16. Recommendations\n\nRecommendations are made only where published evidence and engineering\npractice converge.\n\n- **Use a directory-based coherence protocol**: published 64-128 core\n research and industrial designs consistently use directories; a\n snoop-based approach is not credible at 128 cores.\n- **Plan for multi-channel DRAM**: a single DDR5 channel is unlikely\n to meet aggregate demand; at least 2-4 channels are a reasonable\n starting point if DDR5 is used, with HBM as an alternative if\n higher bandwidth is required.\n- **Provide per-class quality-of-service in the memory controllers**:\n without allocation, fairness is not achievable in a 128-core\n system.\n- **Instrument the design with bandwidth and coherence performance\n counters**: empirical tuning is unavoidable.\n- **Verify coherence state space with a combination of formal\n methods and randomised testing**: full state space enumeration\n is not feasible at 128 cores.\n\nThese are recommendations, not decisions; they are subject to\nrevision as the XH-1 project progresses.\n\n---\n\n## 17. References\n\n- RISC-V International, *The RISC-V Instruction Set Manual, Volume I:\n Unprivileged ISA*, current ratified version. RVWMO is described\n in Volume I.\n- RISC-V International, *The RISC-V Instruction Set Manual, Volume II:\n Privileged Architecture*, current ratified version.\n- JEDEC, *DDR5 SDRAM Standard* (JESD79-5), published release.\n- JEDEC, *High Bandwidth Memory DRAM (HBM3)*, JESD238A.\n- JEDEC, *High Bandwidth Memory DRAM (HBM3E)*, JESD238B-1 (where\n applicable).\n- Hennessy and Patterson, *Computer Architecture: A Quantitative\n Approach*, recent edition, for general memory system modelling\n and AMAT analysis.\n- Sorin, Hill, and Wood, *A Primer on Memory Consistency and Cache\n Coherence*, for coherence protocol analysis.\n- Martin, Hill, and Wood, and successor work, on many-core\n coherence verification and token coherence (where relevant to\n XH-1 verification work).\n\nWhere specific figures in this document depend on vendor or\nJEDEC-published numbers, the source is indicated in-line. Where\nfigures are derived from calculations, this is stated explicitly.\n\n---\n\n## 18. Document Status\n\nThis document is a research proposal and analysis. It is not a\ndescription of an implemented design. Items marked as \"proposal\"\nor \"expected\" have not been formally decided within the XH-1\nproject. Items marked as \"open\" represent unresolved design\nquestions. The document should be re-reviewed as the XH-1\narchitecture specification stabilises.","refusal":null,"reasoning":null}}],"usage":{"prompt_tokens":663,"completion_tokens":4358,"total_tokens":5021,"cost":0,"is_byok":false,"prompt_tokens_details":{"cached_tokens":135,"cache_write_tokens":0,"audio_tokens":0,"video_tokens":0},"cost_details":{"upstream_inference_cost":0,"upstream_inference_prompt_cost":0,"upstream_inference_completions_cost":0},"completion_tokens_details":{"reasoning_tokens":0,"image_tokens":0,"audio_tokens":0}}} |