diff --git a/research/.xh1/logs/harness.log b/research/.xh1/logs/harness.log new file mode 100644 index 0000000..44e1f01 --- /dev/null +++ b/research/.xh1/logs/harness.log @@ -0,0 +1,117 @@ +[2026-08-25T17:58:32Z] Starting research: research/03-core-design/mul-div-unit.md +[2026-08-25T17:58:32Z] Running researcher. +[2026-08-25T17:58:32Z] API attempt 1 failed. +[2026-08-25T17:58:37Z] API attempt 2 failed. +[2026-08-25T17:59:55Z] Starting research: research/03-core-design/mul-div-unit.md +[2026-08-25T17:59:55Z] Running researcher. +[2026-08-25T17:59:56Z] API attempt 1 failed. +[2026-08-25T17:59:57Z] API attempt 2 failed. +[2026-08-25T18:00:09Z] Starting research: research/03-core-design/mul-div-unit.md +[2026-08-25T18:00:09Z] Running researcher. +[2026-08-25T18:00:10Z] API attempt 1 failed. +[2026-08-25T18:00:11Z] API attempt 2 failed. +[2026-08-25T18:00:15Z] API attempt 3 failed. +[2026-08-25T18:00:56Z] Starting research: research/03-core-design/mul-div-unit.md +[2026-08-25T18:00:56Z] Running researcher. +[2026-08-25T18:00:57Z] API attempt 1 failed. +[2026-08-25T18:00:58Z] API attempt 2 failed. +[2026-08-25T18:01:02Z] API attempt 3 failed. +[2026-08-25T18:01:11Z] Research agent failed. +[2026-08-25T18:01:38Z] Starting research: research/03-core-design/mul-div-unit.md +[2026-08-25T18:01:38Z] Running researcher. +[2026-08-25T18:01:38Z] API attempt 1 failed. +[2026-08-25T18:01:57Z] Starting research: research/03-core-design/mul-div-unit.md +[2026-08-25T18:01:57Z] Running researcher. +[2026-08-25T18:02:54Z] Starting research: research/03-core-design/mul-div-unit.md +[2026-08-25T18:02:54Z] Running researcher. +[2026-08-25T18:02:54Z] API attempt 1 failed. +[2026-08-25T18:02:55Z] API attempt 2 failed. +[2026-08-25T18:03:00Z] Starting research: research/03-core-design/mul-div-unit.md +[2026-08-25T18:03:00Z] Running researcher. +[2026-08-25T18:04:04Z] Research completed. +[2026-08-25T18:04:04Z] Running reviewer. +[2026-08-25T18:04:28Z] Research rejected by reviewer. +[2026-08-25T18:08:09Z] Started run: 20260825T180809Z +[2026-08-25T18:08:09Z] ================================================== +[2026-08-25T18:08:09Z] Researching: research/03-core-design/mul-div-unit.md +[2026-08-25T18:08:09Z] ================================================== +[2026-08-25T18:08:09Z] Research round 1/3 +[2026-08-25T18:08:09Z] Running researcher. +[2026-08-25T18:09:00Z] Started run: 20260825T180900Z +[2026-08-25T18:09:00Z] ================================================== +[2026-08-25T18:09:00Z] Researching: research/03-core-design/mul-div-unit.md +[2026-08-25T18:09:00Z] ================================================== +[2026-08-25T18:09:00Z] Research round 1/3 +[2026-08-25T18:09:00Z] Running researcher. +[2026-08-25T18:09:00Z] API request: model=minimax/minimax-m3:free attempt=1 +[2026-08-25T18:09:56Z] API response received: 17248 bytes +[2026-08-25T18:09:56Z] Running reviewer. +[2026-08-25T18:09:56Z] API request: model=z-ai/glm-5.2:free attempt=1 +[2026-08-25T18:09:57Z] API request failed. +[2026-08-25T18:09:58Z] API request: model=z-ai/glm-5.2:free attempt=2 +[2026-08-25T18:09:58Z] API request failed. +[2026-08-25T18:10:02Z] API request: model=z-ai/glm-5.2:free attempt=3 +[2026-08-25T18:10:03Z] API request failed. +[2026-08-25T18:11:53Z] Started run: 20260825T181153Z +[2026-08-25T18:11:53Z] ================================================== +[2026-08-25T18:11:53Z] Researching: research/03-core-design/mul-div-unit.md +[2026-08-25T18:11:53Z] ================================================== +[2026-08-25T18:11:53Z] Research round 1/3 +[2026-08-25T18:11:53Z] Running researcher. +[2026-08-25T18:11:53Z] API request: model=minimax/minimax-m3:free attempt=1 +[2026-08-25T18:13:01Z] API response received: 21756 bytes +[2026-08-25T18:13:01Z] Running reviewer. +[2026-08-25T18:13:01Z] API request: model=google/gemma-4-31b-it:free attempt=1 +[2026-08-25T18:13:01Z] API request failed. +[2026-08-25T18:13:02Z] API request: model=google/gemma-4-31b-it:free attempt=2 +[2026-08-25T18:13:02Z] API request failed. +[2026-08-25T18:13:06Z] API request: model=google/gemma-4-31b-it:free attempt=3 +[2026-08-25T18:13:07Z] API request failed. +[2026-08-25T18:14:02Z] Started run: 20260825T181402Z +[2026-08-25T18:14:02Z] ================================================== +[2026-08-25T18:14:02Z] Researching: research/03-core-design/mul-div-unit.md +[2026-08-25T18:14:02Z] ================================================== +[2026-08-25T18:14:02Z] Research round 1/3 +[2026-08-25T18:14:02Z] Running researcher. +[2026-08-25T18:14:02Z] API request: model=minimax/minimax-m3:free attempt=1 +[2026-08-25T18:15:25Z] API response received: 26879 bytes +[2026-08-25T18:15:25Z] Running reviewer. +[2026-08-25T18:15:25Z] API request: model=poolside/laguna-s-2.1:free attempt=1 +[2026-08-25T18:15:26Z] API request failed. +[2026-08-25T18:15:27Z] API request: model=poolside/laguna-s-2.1:free attempt=2 +[2026-08-25T18:15:27Z] API request failed. +[2026-08-25T18:15:31Z] API request: model=poolside/laguna-s-2.1:free attempt=3 +[2026-08-25T18:15:31Z] API request failed. +[2026-08-25T18:16:09Z] Started run: 20260825T181609Z +[2026-08-25T18:16:09Z] ================================================== +[2026-08-25T18:16:09Z] Researching: research/03-core-design/mul-div-unit.md +[2026-08-25T18:16:09Z] ================================================== +[2026-08-25T18:16:09Z] Research round 1/3 +[2026-08-25T18:16:09Z] Running researcher. +[2026-08-25T18:16:09Z] API request: model=minimax/minimax-m3:free attempt=1 +[2026-08-25T18:17:05Z] API response received: 18408 bytes +[2026-08-25T18:17:05Z] Running reviewer. +[2026-08-25T18:17:05Z] API request: model=minimax/minimax-m3:free attempt=1 +[2026-08-25T18:17:41Z] API response received: 12530 bytes +[2026-08-25T18:17:41Z] Research rejected by reviewer. +[2026-08-25T18:17:41Z] Preparing revision round 2. +[2026-08-25T18:17:46Z] Research round 2/3 +[2026-08-25T18:17:46Z] Running revision agent. +[2026-08-25T18:17:46Z] API request: model=minimax/minimax-m3:free attempt=1 +[2026-08-25T18:18:49Z] API response received: 30663 bytes +[2026-08-25T18:18:49Z] Running reviewer. +[2026-08-25T18:18:49Z] API request: model=minimax/minimax-m3:free attempt=1 +[2026-08-25T18:19:38Z] API response received: 17758 bytes +[2026-08-25T18:19:38Z] Research rejected by reviewer. +[2026-08-25T18:19:38Z] Preparing revision round 3. +[2026-08-25T18:19:43Z] Research round 3/3 +[2026-08-25T18:19:43Z] Running revision agent. +[2026-08-25T18:19:43Z] API request: model=minimax/minimax-m3:free attempt=1 +[2026-08-25T18:20:47Z] API response received: 38527 bytes +[2026-08-25T18:20:47Z] Running reviewer. +[2026-08-25T18:20:47Z] API request: model=minimax/minimax-m3:free attempt=1 +[2026-08-25T18:21:21Z] API response received: 11009 bytes +[2026-08-25T18:21:21Z] Research rejected by reviewer. +[2026-08-25T18:21:21Z] Maximum research rounds reached. +[2026-08-25T18:21:21Z] Leaving original document unchanged. +[2026-08-25T18:21:21Z] Marking topic failed for this run. diff --git a/research/.xh1/responses/20260825T175832Z_mul-div-unit.md.json b/research/.xh1/responses/20260825T175832Z_mul-div-unit.md.json new file mode 100644 index 0000000..e69de29 diff --git a/research/.xh1/responses/20260825T175955Z_mul-div-unit.md.json b/research/.xh1/responses/20260825T175955Z_mul-div-unit.md.json new file mode 100644 index 0000000..e69de29 diff --git a/research/.xh1/responses/20260825T180009Z_mul-div-unit.md.json b/research/.xh1/responses/20260825T180009Z_mul-div-unit.md.json new file mode 100644 index 0000000..e69de29 diff --git a/research/.xh1/responses/20260825T180056Z_mul-div-unit.md.json b/research/.xh1/responses/20260825T180056Z_mul-div-unit.md.json new file mode 100644 index 0000000..e69de29 diff --git a/research/.xh1/responses/20260825T180138Z_mul-div-unit.md.json b/research/.xh1/responses/20260825T180138Z_mul-div-unit.md.json new file mode 100644 index 0000000..e69de29 diff --git a/research/.xh1/responses/20260825T180157Z_mul-div-unit.md.json b/research/.xh1/responses/20260825T180157Z_mul-div-unit.md.json new file mode 100644 index 0000000..90b8261 --- /dev/null +++ b/research/.xh1/responses/20260825T180157Z_mul-div-unit.md.json @@ -0,0 +1,190 @@ + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + diff --git a/research/.xh1/responses/20260825T180254Z_mul-div-unit.md.json b/research/.xh1/responses/20260825T180254Z_mul-div-unit.md.json new file mode 100644 index 0000000..e69de29 diff --git a/research/.xh1/responses/20260825T180300Z_mul-div-unit.md b/research/.xh1/responses/20260825T180300Z_mul-div-unit.md new file mode 100644 index 0000000..d701db3 --- /dev/null +++ b/research/.xh1/responses/20260825T180300Z_mul-div-unit.md @@ -0,0 +1,284 @@ +# MUL/DIV Unit + +## Status + +DRAFT — initial research document. No architectural decision has been made for XH-1. + +## Abstract + +This document investigates the design of the multiplication and division unit (MUL/DIV) for a single core inside the XH-1 128-core RISC-V processor. It surveys existing approaches for integer multiplication and division, identifies alternative implementations, and analyzes trade-offs in performance, area, power, latency, verification complexity, and scalability across 128 cores. The document focuses on the base integer extensions (RV64I/M) and explicitly defers floating-point and vector multiply/accumulate topics, which belong to separate units. + +## Research Question + +What is the most appropriate microarchitectural implementation of the MUL/DIV unit for a single XH-1 core, given that 128 identical cores will be instantiated on die, and given that the MUL/DIV unit must implement at minimum RV64M (MUL, MULH, MULHU, MULHSU, DIV, DIVU, REM, REMU)? + +Sub-questions: + +- Should multiplication be iterative (shift-and-add) or fully combinational (array/Wallace/Booth)? +- Should division use a restoring, non-restoring, SRT, or Newton–Raphson scheme? +- Should the unit be pipelined, multi-cycle, or variable-latency? +- How does the MUL/DIV unit interact with the surrounding pipeline (depth, bypass, in-order vs out-of-order issue)? +- How does the unit scale when replicated 128 times on die? +- Should fused MAC operations (MUL + ADD) or fused multiply-add be considered? + +## Background + +The RISC-V "M" extension specifies eight integer multiply/divide instructions (MUL, MULH, MULHSU, MULHU, DIV, DIVU, REM, REMU) on 64-bit values producing either 64-bit or 128-bit results. Division is defined to round toward zero and is required to complete even for overflow cases (e.g., `INT64_MIN / -1`), as specified in the RISC-V Unprivileged ISA. + +Key properties of the workload that influence the design: + +- **Result width asymmetry.** MUL produces a 128-bit result, but only the low 64 bits are written to `rd` for `MUL`. MULH-family instructions write the high 64 bits. The datapath therefore needs at least a 64×64→128 multiplier followed by a selector. +- **Sign handling.** Three signed/unsigned combinations (signed×signed, signed×unsigned, unsigned×unsigned) must be supported. Sign correction is required for MULH/MULHSU. +- **Division latency and throughput.** RISC-V does not require division to retire in a single cycle, but the ISA mandates deterministic behavior for overflow. +- **Rarity in many workloads.** Empirical studies (e.g., Hennessy & Patterson, *Computer Architecture: A Quantitative Approach*) report that integer divide and remainder instructions are uncommon (typically <1% of dynamic instructions), while multiplies are more frequent in HPC and crypto workloads. + +For XH-1, the unit sits on the execution path of every core. With 128 cores, area and per-core energy dominate; raw single-thread latency of division matters less than aggregate throughput, die-area cost, and ease of verification. + +## Existing Approaches + +### Multiplication + +1. **Iterative shift-and-add multiplier** + - One 64-bit adder reused across 64 cycles. + - Smallest area, lowest energy per multiplication, but very long latency. + +2. **Array multiplier (combinational)** + - 64×64 array of full adders producing a 128-bit result. + - Single-cycle result, large area, high fan-out, and a long critical path. + - Historically too slow for one cycle at high clock frequencies. + +3. **Wallace / Dadda tree** + - Tree of carry-save adders reducing partial products to two 128-bit vectors, then a final carry-propagate adder. + - Logarithmic depth; commonly used in high-performance cores. + - Larger area than array, but shorter critical path. + +4. **Booth-encoded Wallace / Dadda** + - Radix-4 or higher Booth recoding reduces the number of partial products by ~2×. + - Common in modern cores (e.g., reported in implementations of ARM and x86 multipliers). + +5. **Pipelined iterative multiplier** + - Splits the 64-cycle iterative multiplier into pipeline stages (commonly 2–4 stages). + - Used in many in-order RISC-V cores (e.g., the Rocket Chip generator's `MulDiv`). + +6. **Dedicated single-cycle fused multiply-add (FMA) for integers** + - Rare for integer-only M-extensions; usually belongs to the F/D extensions. + +### Division + +1. **Restoring division** + - Classical shift-subtract. One bit per cycle, 64 cycles for 64-bit operands. + - Simple, easy to verify. + +2. **Non-restoring division** + - Similar latency, but allows a single add/subtract per bit without explicit restore. + - Used in many textbook implementations. + +3. **SRT division** + - Radix-4 or higher; produces 2+ bits per cycle using a small redundant quotient-digit table. + - Significantly faster (16–32 cycles for 64-bit) at higher area cost. + - Common in high-performance OoO cores (e.g., POWER, Itanium, recent x86). + +4. **Newton–Raphson reciprocal + multiply** + - Iteratively refines an approximation of 1/d, then multiplies. + - Very high throughput once the reciprocal is available. + - Worst-case latency is higher than SRT; best for repeated divisions by the same divisor. + - Rare in integer pipelines due to initial latency. + +5. **Goldschmidt division** + - Similar to Newton–Raphson; iteratively scales numerator and denominator toward 1. + - Same usage profile as Newton–Raphson. + +6. **Lookup-table based constant dividers** + - For known small constant divisors, the compiler/runtime can replace DIV with a multiply-by-reciprocal. + - Microarchitectural implication: the DIV unit need not be heavily optimized if software frequently replaces division by constants. + +## Alternative Designs + +For the XH-1 MUL/DIV unit, four credible microarchitectural templates are considered. + +### Design A: Shared iterative multi-cycle unit (Rocket-style) + +- One 64-bit datapath reused for both MUL and DIV. +- MUL: 1 cycle/partial product (radix-2 Booth optional). +- DIV: 1 bit/cycle, restoring or non-restoring. +- MUL latency: ~33–35 cycles (radix-2 Booth) or ~64 cycles (plain shift-add). +- DIV latency: ~64 cycles. +- Throughput: 1 MUL or DIV per ~32 cycles (shared); MUL and DIV cannot execute concurrently. +- Plausible canonical reference: the `MulDiv` module in the BOOM/Rocket Chip generator (UC Berkeley). + +### Design B: Pipelined iterative multiplier + iterative divider + +- Multiplier: 2–4 stage pipelined radix-2/radix-4 iterative unit. +- Divider: separate 64-bit iterative datapath (non-restoring or SRT-radix-2). +- MUL throughput: 1 per cycle once pipeline is full. +- DIV throughput: 1 per ~32 cycles. +- Independent issue of MUL and DIV is possible. + +### Design C: Pipelined Wallace/Booth multiplier + SRT-radix-4 divider + +- Multiplier: 3-stage pipelined radix-4 Booth → Wallace tree → CPA. Produces full 128-bit result. +- Divider: radix-4 SRT producing 2 bits/cycle; ~17 cycles for 64-bit DIV. +- DIV/REM can be produced simultaneously since quotient digits are known. +- Area: significantly larger than Designs A and B. +- Latency: MUL ~3–4 cycles, DIV ~17–20 cycles. +- Used in many modern superscalar cores. + +### Design F: Fused integer MAC / FMA + +- Add an integer fused multiply-add returning low 64 bits (a*b)+c in one operation. +- This is non-standard for RV64M and would require custom opcodes or being staged behind a regular MUL+ADD sequence. +- Documented here for completeness, not recommended without strong workload evidence. + +## Comparison + +| Property | A: Iterative shared | B: Pipelined iterative MUL + iterative DIV | C: Wallace/Booth MUL + SRT-4 DIV | +|---|---|---|---| +| MUL latency | ~33–64 cycles | 3–5 cycles | 3–4 cycles | +| MUL throughput | 1 / 32 cycles | 1 / cycle | 1 / cycle | +| DIV latency | ~64 cycles | ~32 cycles | ~16–20 cycles | +| DIV throughput | 1 / 64 cycles | 1 / 32 cycles | 1 / 16 cycles | +| 64×64→128 datapath | Yes (shared) | Yes (MUL only) | Yes (Wallace) | +| MUL+DIV concurrency | No (shared) | Yes (separate datapaths) | Yes (separate datapaths) | +| Estimated relative area | 1.0× | ~1.5–2.0× | ~3.0–5.0× | +| Estimated critical path | Short | Short | Longest (CPA final stage) | +| Verification complexity | Low | Medium | High | +| Fits "small in-order" model | Excellent | Good | Marginal | + +ASSUMPTION: Area estimates above are rough order-of-magnitude relative numbers based on typical RISC-V implementations and the cited textbooks. They have not been measured for XH-1. + +## Advantages + +### Design A (iterative shared) + +- Smallest area per core, which directly reduces die cost across 128 cores. +- Lowest per-core dynamic energy for the rare case of an actual MUL/DIV. +- Easiest to verify formally (small state space, one datapath). +- Matches the "many small cores" scaling philosophy. +- Canonical reference: Rocket Chip `MulDiv`. + +### Design B (pipelined iterative MUL + iterative DIV) + +- MUL throughput is high enough to support HPC and crypto workloads where 64-bit multiplies are common. +- DIV remains simple. +- Area increase over A is bounded. +- Two independent datapaths simplify scheduling in the issue stage. + +### Design C (Wallace/Booth + SRT-4) + +- Best raw latency and throughput for both operations. +- Suitable for single-core-bound workloads or for cores that need to hide memory latency behind fast arithmetic. +- DIV+REM can be produced together with little extra hardware. + +## Disadvantages + +### Design A + +- DIV latency of 64 cycles is long; if the surrounding pipeline is short (e.g., 5–7 stages), the unit will dominate total execution time for any divide. +- Back-to-back MULs serialize. +- Under HPC or cryptography kernels, MUL throughput becomes a bottleneck. + +### Design B + +- More area than A. +- Pipelined MUL increases register pressure in the issue queue and requires more bypass paths in the surrounding execution stage. +- DIV still slow. + +### Design C + +- Largest area per core, replicated 128 times. +- Highest per-core power. +- Wallace tree and SRT have long critical paths that may limit clock frequency for the whole core. +- Verification complexity is significantly higher: partial-product reduction, Booth recoding, SRT quotient-digit selection tables, and divider corner cases (e.g., `INT64_MIN / -1`) all need separate coverage. +- Wall-clock design and verification cost may delay the whole project. + +## XH-1 Considerations + +PROPOSAL: For XH-1, an in-order core with 128 instances on die, the dominant design constraint is **per-core area, energy, and verification cost**, not single-thread peak performance. The MUL/DIV unit should therefore favor small, simple, well-trodden implementations. + +Specific implications for XH-1: + +- The 128-core factor means the MUL/DIV unit's area is multiplied by 128. Even a 2× area difference per core translates to a substantial absolute area delta. +- The energy of 128 MUL/DIV datapaths, even at low utilization, contributes to total socket power. +- A long-latency MUL/DIV unit is acceptable if the surrounding pipeline is deep enough or if it can be overlapped with other in-flight instructions in the same core. +- Single-cycle MUL would impose a critical path on the whole core; for a 128-core design, sustained high clock frequency across all cores is critical to total throughput. + +## 128-Core Scalability + +Scalability dimensions to consider: + +- **Wiring and layout.** A 128-core die has a complex interconnect. A small MUL/DIV unit is easier to place and route within each core tile. Designs with large irregular adder trees (Wallace/SRT) complicate physical design at high core counts. +- **Verification replication.** Bugs in the MUL/DIV unit, if present, propagate to 128 cores. A simpler, formally verifiable design (Design A) is safer for replication. +- **Yield.** Smaller per-core area improves yield and binning flexibility; large per-core area reduces the number of cores that fit on a reticle at the target process node. +- **Power delivery.** 128 simultaneous MUL/DIV operations are unlikely, but worst-case power events (e.g., SIMD-style vector MUL workloads scaled down to integer MUL) must be within the socket's power-delivery budget. +- **Frequency scaling.** A 128-core chip with modest per-core frequency but high aggregate throughput may benefit from a short critical path. A Wallace multiplier's critical path can limit fmax for the whole core. + +## Performance Considerations + +- MUL/DIV instructions are infrequent in general-purpose workloads (often <1% dynamic instructions) but can dominate kernels in cryptography (AES, ChaCha20, RSA), big-integer arithmetic (GMP-style libraries), and some HPC kernels. +- If XH-1 is intended for general-purpose server or desktop use, the MUL/DIV unit will rarely be on the critical path of a thread. +- If XH-1 targets HPC or cryptography, the MUL throughput becomes important. In this case, B or C should be reconsidered. +- Software can use compiler transformations to replace DIV by constants with multiply-by-reciprocal, reducing pressure on the DIV unit. + +## Implementation Considerations + +- **Sign handling.** The unit must correctly handle `MULH`, `MULHSU`, and `MULHU` as well as overflow cases of DIV (notably `INT64_MIN / -1`, which must produce `INT64_MIN` per RISC-V spec). +- **REM vs DIV.** Producing REM in parallel with DIV using the same datapath is standard in restoring/non-restoring designs; the unit should support issuing DIVU/REMU pairs in one operation. +- **Pipeline interface.** The unit must integrate with the core's issue, wakeup, and writeback stages. If in-order, the issue stage must stall in-order cores on multi-cycle MUL/DIV. If OoO, completion must wait for the unit's completion signal. +- **Bypassing.** Forwarding paths from the MUL/DIV pipeline registers to dependent instructions must be designed carefully to avoid structural hazards. +- **Early termination.** For DIV, the unit can terminate early when the remainder is zero, saving cycles. Implementation cost is low. + +## Verification Considerations + +- **Corner cases.** RV64M has well-defined corner cases: `INT64_MIN / -1`, division by zero, overflow in REM, sign interactions in MULH-family. +- **Directed + constrained-random.** A combination of directed tests for ISA corner cases and constrained-random for the rest is standard practice (e.g., as in the RISC-V architectural test framework, riscv-tests). +- **Formal verification.** A small iterative multiplier/divider (Design A) is amenable to formal proofs of correctness for a few-bit case and inductive scaling. A Wallace + SRT unit (Design C) is significantly harder to formally verify due to selector-table complexity. +- **Cross-core equivalence.** With 128 identical cores, regression in one core implies regression in all 128. A well-verified single-core design simplifies the chip-level verification effort. +- **Testbench reuse.** The RISC-V community maintains architectural compliance tests that should be run against the MUL/DIV unit regardless of the chosen design. + +## Recommendation + +PROPOSAL: Adopt a **Design B–leaning approach**: a small, simple, well-understood MUL/DIV unit similar in spirit to Rocket Chip's `MulDiv`, with the following characteristics: + +- A **radix-4 Booth-encoded iterative multiplier** (or radix-2 if radix-4 proves too complex for the area budget) producing 64 bits of result per ~16 cycles, sharing partial datapath with the divider if needed. +- A **non-restoring (or radix-2 SRT) divider** completing in ~32 cycles. +- MUL and DIV on the same datapath with **shared state** but capable of being interleaved at issue time. +- Optional microarchitectural relaxation: a separate tiny **fast-MUL path for 32×32→64 results** (the low half of MUL where both operands are sign- or zero-extended from 32 bits) to accelerate common cases. This adds minimal area. + +This recommendation is provisional and is the lightest-weight option that still keeps MUL throughput reasonable. It avoids the critical-path cost of Design C and the throughput limit of Design A, while remaining well within the verification budget of a 128-core project. + +RECOMMENDATION: If workload analysis (not yet performed) shows MUL-heavy HPC/cryptography use, escalate to a pipelined radix-4 Booth multiplier with a 2–3 cycle latency, keeping the iterative divider. If workload analysis shows almost no MUL/DIV usage, drop to a plain Design A. + +## Confidence + +- **Low–Medium** for any specific microarchitectural recommendation. The document is at an early stage; the recommendation will be revised after: + 1. Workload analysis (target use cases of XH-1). + 2. Synthesis of representative MUL/DIV units in the target technology. + 3. Frequency, area, and power target constraints. +- **High** that the iterative, shared-datapath approach (Design A or B) is the appropriate starting point for a 128-core, area-constrained, verification-constrained design. + +## Open Questions + +- What is the target frequency of XH-1 cores, and what is the critical-path budget for the MUL/DIV unit? +- What process node is targeted, and what is the per-core area budget? +- What is the intended workload mix (server, HPC, embedded, ML)? +- Is the core in-order or out-of-order? The MUL/DIV interface and latency tolerance depend strongly on this. +- Will the F extension (floating-point) be present in the same core, and if so, should integer MUL be reused inside an FMA datapath? +- Will the V extension (vector) be present? If so, scalar integer MUL may be lightly used and the scalar MUL/DIV unit can be minimal. +- Are fused integer MAC operations required by any target workload? +- What level of formal verification is mandated for XH-1? + +## Sources + +Primary and authoritative references used or cited in this document: + +- RISC-V International, *The RISC-V Instruction Set Manual, Volume I: Unprivileged Architecture* — official definition of RV64M (MUL, MULH, MULHSU, MULHU, DIV, DIVU, REM, REMU) and division overflow semantics. +- RISC-V International, *Architectural Compatibility Test Suite* (riscv-tests, riscv-arch-test) — official compliance test references. +- UC Berkeley Architecture Research, *Rocket Chip Generator* documentation — reference for the small iterative `MulDiv` module. +- UC Berkeley Architecture Research, *BOOM Out-of-Order Processor* documentation — reference for SRT-class dividers and pipelined multipliers in BOOM v2/v3. +- Hennessy & Patterson, *Computer Architecture: A Quantitative Approach* (recent editions) — workload frequency of MUL/DIV, energy/area considerations. +- Ercegovac & Lang, *Digital Arithmetic* — comprehensive treatment of shift-add, Booth, Wallace, SRT, and Newton–Raphson dividers. +- Parhami, *Computer Arithmetic: Algorithms and Hardware Designs* — additional reference for multiplier and divider architectures. + +ASSUMPTION: Specific page numbers and edition identifiers for Hennessy & Patterson, Ercegovac & Lang, and Parhami have not been quoted above because the exact editions in the XH-1 research library have not been recorded in this document. They should be cited precisely when this document is finalized. + +INSUFFICIENT EVIDENCE: No synthesis, layout, or PPA data for the target process node is yet available for any of the four design candidates. The relative area and energy figures are qualitative estimates only. diff --git a/research/.xh1/responses/20260825T180300Z_mul-div-unit.md.json b/research/.xh1/responses/20260825T180300Z_mul-div-unit.md.json new file mode 100644 index 0000000..bbff4f2 --- /dev/null +++ b/research/.xh1/responses/20260825T180300Z_mul-div-unit.md.json @@ -0,0 +1,305 @@ + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +{"id":"gen-1787680980-y6wlPzNkplZYwR1xpnVk","object":"chat.completion","created":1787680980,"model":"minimax/minimax-m3:free","provider":"GMICloud","system_fingerprint":null,"service_tier":null,"choices":[{"index":0,"logprobs":null,"finish_reason":"stop","native_finish_reason":"stop","message":{"role":"assistant","content":"# MUL/DIV Unit\n\n## Status\n\nDRAFT — initial research document. No architectural decision has been made for XH-1.\n\n## Abstract\n\nThis document investigates the design of the multiplication and division unit (MUL/DIV) for a single core inside the XH-1 128-core RISC-V processor. It surveys existing approaches for integer multiplication and division, identifies alternative implementations, and analyzes trade-offs in performance, area, power, latency, verification complexity, and scalability across 128 cores. The document focuses on the base integer extensions (RV64I/M) and explicitly defers floating-point and vector multiply/accumulate topics, which belong to separate units.\n\n## Research Question\n\nWhat is the most appropriate microarchitectural implementation of the MUL/DIV unit for a single XH-1 core, given that 128 identical cores will be instantiated on die, and given that the MUL/DIV unit must implement at minimum RV64M (MUL, MULH, MULHU, MULHSU, DIV, DIVU, REM, REMU)?\n\nSub-questions:\n\n- Should multiplication be iterative (shift-and-add) or fully combinational (array/Wallace/Booth)?\n- Should division use a restoring, non-restoring, SRT, or Newton–Raphson scheme?\n- Should the unit be pipelined, multi-cycle, or variable-latency?\n- How does the MUL/DIV unit interact with the surrounding pipeline (depth, bypass, in-order vs out-of-order issue)?\n- How does the unit scale when replicated 128 times on die?\n- Should fused MAC operations (MUL + ADD) or fused multiply-add be considered?\n\n## Background\n\nThe RISC-V \"M\" extension specifies eight integer multiply/divide instructions (MUL, MULH, MULHSU, MULHU, DIV, DIVU, REM, REMU) on 64-bit values producing either 64-bit or 128-bit results. Division is defined to round toward zero and is required to complete even for overflow cases (e.g., `INT64_MIN / -1`), as specified in the RISC-V Unprivileged ISA.\n\nKey properties of the workload that influence the design:\n\n- **Result width asymmetry.** MUL produces a 128-bit result, but only the low 64 bits are written to `rd` for `MUL`. MULH-family instructions write the high 64 bits. The datapath therefore needs at least a 64×64→128 multiplier followed by a selector.\n- **Sign handling.** Three signed/unsigned combinations (signed×signed, signed×unsigned, unsigned×unsigned) must be supported. Sign correction is required for MULH/MULHSU.\n- **Division latency and throughput.** RISC-V does not require division to retire in a single cycle, but the ISA mandates deterministic behavior for overflow.\n- **Rarity in many workloads.** Empirical studies (e.g., Hennessy & Patterson, *Computer Architecture: A Quantitative Approach*) report that integer divide and remainder instructions are uncommon (typically <1% of dynamic instructions), while multiplies are more frequent in HPC and crypto workloads.\n\nFor XH-1, the unit sits on the execution path of every core. With 128 cores, area and per-core energy dominate; raw single-thread latency of division matters less than aggregate throughput, die-area cost, and ease of verification.\n\n## Existing Approaches\n\n### Multiplication\n\n1. **Iterative shift-and-add multiplier**\n - One 64-bit adder reused across 64 cycles.\n - Smallest area, lowest energy per multiplication, but very long latency.\n\n2. **Array multiplier (combinational)**\n - 64×64 array of full adders producing a 128-bit result.\n - Single-cycle result, large area, high fan-out, and a long critical path.\n - Historically too slow for one cycle at high clock frequencies.\n\n3. **Wallace / Dadda tree**\n - Tree of carry-save adders reducing partial products to two 128-bit vectors, then a final carry-propagate adder.\n - Logarithmic depth; commonly used in high-performance cores.\n - Larger area than array, but shorter critical path.\n\n4. **Booth-encoded Wallace / Dadda**\n - Radix-4 or higher Booth recoding reduces the number of partial products by ~2×.\n - Common in modern cores (e.g., reported in implementations of ARM and x86 multipliers).\n\n5. **Pipelined iterative multiplier**\n - Splits the 64-cycle iterative multiplier into pipeline stages (commonly 2–4 stages).\n - Used in many in-order RISC-V cores (e.g., the Rocket Chip generator's `MulDiv`).\n\n6. **Dedicated single-cycle fused multiply-add (FMA) for integers**\n - Rare for integer-only M-extensions; usually belongs to the F/D extensions.\n\n### Division\n\n1. **Restoring division**\n - Classical shift-subtract. One bit per cycle, 64 cycles for 64-bit operands.\n - Simple, easy to verify.\n\n2. **Non-restoring division**\n - Similar latency, but allows a single add/subtract per bit without explicit restore.\n - Used in many textbook implementations.\n\n3. **SRT division**\n - Radix-4 or higher; produces 2+ bits per cycle using a small redundant quotient-digit table.\n - Significantly faster (16–32 cycles for 64-bit) at higher area cost.\n - Common in high-performance OoO cores (e.g., POWER, Itanium, recent x86).\n\n4. **Newton–Raphson reciprocal + multiply**\n - Iteratively refines an approximation of 1/d, then multiplies.\n - Very high throughput once the reciprocal is available.\n - Worst-case latency is higher than SRT; best for repeated divisions by the same divisor.\n - Rare in integer pipelines due to initial latency.\n\n5. **Goldschmidt division**\n - Similar to Newton–Raphson; iteratively scales numerator and denominator toward 1.\n - Same usage profile as Newton–Raphson.\n\n6. **Lookup-table based constant dividers**\n - For known small constant divisors, the compiler/runtime can replace DIV with a multiply-by-reciprocal.\n - Microarchitectural implication: the DIV unit need not be heavily optimized if software frequently replaces division by constants.\n\n## Alternative Designs\n\nFor the XH-1 MUL/DIV unit, four credible microarchitectural templates are considered.\n\n### Design A: Shared iterative multi-cycle unit (Rocket-style)\n\n- One 64-bit datapath reused for both MUL and DIV.\n- MUL: 1 cycle/partial product (radix-2 Booth optional).\n- DIV: 1 bit/cycle, restoring or non-restoring.\n- MUL latency: ~33–35 cycles (radix-2 Booth) or ~64 cycles (plain shift-add).\n- DIV latency: ~64 cycles.\n- Throughput: 1 MUL or DIV per ~32 cycles (shared); MUL and DIV cannot execute concurrently.\n- Plausible canonical reference: the `MulDiv` module in the BOOM/Rocket Chip generator (UC Berkeley).\n\n### Design B: Pipelined iterative multiplier + iterative divider\n\n- Multiplier: 2–4 stage pipelined radix-2/radix-4 iterative unit.\n- Divider: separate 64-bit iterative datapath (non-restoring or SRT-radix-2).\n- MUL throughput: 1 per cycle once pipeline is full.\n- DIV throughput: 1 per ~32 cycles.\n- Independent issue of MUL and DIV is possible.\n\n### Design C: Pipelined Wallace/Booth multiplier + SRT-radix-4 divider\n\n- Multiplier: 3-stage pipelined radix-4 Booth → Wallace tree → CPA. Produces full 128-bit result.\n- Divider: radix-4 SRT producing 2 bits/cycle; ~17 cycles for 64-bit DIV.\n- DIV/REM can be produced simultaneously since quotient digits are known.\n- Area: significantly larger than Designs A and B.\n- Latency: MUL ~3–4 cycles, DIV ~17–20 cycles.\n- Used in many modern superscalar cores.\n\n### Design F: Fused integer MAC / FMA\n\n- Add an integer fused multiply-add returning low 64 bits (a*b)+c in one operation.\n- This is non-standard for RV64M and would require custom opcodes or being staged behind a regular MUL+ADD sequence.\n- Documented here for completeness, not recommended without strong workload evidence.\n\n## Comparison\n\n| Property | A: Iterative shared | B: Pipelined iterative MUL + iterative DIV | C: Wallace/Booth MUL + SRT-4 DIV |\n|---|---|---|---|\n| MUL latency | ~33–64 cycles | 3–5 cycles | 3–4 cycles |\n| MUL throughput | 1 / 32 cycles | 1 / cycle | 1 / cycle |\n| DIV latency | ~64 cycles | ~32 cycles | ~16–20 cycles |\n| DIV throughput | 1 / 64 cycles | 1 / 32 cycles | 1 / 16 cycles |\n| 64×64→128 datapath | Yes (shared) | Yes (MUL only) | Yes (Wallace) |\n| MUL+DIV concurrency | No (shared) | Yes (separate datapaths) | Yes (separate datapaths) |\n| Estimated relative area | 1.0× | ~1.5–2.0× | ~3.0–5.0× |\n| Estimated critical path | Short | Short | Longest (CPA final stage) |\n| Verification complexity | Low | Medium | High |\n| Fits \"small in-order\" model | Excellent | Good | Marginal |\n\nASSUMPTION: Area estimates above are rough order-of-magnitude relative numbers based on typical RISC-V implementations and the cited textbooks. They have not been measured for XH-1.\n\n## Advantages\n\n### Design A (iterative shared)\n\n- Smallest area per core, which directly reduces die cost across 128 cores.\n- Lowest per-core dynamic energy for the rare case of an actual MUL/DIV.\n- Easiest to verify formally (small state space, one datapath).\n- Matches the \"many small cores\" scaling philosophy.\n- Canonical reference: Rocket Chip `MulDiv`.\n\n### Design B (pipelined iterative MUL + iterative DIV)\n\n- MUL throughput is high enough to support HPC and crypto workloads where 64-bit multiplies are common.\n- DIV remains simple.\n- Area increase over A is bounded.\n- Two independent datapaths simplify scheduling in the issue stage.\n\n### Design C (Wallace/Booth + SRT-4)\n\n- Best raw latency and throughput for both operations.\n- Suitable for single-core-bound workloads or for cores that need to hide memory latency behind fast arithmetic.\n- DIV+REM can be produced together with little extra hardware.\n\n## Disadvantages\n\n### Design A\n\n- DIV latency of 64 cycles is long; if the surrounding pipeline is short (e.g., 5–7 stages), the unit will dominate total execution time for any divide.\n- Back-to-back MULs serialize.\n- Under HPC or cryptography kernels, MUL throughput becomes a bottleneck.\n\n### Design B\n\n- More area than A.\n- Pipelined MUL increases register pressure in the issue queue and requires more bypass paths in the surrounding execution stage.\n- DIV still slow.\n\n### Design C\n\n- Largest area per core, replicated 128 times.\n- Highest per-core power.\n- Wallace tree and SRT have long critical paths that may limit clock frequency for the whole core.\n- Verification complexity is significantly higher: partial-product reduction, Booth recoding, SRT quotient-digit selection tables, and divider corner cases (e.g., `INT64_MIN / -1`) all need separate coverage.\n- Wall-clock design and verification cost may delay the whole project.\n\n## XH-1 Considerations\n\nPROPOSAL: For XH-1, an in-order core with 128 instances on die, the dominant design constraint is **per-core area, energy, and verification cost**, not single-thread peak performance. The MUL/DIV unit should therefore favor small, simple, well-trodden implementations.\n\nSpecific implications for XH-1:\n\n- The 128-core factor means the MUL/DIV unit's area is multiplied by 128. Even a 2× area difference per core translates to a substantial absolute area delta.\n- The energy of 128 MUL/DIV datapaths, even at low utilization, contributes to total socket power.\n- A long-latency MUL/DIV unit is acceptable if the surrounding pipeline is deep enough or if it can be overlapped with other in-flight instructions in the same core.\n- Single-cycle MUL would impose a critical path on the whole core; for a 128-core design, sustained high clock frequency across all cores is critical to total throughput.\n\n## 128-Core Scalability\n\nScalability dimensions to consider:\n\n- **Wiring and layout.** A 128-core die has a complex interconnect. A small MUL/DIV unit is easier to place and route within each core tile. Designs with large irregular adder trees (Wallace/SRT) complicate physical design at high core counts.\n- **Verification replication.** Bugs in the MUL/DIV unit, if present, propagate to 128 cores. A simpler, formally verifiable design (Design A) is safer for replication.\n- **Yield.** Smaller per-core area improves yield and binning flexibility; large per-core area reduces the number of cores that fit on a reticle at the target process node.\n- **Power delivery.** 128 simultaneous MUL/DIV operations are unlikely, but worst-case power events (e.g., SIMD-style vector MUL workloads scaled down to integer MUL) must be within the socket's power-delivery budget.\n- **Frequency scaling.** A 128-core chip with modest per-core frequency but high aggregate throughput may benefit from a short critical path. A Wallace multiplier's critical path can limit fmax for the whole core.\n\n## Performance Considerations\n\n- MUL/DIV instructions are infrequent in general-purpose workloads (often <1% dynamic instructions) but can dominate kernels in cryptography (AES, ChaCha20, RSA), big-integer arithmetic (GMP-style libraries), and some HPC kernels.\n- If XH-1 is intended for general-purpose server or desktop use, the MUL/DIV unit will rarely be on the critical path of a thread.\n- If XH-1 targets HPC or cryptography, the MUL throughput becomes important. In this case, B or C should be reconsidered.\n- Software can use compiler transformations to replace DIV by constants with multiply-by-reciprocal, reducing pressure on the DIV unit.\n\n## Implementation Considerations\n\n- **Sign handling.** The unit must correctly handle `MULH`, `MULHSU`, and `MULHU` as well as overflow cases of DIV (notably `INT64_MIN / -1`, which must produce `INT64_MIN` per RISC-V spec).\n- **REM vs DIV.** Producing REM in parallel with DIV using the same datapath is standard in restoring/non-restoring designs; the unit should support issuing DIVU/REMU pairs in one operation.\n- **Pipeline interface.** The unit must integrate with the core's issue, wakeup, and writeback stages. If in-order, the issue stage must stall in-order cores on multi-cycle MUL/DIV. If OoO, completion must wait for the unit's completion signal.\n- **Bypassing.** Forwarding paths from the MUL/DIV pipeline registers to dependent instructions must be designed carefully to avoid structural hazards.\n- **Early termination.** For DIV, the unit can terminate early when the remainder is zero, saving cycles. Implementation cost is low.\n\n## Verification Considerations\n\n- **Corner cases.** RV64M has well-defined corner cases: `INT64_MIN / -1`, division by zero, overflow in REM, sign interactions in MULH-family.\n- **Directed + constrained-random.** A combination of directed tests for ISA corner cases and constrained-random for the rest is standard practice (e.g., as in the RISC-V architectural test framework, riscv-tests).\n- **Formal verification.** A small iterative multiplier/divider (Design A) is amenable to formal proofs of correctness for a few-bit case and inductive scaling. A Wallace + SRT unit (Design C) is significantly harder to formally verify due to selector-table complexity.\n- **Cross-core equivalence.** With 128 identical cores, regression in one core implies regression in all 128. A well-verified single-core design simplifies the chip-level verification effort.\n- **Testbench reuse.** The RISC-V community maintains architectural compliance tests that should be run against the MUL/DIV unit regardless of the chosen design.\n\n## Recommendation\n\nPROPOSAL: Adopt a **Design B–leaning approach**: a small, simple, well-understood MUL/DIV unit similar in spirit to Rocket Chip's `MulDiv`, with the following characteristics:\n\n- A **radix-4 Booth-encoded iterative multiplier** (or radix-2 if radix-4 proves too complex for the area budget) producing 64 bits of result per ~16 cycles, sharing partial datapath with the divider if needed.\n- A **non-restoring (or radix-2 SRT) divider** completing in ~32 cycles.\n- MUL and DIV on the same datapath with **shared state** but capable of being interleaved at issue time.\n- Optional microarchitectural relaxation: a separate tiny **fast-MUL path for 32×32→64 results** (the low half of MUL where both operands are sign- or zero-extended from 32 bits) to accelerate common cases. This adds minimal area.\n\nThis recommendation is provisional and is the lightest-weight option that still keeps MUL throughput reasonable. It avoids the critical-path cost of Design C and the throughput limit of Design A, while remaining well within the verification budget of a 128-core project.\n\nRECOMMENDATION: If workload analysis (not yet performed) shows MUL-heavy HPC/cryptography use, escalate to a pipelined radix-4 Booth multiplier with a 2–3 cycle latency, keeping the iterative divider. If workload analysis shows almost no MUL/DIV usage, drop to a plain Design A.\n\n## Confidence\n\n- **Low–Medium** for any specific microarchitectural recommendation. The document is at an early stage; the recommendation will be revised after:\n 1. Workload analysis (target use cases of XH-1).\n 2. Synthesis of representative MUL/DIV units in the target technology.\n 3. Frequency, area, and power target constraints.\n- **High** that the iterative, shared-datapath approach (Design A or B) is the appropriate starting point for a 128-core, area-constrained, verification-constrained design.\n\n## Open Questions\n\n- What is the target frequency of XH-1 cores, and what is the critical-path budget for the MUL/DIV unit?\n- What process node is targeted, and what is the per-core area budget?\n- What is the intended workload mix (server, HPC, embedded, ML)?\n- Is the core in-order or out-of-order? The MUL/DIV interface and latency tolerance depend strongly on this.\n- Will the F extension (floating-point) be present in the same core, and if so, should integer MUL be reused inside an FMA datapath?\n- Will the V extension (vector) be present? If so, scalar integer MUL may be lightly used and the scalar MUL/DIV unit can be minimal.\n- Are fused integer MAC operations required by any target workload?\n- What level of formal verification is mandated for XH-1?\n\n## Sources\n\nPrimary and authoritative references used or cited in this document:\n\n- RISC-V International, *The RISC-V Instruction Set Manual, Volume I: Unprivileged Architecture* — official definition of RV64M (MUL, MULH, MULHSU, MULHU, DIV, DIVU, REM, REMU) and division overflow semantics.\n- RISC-V International, *Architectural Compatibility Test Suite* (riscv-tests, riscv-arch-test) — official compliance test references.\n- UC Berkeley Architecture Research, *Rocket Chip Generator* documentation — reference for the small iterative `MulDiv` module.\n- UC Berkeley Architecture Research, *BOOM Out-of-Order Processor* documentation — reference for SRT-class dividers and pipelined multipliers in BOOM v2/v3.\n- Hennessy & Patterson, *Computer Architecture: A Quantitative Approach* (recent editions) — workload frequency of MUL/DIV, energy/area considerations.\n- Ercegovac & Lang, *Digital Arithmetic* — comprehensive treatment of shift-add, Booth, Wallace, SRT, and Newton–Raphson dividers.\n- Parhami, *Computer Arithmetic: Algorithms and Hardware Designs* — additional reference for multiplier and divider architectures.\n\nASSUMPTION: Specific page numbers and edition identifiers for Hennessy & Patterson, Ercegovac & Lang, and Parhami have not been quoted above because the exact editions in the XH-1 research library have not been recorded in this document. They should be cited precisely when this document is finalized.\n\nINSUFFICIENT EVIDENCE: No synthesis, layout, or PPA data for the target process node is yet available for any of the four design candidates. The relative area and energy figures are qualitative estimates only.","refusal":null,"reasoning":null}}],"usage":{"prompt_tokens":777,"completion_tokens":4495,"total_tokens":5272,"cost":0,"is_byok":false,"prompt_tokens_details":{"cached_tokens":762,"cache_write_tokens":0,"audio_tokens":0,"video_tokens":0},"cost_details":{"upstream_inference_cost":0,"upstream_inference_prompt_cost":0,"upstream_inference_completions_cost":0},"completion_tokens_details":{"reasoning_tokens":0,"image_tokens":0,"audio_tokens":0}}} \ No newline at end of file diff --git a/research/.xh1/responses/20260825T180300Z_mul-div-unit.md.review b/research/.xh1/responses/20260825T180300Z_mul-div-unit.md.review new file mode 100644 index 0000000..e0db023 --- /dev/null +++ b/research/.xh1/responses/20260825T180300Z_mul-div-unit.md.review @@ -0,0 +1,37 @@ +VERDICT: FAIL + +ISSUES: +- **Design B–leaning recommendation contradicts the stated proposal and the comparison table.** The "XH-1 Considerations" and "128-Core Scalability" sections argue decisively that area, energy, verification, and 128× replication favor the smallest, simplest design (Design A). The recommendation then promotes Design B (or a Design B–leaning hybrid) without reconciling the conflict. The justification "avoids the critical-path cost of Design C and the throughput limit of Design A" is hand-waving; the comparison table shows Design A's critical path is shortest. +- **The four candidate designs are labeled inconsistently.** The "Existing Approaches" section enumerates multiplication and division schemes separately, but the "Alternative Designs" section labels them A, B, C, then jumps to "Design F" (fused MAC) without a Design D or E. This is sloppy and obscures the design space — e.g., where is the pipelined iterative multiplier with SRT-radix-4 divider option? Where is a non-pipelined Wallace+non-restoring hybrid? +- **Design B is mischaracterized in the comparison table.** "Pipelined iterative MUL + iterative DIV" with "MUL latency 3–5 cycles" and "1/cycle" throughput is not really an "iterative" multiplier — that description fits a pipelined combinational or partially combinational datapath. The terminology is inconsistent with the rest of the document and with how Rocket's `MulDiv` actually works (which is closer to Design A, not B). +- **Unsupported area/energy claims.** "Estimated relative area 1.0× / ~1.5–2.0× / ~3.0–5.0×" and the corresponding energy numbers are presented as quasi-data with no synthesis backing, no source citation, and no process node. The document acknowledges this in a footnote-style ASSUMPTION, but the numbers are still used as if comparable in the comparison table and in the recommendation. For a research document whose central trade-off is area-per-core × 128, this is a serious gap. +- **Missing alternative: constant-divider optimization (magic-number multiplication).** The "Existing Approaches" section mentions lookup-table / multiply-by-reciprocal constant dividers in passing, but no design candidate explores offloading DIV-by-constant entirely to software or a dedicated microarchitectural helper. Given that compilers routinely do this and the document explicitly states MUL/DIV are rare, omitting this as a real alternative is a gap. +- **Missing alternative: shared cluster-level MUL/DIV unit.** For a 128-core design, a recurring microarchitectural option is to share one (or a few) high-performance MUL/DIV units across a cluster of cores (e.g., 4 or 8 cores share a Wallace/SRT unit), trading per-core latency for die area. The document does not consider this at all, even though it is highly relevant to the stated scaling concern. This is a meaningful missing alternative. +- **Missing alternative: vector / SIMD reuse.** The document explicitly defers vector MUL/MAC but does not consider whether the scalar MUL/DIV unit should be designed knowing that the V extension (if present) would dominate aggregate multiply throughput. The "Open Questions" section flags this but it should be in the alternatives analysis. +- **Failure to consider 128 cores in the recommendation logic.** The recommendation says "B-leaning" while the 128-core analysis section argues for "A." The document never reconciles these; it just presents both without resolving. This is an internal contradiction. +- **The "fast 32×32→64 MUL path" recommendation is unsupported.** Stated as "minimal area" with no area estimate and no discussion of how it interacts with the shared datapath, the issue logic, the partial-product generator, or verification. It is a non-trivial addition that needs justification. +- **In-order vs out-of-order core assumption is missing.** The recommendation is presented as if the core is in-order, but the document never states the core microarchitecture. Most of the latency-vs-throughput trade-off and the Wallace/SRT critical-path analysis hinges on this. This is a critical missing assumption. +- **"Short critical path" claim for the iterative divider is unsupported.** A 64-bit non-restoring divider has a 64-bit adder in its critical path, which is comparable to other units in the core. The document asserts the iterative approach has the shortest critical path without discussing the actual datapath depth of the 64-bit CPA inside the divider loop. +- **"DIV/REM produced simultaneously" claim for SRT-radix-4 is misleading.** SRT produces quotient digits; the remainder is recovered at the end from a redundant form and typically requires a correction step. Stating "DIV/REM can be produced simultaneously since quotient digits are known" elides this and is borderline incorrect. +- **Wallace tree is described as having a "longer critical path" than array, which is backwards.** A Wallace tree has a *logarithmic* depth and a strictly shorter critical path than the linear array; the trade-off is area/routing, not critical path. This is a factual error. +- **"Radix-4 Booth reduces partial products by ~2×" is sloppy.** Radix-4 Booth recodes 64 bits into 33 signed digits, reducing partial products from 64 to 33 — roughly 2× fewer, but the document should state the exact number. Minor, but symptomatic. +- **Citation hygiene is poor.** "Hennessy & Patterson" and "Ercegovac & Lang" are cited with no specific edition, chapter, page, or even which textbook (H&P has multiple versions; "recent editions" is not a citation). For a research document, this is a serious sourcing weakness, and the document itself acknowledges the gap. +- **"Reported in implementations of ARM and x86 multipliers"** — no specific ARM core or x86 microarchitecture is named. This is a hallucination-prone claim with no source. +- **"Empirical studies (e.g., Hennessy & Patterson) report that integer divide and remainder instructions are uncommon (typically <1%)"** — the specific number and study are not cited. Different workloads vary widely; some SPEC int workloads have higher divide frequency. Unsupported. +- **Workload analysis is repeatedly deferred but recommendations are made anyway.** The recommendation, the proposal, and the open questions all say "workload analysis pending" yet a specific microarchitectural recommendation is still made. This is unjustified given the document's own caveats. +- **No quantitative power analysis at all.** For a 128-core design, dynamic and leakage power of replicated MUL/DIV units should be at least estimated, even at a high level. The document says "energy" repeatedly but provides no numbers. +- **The 128-core verification claim ("formal proofs of correctness for a few-bit case and inductive scaling") is a hand-wave.** Inductive scaling proofs of arbitrary Booth/SRT/non-restoring dividers are not standard practice and are not trivial. This overstates the state of formal verification for these units. + +REQUIRED FIXES: +- Resolve the contradiction between the "XH-1 Considerations" argument (which favors Design A) and the recommendation (which favors Design B). Either justify the deviation explicitly or align the recommendation with the analysis. +- Fix the factual error about Wallace trees (shorter, not longer, critical path than array multipliers) and the misleading SRT remainder-claim. +- Renumber design candidates consistently (A, B, C, D, …) and include the missing alternatives: cluster-shared MUL/DIV, constant-divider / magic-number offload, and a hybrid Wallace + non-restoring option. +- Provide actual synthesis, area, and energy estimates from at least one reference technology (e.g., a published RISC-V core, or a synthesized estimate from open-source IP like Rocket/BOOM) rather than qualitative multipliers with no anchor. +- Add a discussion of how the recommendation depends on in-order vs out-of-order core microarchitecture, and explicitly state the assumed core type for this document. +- Provide specific citations (edition, chapter/page) for Hennessy & Patterson, Ercegovac & Lang, and Parhami, or remove the claims that depend on them. +- Either substantiate or remove the "32×32 fast MUL" recommendation; it is currently an unjustified adder. +- Include a power/energy analysis (even rough) for 128-core replication; this is central to the stated design constraint. +- Quantify or remove the "1% dynamic instructions" claim; cite the specific study and workload. +- Address verification claims more carefully: do not overstate the practicality of inductive formal proofs for divider corner cases, and note the realistic state of formal vs constrained-random coverage for each design. + +CONFIDENCE: HIGH diff --git a/research/.xh1/responses/20260825T180300Z_mul-div-unit.md.review.json b/research/.xh1/responses/20260825T180300Z_mul-div-unit.md.review.json new file mode 100644 index 0000000..cd58977 --- /dev/null +++ b/research/.xh1/responses/20260825T180300Z_mul-div-unit.md.review.json @@ -0,0 +1,111 @@ + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +{"id":"gen-1787681044-rDBl26S3P1YgOhoaSWIL","object":"chat.completion","created":1787681044,"model":"minimax/minimax-m3:free","provider":"GMICloud","system_fingerprint":null,"service_tier":null,"choices":[{"index":0,"logprobs":null,"finish_reason":"stop","native_finish_reason":"stop","message":{"role":"assistant","content":"VERDICT: FAIL\n\nISSUES:\n- **Design B–leaning recommendation contradicts the stated proposal and the comparison table.** The \"XH-1 Considerations\" and \"128-Core Scalability\" sections argue decisively that area, energy, verification, and 128× replication favor the smallest, simplest design (Design A). The recommendation then promotes Design B (or a Design B–leaning hybrid) without reconciling the conflict. The justification \"avoids the critical-path cost of Design C and the throughput limit of Design A\" is hand-waving; the comparison table shows Design A's critical path is shortest.\n- **The four candidate designs are labeled inconsistently.** The \"Existing Approaches\" section enumerates multiplication and division schemes separately, but the \"Alternative Designs\" section labels them A, B, C, then jumps to \"Design F\" (fused MAC) without a Design D or E. This is sloppy and obscures the design space — e.g., where is the pipelined iterative multiplier with SRT-radix-4 divider option? Where is a non-pipelined Wallace+non-restoring hybrid?\n- **Design B is mischaracterized in the comparison table.** \"Pipelined iterative MUL + iterative DIV\" with \"MUL latency 3–5 cycles\" and \"1/cycle\" throughput is not really an \"iterative\" multiplier — that description fits a pipelined combinational or partially combinational datapath. The terminology is inconsistent with the rest of the document and with how Rocket's `MulDiv` actually works (which is closer to Design A, not B).\n- **Unsupported area/energy claims.** \"Estimated relative area 1.0× / ~1.5–2.0× / ~3.0–5.0×\" and the corresponding energy numbers are presented as quasi-data with no synthesis backing, no source citation, and no process node. The document acknowledges this in a footnote-style ASSUMPTION, but the numbers are still used as if comparable in the comparison table and in the recommendation. For a research document whose central trade-off is area-per-core × 128, this is a serious gap.\n- **Missing alternative: constant-divider optimization (magic-number multiplication).** The \"Existing Approaches\" section mentions lookup-table / multiply-by-reciprocal constant dividers in passing, but no design candidate explores offloading DIV-by-constant entirely to software or a dedicated microarchitectural helper. Given that compilers routinely do this and the document explicitly states MUL/DIV are rare, omitting this as a real alternative is a gap.\n- **Missing alternative: shared cluster-level MUL/DIV unit.** For a 128-core design, a recurring microarchitectural option is to share one (or a few) high-performance MUL/DIV units across a cluster of cores (e.g., 4 or 8 cores share a Wallace/SRT unit), trading per-core latency for die area. The document does not consider this at all, even though it is highly relevant to the stated scaling concern. This is a meaningful missing alternative.\n- **Missing alternative: vector / SIMD reuse.** The document explicitly defers vector MUL/MAC but does not consider whether the scalar MUL/DIV unit should be designed knowing that the V extension (if present) would dominate aggregate multiply throughput. The \"Open Questions\" section flags this but it should be in the alternatives analysis.\n- **Failure to consider 128 cores in the recommendation logic.** The recommendation says \"B-leaning\" while the 128-core analysis section argues for \"A.\" The document never reconciles these; it just presents both without resolving. This is an internal contradiction.\n- **The \"fast 32×32→64 MUL path\" recommendation is unsupported.** Stated as \"minimal area\" with no area estimate and no discussion of how it interacts with the shared datapath, the issue logic, the partial-product generator, or verification. It is a non-trivial addition that needs justification.\n- **In-order vs out-of-order core assumption is missing.** The recommendation is presented as if the core is in-order, but the document never states the core microarchitecture. Most of the latency-vs-throughput trade-off and the Wallace/SRT critical-path analysis hinges on this. This is a critical missing assumption.\n- **\"Short critical path\" claim for the iterative divider is unsupported.** A 64-bit non-restoring divider has a 64-bit adder in its critical path, which is comparable to other units in the core. The document asserts the iterative approach has the shortest critical path without discussing the actual datapath depth of the 64-bit CPA inside the divider loop.\n- **\"DIV/REM produced simultaneously\" claim for SRT-radix-4 is misleading.** SRT produces quotient digits; the remainder is recovered at the end from a redundant form and typically requires a correction step. Stating \"DIV/REM can be produced simultaneously since quotient digits are known\" elides this and is borderline incorrect.\n- **Wallace tree is described as having a \"longer critical path\" than array, which is backwards.** A Wallace tree has a *logarithmic* depth and a strictly shorter critical path than the linear array; the trade-off is area/routing, not critical path. This is a factual error.\n- **\"Radix-4 Booth reduces partial products by ~2×\" is sloppy.** Radix-4 Booth recodes 64 bits into 33 signed digits, reducing partial products from 64 to 33 — roughly 2× fewer, but the document should state the exact number. Minor, but symptomatic.\n- **Citation hygiene is poor.** \"Hennessy & Patterson\" and \"Ercegovac & Lang\" are cited with no specific edition, chapter, page, or even which textbook (H&P has multiple versions; \"recent editions\" is not a citation). For a research document, this is a serious sourcing weakness, and the document itself acknowledges the gap.\n- **\"Reported in implementations of ARM and x86 multipliers\"** — no specific ARM core or x86 microarchitecture is named. This is a hallucination-prone claim with no source.\n- **\"Empirical studies (e.g., Hennessy & Patterson) report that integer divide and remainder instructions are uncommon (typically <1%)\"** — the specific number and study are not cited. Different workloads vary widely; some SPEC int workloads have higher divide frequency. Unsupported.\n- **Workload analysis is repeatedly deferred but recommendations are made anyway.** The recommendation, the proposal, and the open questions all say \"workload analysis pending\" yet a specific microarchitectural recommendation is still made. This is unjustified given the document's own caveats.\n- **No quantitative power analysis at all.** For a 128-core design, dynamic and leakage power of replicated MUL/DIV units should be at least estimated, even at a high level. The document says \"energy\" repeatedly but provides no numbers.\n- **The 128-core verification claim (\"formal proofs of correctness for a few-bit case and inductive scaling\") is a hand-wave.** Inductive scaling proofs of arbitrary Booth/SRT/non-restoring dividers are not standard practice and are not trivial. This overstates the state of formal verification for these units.\n\nREQUIRED FIXES:\n- Resolve the contradiction between the \"XH-1 Considerations\" argument (which favors Design A) and the recommendation (which favors Design B). Either justify the deviation explicitly or align the recommendation with the analysis.\n- Fix the factual error about Wallace trees (shorter, not longer, critical path than array multipliers) and the misleading SRT remainder-claim.\n- Renumber design candidates consistently (A, B, C, D, …) and include the missing alternatives: cluster-shared MUL/DIV, constant-divider / magic-number offload, and a hybrid Wallace + non-restoring option.\n- Provide actual synthesis, area, and energy estimates from at least one reference technology (e.g., a published RISC-V core, or a synthesized estimate from open-source IP like Rocket/BOOM) rather than qualitative multipliers with no anchor.\n- Add a discussion of how the recommendation depends on in-order vs out-of-order core microarchitecture, and explicitly state the assumed core type for this document.\n- Provide specific citations (edition, chapter/page) for Hennessy & Patterson, Ercegovac & Lang, and Parhami, or remove the claims that depend on them.\n- Either substantiate or remove the \"32×32 fast MUL\" recommendation; it is currently an unjustified adder.\n- Include a power/energy analysis (even rough) for 128-core replication; this is central to the stated design constraint.\n- Quantify or remove the \"1% dynamic instructions\" claim; cite the specific study and workload.\n- Address verification claims more carefully: do not overstate the practicality of inductive formal proofs for divider corner cases, and note the realistic state of formal vs constrained-random coverage for each design.\n\nCONFIDENCE: HIGH","refusal":null,"reasoning":null}}],"usage":{"prompt_tokens":4814,"completion_tokens":1841,"total_tokens":6655,"cost":0,"is_byok":false,"prompt_tokens_details":{"cached_tokens":134,"cache_write_tokens":0,"audio_tokens":0,"video_tokens":0},"cost_details":{"upstream_inference_cost":0,"upstream_inference_prompt_cost":0,"upstream_inference_completions_cost":0},"completion_tokens_details":{"reasoning_tokens":0,"image_tokens":0,"audio_tokens":0}}} \ No newline at end of file diff --git a/research/.xh1/runs/20260825T180809Z/attempts b/research/.xh1/runs/20260825T180809Z/attempts new file mode 100644 index 0000000..e69de29 diff --git a/research/.xh1/runs/20260825T180809Z/completed b/research/.xh1/runs/20260825T180809Z/completed new file mode 100644 index 0000000..e69de29 diff --git a/research/.xh1/runs/20260825T180809Z/failed b/research/.xh1/runs/20260825T180809Z/failed new file mode 100644 index 0000000..e69de29 diff --git a/research/.xh1/runs/20260825T180900Z/attempts b/research/.xh1/runs/20260825T180900Z/attempts new file mode 100644 index 0000000..af7dc9a --- /dev/null +++ b/research/.xh1/runs/20260825T180900Z/attempts @@ -0,0 +1,2 @@ +2026-08-25T18:09:56Z research/03-core-design/mul-div-unit.md 1 research completed +2026-08-25T18:10:12Z research/03-core-design/mul-div-unit.md 1 review api-failure diff --git a/research/.xh1/runs/20260825T180900Z/completed b/research/.xh1/runs/20260825T180900Z/completed new file mode 100644 index 0000000..e69de29 diff --git a/research/.xh1/runs/20260825T180900Z/failed b/research/.xh1/runs/20260825T180900Z/failed new file mode 100644 index 0000000..e427ae2 --- /dev/null +++ b/research/.xh1/runs/20260825T180900Z/failed @@ -0,0 +1 @@ +research/03-core-design/mul-div-unit.md diff --git a/research/.xh1/runs/20260825T180900Z/research_03-core-design_mul-div-unit.md.current.md b/research/.xh1/runs/20260825T180900Z/research_03-core-design_mul-div-unit.md.current.md new file mode 100644 index 0000000..3217b20 --- /dev/null +++ b/research/.xh1/runs/20260825T180900Z/research_03-core-design_mul-div-unit.md.current.md @@ -0,0 +1,194 @@ +# Multiplier–Divider Unit for the XH-1 Core + +## Status +DRAFT — Engineering research, no silicon data. + +## Abstract +This document investigates the design of an integer multiply/divide (MUL/DIV) unit suitable for one tile of the XH-1 128-core RISC-V processor. It frames the problem space (RV32M/RV64M, latency vs. throughput, area and energy), surveys the canonical microarchitectural options (iterative shift-and-add, array multipliers, Booth/Wallace/Dadda trees, reciprocal/sRT dividers, radix-2/4/8 dividers, and merged MAC-fused units), and maps them against the constraints implied by the other XH-1 core documents (small per-core budget, 128-core replication, expected integration with the pipeline, register file, and load–store unit). The document is written so that downstream choices about pipeline depth, core clock target, and shared vs. private FP/MAC can be made coherently with the MUL/DIV decision. + +## Research Question +What microarchitecture for the integer multiply/divide unit best fits one XH-1 core, given that the core will be replicated 128 times, must support the M extension on RV32 and/or RV64, and must coexist with a pipeline, register file, and load–store unit whose budgets are not yet fixed? + +Sub-questions: +1. Should the unit be fully combinatorial, multi-cycle iterative, or pipelined? +2. Which radix and which encoding (Booth-2, Booth-3, modified Booth-4) is appropriate for the expected operand width? +3. Should MUL and DIV share silicon, or be separate datapaths? +4. How does the choice interact with 128-core replication (area amortization, frequency, yield)? +5. What are the verification implications of each choice? + +## Background +The RISC-V M extension defines four signed/unsigned multiply variants producing 2X-bit results (MUL, MULH, MULHU, MULHSU), four matching multiply-high variants, and signed/unsigned divide and remainder (DIV, DIVU, REM, REMU). On RV64 there is an additional MULW/DIVW/REM family that operates on 32-bit values and sign-extends. + +Key microarchitectural properties that drive the design: +- Latency-tolerance: in a deep pipeline, multi-cycle iterative units are acceptable as long as the structural hazard is bounded and the scoreboard/issue logic handles in-flight multiplies. +- Throughput: the M extension opcodes are infrequent in many workloads but dominate in linear-algebra kernels, crypto, and hash functions. +- Operand width: an RV64 core needs both 64-bit and 32-bit paths. A 64×64→128 multiplier is roughly 4× the area of a 32×32→64 multiplier. +- Result width: MULH-family instructions require the full 2X-bit result; MUL only requires the low X bits. Most designs share the upper datapath and select the low half. + +ASSUMPTION: XH-1 cores implement the M extension. If only the I extension is required, the MUL/DIV unit collapses dramatically. The remainder of this document assumes M is in scope. +ASSUMPTION: XH-1 is RV64. If RV32 only, all 64-bit-specific considerations below can be relaxed. + +## Existing Approaches + +### Multipliers +1. Iterative shift-and-add multiplier +2. Array (braid) multiplier +3. Wallace tree multiplier +4. Dadda tree multiplier +5. Booth-encoded tree multiplier (radix-4, modified Booth) +6. Higher-radix Booth (radix-8/16) with 4:2/5:2 compressor trees +7. Pipelined versions of any of the above, with stage counts from 1 to N + +### Dividers +1. Restoring divider +2. Non-restoring divider +3. SRT divider (s radix-2, radix-4, radix-8, radix-16) +4. Newton–Raphson reciprocal + multiply (software or hardware) +5. Goldschmidt divider +6. Digit-recurrence with prescaling (for faster convergence) +7. Lookup-table (LUT) assisted dividers (small LUT, big LUT) + +PROPOSAL: Classify each option along four axes — latency (cycles), throughput (1/n per cycle), area (gate equivalent or µm² estimate), and design complexity (verification effort, corner cases). + +## Alternative Designs + +### Option A: Pipelined iterative multiplier + iterative non-restoring divider +- One shared 64-bit datapath, Booth-2 radix-4, latency ≈ 3–4 cycles for MUL, fully pipelined at 1 result/cycle. +- Divider: non-restoring, ≈ 32–64 cycles for RV64 DIV, variable. +- Area: smallest among the options; one of the smallest in published small-core implementations. +- Verification: standard; iterative state machines are well understood. + +### Option B: Fully combinational 64×64 array multiplier + SRT-4 divider +- MUL latency: 1 cycle, but very high area and long critical path. +- DIV: SRT-4, ≈ 16 cycles typical. +- Area: largest. +- Verification: simple timing closure problem, but 128 cores × this area is likely prohibitive. + +### Option C: Pipelined Wallace/Booth tree (3-stage) + SRT-4 divider +- MUL: 3-cycle pipelined, 1 result/cycle sustained. +- DIV: 16–20 cycles, fully pipelined divider. +- Area: moderate to large; tree is irregular. +- Verification: irregular partial-product reduction is harder to verify than array. + +### Option D: Fused multiply–accumulate (MAC) with Booth-3 (radix-8) and 4:2 compressors +- MUL: 3-cycle, also accepts an accumulate operand each cycle. +- Provides a useful primitive for dot-product kernels and soft-fp libraries. +- Divider: separate iterative non-restoring path. +- Area: similar to C, but the accumulator latch and bypass network add cost. + +### Option E: Truncated 32×32→64 only, with MULHW/DIVW synthesized in microcode +- Illegal in general: software synthesis of MULH is too slow to be acceptable. +- Listed for completeness and to discard. + +## Comparison + +| Option | MUL latency | MUL throughput | DIV latency (RV64) | Area (relative) | Design complexity | +|--------|-------------|----------------|---------------------|------------------|--------------------| +| A — Iterative Booth + NR div | 3–4 cyc | 1/cyc pipelined | 32–64 cyc | 0.6–0.8× | Low | +| B — Array + SRT-4 | 1 cyc | 1/cyc | 16 cyc | 2.5–3.0× | Medium (timing) | +| C — Wallace/Booth + SRT-4 | 3 cyc | 1/cyc | 16–20 cyc | 1.0–1.2× | High (tree) | +| D — MAC (Booth-3, 4:2) + NR div | 3 cyc | 1/cyc (with acc) | 32–64 cyc | 1.1–1.3× | High | +| E — Truncated, microcoded | n/a | n/a | n/a | smallest | n/a (incomplete) | + +Numbers above are ORDER-OF-MAGNITUDE ESTIMATES based on published small-core implementations; INSUFFICIENT EVIDENCE exists to claim a specific gate count or µm² without synthesis at a known target node. + +## Advantages +- Option A: smallest area per core → most 128-core replication headroom; trivially fits any pipeline depth; well-understood verification. +- Option B: best single-cycle latency; useful for OoO cores with tight issue windows. +- Option C: good balance of latency and area; widely used in commercial OoO cores. +- Option D: enables efficient dot-product and SIMD-style software; helps a soft FP stack. + +## Disadvantages +- Option A: lower peak MUL throughput than pipelined trees; multi-cycle DIV may stall the issue queue on back-to-back DIVs. +- Option B: area is prohibitive at 128 cores; long critical path forces a slow core clock or deep pipelining (defeating the latency benefit). +- Option C: Wallace/Dadda partial-product reduction has irregular carry-save structures that are harder to formally verify and harder to fix in ECO. +- Option D: accumulator adds bypass and forwarding complexity into the pipeline and register-file writeback path; benefits only workloads that can be rewritten to use the MAC. +- Option E: not viable for an M-extension-compliant core. + +## XH-1 Considerations +ASSUMPTION: XH-1 is a tiled, replicated design where per-core area is a first-class constraint because the whole 128-core array must fit in the package, power, and yield envelope. A small per-core MUL/DIV is therefore a high-value design point. + +PROPOSAL: Treat MUL/DIV as one of the "shared-tile resource" candidates. Specifically: +- If 128 cores × 1 MUL/DIV per core exceeds the area budget, fall back to a per-cluster shared unit (e.g., 1 MUL/DIV per 4 or 8 cores) behind a dedicated interconnect port. +- The crossbar/coherence documents in the repository should be checked before committing to a per-core vs. shared decision. +- A shared unit complicates the scoreboard: the issue queue must track remote MUL/DIV latency, which can be 4–8× the per-core MUL latency. + +## 128-Core Scalability +- Per-core area: Option A scales best; Option B is the worst case. +- Frequency: Option A and Option C both close timing at typical small-core targets; Option B is the only option likely to force a slower core clock. +- Yield: small per-core datapath → higher core yield, fewer fatal defects per die. +- Coherence traffic: a long-latency MUL/DIV keeps the core stalled but does not generate coherence traffic; an off-core MUL/DIV (shared) does, because the issuing core may continue past the result and the unit must return through the interconnect. +- Verification: per-core unit × 128 = 128 instances × one verification suite. This is a strong argument for the simplest microarchitecture that meets the latency target. + +## Performance Considerations +- For most non-numeric workloads, MUL/DIV throughput does not bound IPC; latency does not either, because the operand is rarely on the critical path. A 3–4 cycle MUL and a 32–64 cycle DIV is acceptable. +- For numeric, crypto, and hash kernels, MUL throughput is critical; the pipelined-tree options (C, D) win by roughly 2–3× peak throughput. +- DIV is rarely on the hot path; a 32–64 cycle iterative divider is almost always sufficient. + +RECOMMENDATION (conditional): If XH-1 targets general-purpose + occasional numeric, prefer Option A. If XH-1 targets numeric/CV/crypto explicitly, prefer Option C. + +## Area Considerations +- A 64×64→128 array multiplier at a modern node is roughly the size of the register file's writeback port plus the ALU; this is the dominant per-core cost in Option B. +- A pipelined 3-stage tree (Option C) shrinks the per-stage critical path at the cost of three sets of partial-product reduction and accumulation latches. +- An iterative multiplier (Option A) is dominated by a single 33–64-bit adder and a shift register, and is the smallest. +- PROPOSAL: Capture the per-option area in the implementation document once the target node is fixed. Until then, treat the "relative" column in the comparison table as the working estimate. + +## Power and Energy Considerations +- Combinational array multipliers (Option B) toggle the entire partial-product array every cycle; energy per MUL is highest. +- Pipelined tree multipliers (Option C) distribute the switching across pipeline registers, lowering per-cycle peak power but with similar energy per operation. +- Iterative shift-and-add (Option A) is the lowest per-operation energy because the active datapath per cycle is small (one adder stage). +- 128-core replication: peak power is the product of per-core dynamic power and number of active cores. Option A keeps per-core power lowest, which is the most important axis for a 128-core envelope. + +## Implementation Considerations +- Sign handling: MULHSU requires signed×unsigned with full sign extension. The most common bug source in MUL/DIV is signed/unsigned mode selection; design the control path so that mode bits are sourced from the decoder, never from a sticky register. +- Zero-detection: DIV by zero must complete in bounded time and write a defined result. The IEEE / RISC-V rule is that DIV/REM by zero return −1 / the dividend. The unit must NOT trap on divide-by-zero; that is a software choice. +- Overflow: for DIV, the only overflow case is INT_MIN / −1. RISC-V mandates that this returns INT_MIN. The divider must detect this explicitly; iterative non-restoring dividers do, but SRT designs must be checked. +- Latency variability: DIV latency is data-dependent in non-restoring designs (it is fixed in SRT). If the pipeline assumes a fixed MUL/DIV latency, prefer an SRT divider or a constant-iteration iterative divider. +- Reset and scan: a 128-core replication multiplies the scan chain length. PROPOSAL: gate scan on the MUL/DIV unit to limit shift power, at the cost of reduced fault coverage. This is a verification trade-off. + +## Verification Considerations +- The MUL/DIV unit has the highest ratio of corner cases to lines of RTL of any execution unit. Famous verification pitfalls: + - Signed multiplication: 0 × INT_MIN, INT_MIN × INT_MIN. + - MULH overflow into the high half with sign extension. + - MULHSU sign/unsigned mixing. + - DIV by zero, REM by zero. + - DIV overflow (INT_MIN / −1, INT_MIN % −1). + - All 16 combinations of signed/unsigned × 4 ops × 2 widths on RV64. +- PROPOSAL: Maintain a directed-test corpus at the unit level that exhausts the (sign × op × corner-operand) matrix, and a constrained-random suite at the core level. +- ASSUMPTION: XH-1 follows the riscv-formal convention of writing the MUL/DIV shadow model in a functional language. If so, the shadow model is the most expensive deliverable in the verification flow and is the strongest argument for picking the simplest microarchitecture. + +## Software Considerations +- The compiler's latency model for MUL/DIV must match the hardware's; otherwise the scheduler will insert unnecessary stalls or miss scheduling opportunities. If Option A is chosen, the compiler should treat MUL as a 3–4 cycle latency op and DIV as a 32–64 cycle latency op. +- If MUL/DIV is moved off-core (shared), the toolchain needs an intrinsic or scheduling model that reflects the round-trip latency through the interconnect. This is non-trivial; most toolchains do not model non-uniform functional unit latency across cores. +- The Linux kernel's alternatives patching and the C library's soft-float paths sometimes use MUL/DIV in hot paths; verify that the chosen unit's latency is acceptable for the kernel configurations that the project intends to boot. + +## Recommendation +PROPOSAL: For XH-1's first-pass implementation, adopt **Option A** (iterative Booth-2 multiplier pipelined to 1 result/cycle, non-restoring divider at 1 bit/cycle) for the following reasons: +1. Per-core area is minimized, which is the dominant axis in a 128-core replication. +2. The verification footprint is the smallest among the viable options. +3. The latency/throughput profile is acceptable for the majority of non-numeric workloads. +4. The microarchitecture is the easiest to ECO and to port across nodes. + +PROPOSAL: If benchmarks (TBD) show that MUL/DIV throughput is on the critical path, migrate to **Option C** (pipelined 3-stage Booth/Wallace tree + SRT-4 divider) as a follow-on revision. Reserve **Option B** (fully combinational) for a hypothetical single-core "XH-1 Big" variant where replication is not the binding constraint. + +ASSUMPTION: This recommendation is contingent on the pipeline depth, target clock, and per-core area budget not being fixed by other documents in the repository. If, for example, the pipeline document commits to a 5+ GHz target, Option A may not close timing and Option C becomes the floor rather than the upgrade. + +## Confidence +- Low-to-medium on the relative area numbers; they are order-of-magnitude estimates from published small-core work, not XH-1 synthesis. +- High on the qualitative trade-offs (area, verification complexity, replication cost). +- Low on the latency numbers, which depend on the target node and the adder architecture that is not yet specified in `alu.md` or `datapath.md`. + +## Open Questions +- What is the target node and the target core clock? This gates the latency-vs-area decision. +- Is XH-1 RV32 or RV64? If RV32, the 32×32→64 multiplier is dramatically smaller. +- Is the M extension mandatory for the first tape-out, or can it be gated behind a config bit? +- Is there a plan for a shared per-cluster MAC/vector unit that subsumes the MUL/DIV? If so, the per-core MUL/DIV may be reduced to a minimal subset. +- What is the pipeline depth (see `pipeline.md`, currently SOON)? This determines whether a 1-cycle, 3-cycle, or N-cycle MUL fits cleanly into the execute stage. +- What is the issue model — in-order (single-issue, dual-issue, etc.)? The structural hazard model for MUL/DIV depends on it. +- Does the project plan to use `riscv-formal` for verification, or a UVM/constrained-random flow? The two flows have different optimal unit-level architectures. +- What are the integer workload targets — server, embedded, HPC, ML inference? Each implies a different MUL/DIV throughput requirement. + +## Sources +- The RISC-V "Unprivileged ISA" specification, M-extension chapter — INSUFFICIENT EVIDENCE to cite a specific revision without checking the repository. The relevant normative requirements are: MUL/MULH/MULHU/MULHSU produce full 2X-bit results, DIV/DIVU/REM/REMU round toward zero, DIV/REM by zero write back defined results, and INT_MIN / −1 is defined. +- Canonical microarchitecture textbooks (Hennessy & Patterson; Parhami; Ercegovac & Lang) — referenced only at the level of well-known technique names (Booth, Wallace, Dadda, SRT, non-restoring division). No specific page or claim is attributed. +- INSUFFICIENT EVIDENCE: any XH-1-specific synthesis, layout, or benchmark data. The comparison table is built from public-domain technique properties and must be re-validated with synthesis at the target node before tape-out. diff --git a/research/.xh1/runs/20260825T180900Z/research_03-core-design_mul-div-unit.md.review.md.json b/research/.xh1/runs/20260825T180900Z/research_03-core-design_mul-div-unit.md.review.md.json new file mode 100644 index 0000000..e69de29 diff --git a/research/.xh1/runs/20260825T180900Z/research_03-core-design_mul-div-unit.md.round1.md b/research/.xh1/runs/20260825T180900Z/research_03-core-design_mul-div-unit.md.round1.md new file mode 100644 index 0000000..3217b20 --- /dev/null +++ b/research/.xh1/runs/20260825T180900Z/research_03-core-design_mul-div-unit.md.round1.md @@ -0,0 +1,194 @@ +# Multiplier–Divider Unit for the XH-1 Core + +## Status +DRAFT — Engineering research, no silicon data. + +## Abstract +This document investigates the design of an integer multiply/divide (MUL/DIV) unit suitable for one tile of the XH-1 128-core RISC-V processor. It frames the problem space (RV32M/RV64M, latency vs. throughput, area and energy), surveys the canonical microarchitectural options (iterative shift-and-add, array multipliers, Booth/Wallace/Dadda trees, reciprocal/sRT dividers, radix-2/4/8 dividers, and merged MAC-fused units), and maps them against the constraints implied by the other XH-1 core documents (small per-core budget, 128-core replication, expected integration with the pipeline, register file, and load–store unit). The document is written so that downstream choices about pipeline depth, core clock target, and shared vs. private FP/MAC can be made coherently with the MUL/DIV decision. + +## Research Question +What microarchitecture for the integer multiply/divide unit best fits one XH-1 core, given that the core will be replicated 128 times, must support the M extension on RV32 and/or RV64, and must coexist with a pipeline, register file, and load–store unit whose budgets are not yet fixed? + +Sub-questions: +1. Should the unit be fully combinatorial, multi-cycle iterative, or pipelined? +2. Which radix and which encoding (Booth-2, Booth-3, modified Booth-4) is appropriate for the expected operand width? +3. Should MUL and DIV share silicon, or be separate datapaths? +4. How does the choice interact with 128-core replication (area amortization, frequency, yield)? +5. What are the verification implications of each choice? + +## Background +The RISC-V M extension defines four signed/unsigned multiply variants producing 2X-bit results (MUL, MULH, MULHU, MULHSU), four matching multiply-high variants, and signed/unsigned divide and remainder (DIV, DIVU, REM, REMU). On RV64 there is an additional MULW/DIVW/REM family that operates on 32-bit values and sign-extends. + +Key microarchitectural properties that drive the design: +- Latency-tolerance: in a deep pipeline, multi-cycle iterative units are acceptable as long as the structural hazard is bounded and the scoreboard/issue logic handles in-flight multiplies. +- Throughput: the M extension opcodes are infrequent in many workloads but dominate in linear-algebra kernels, crypto, and hash functions. +- Operand width: an RV64 core needs both 64-bit and 32-bit paths. A 64×64→128 multiplier is roughly 4× the area of a 32×32→64 multiplier. +- Result width: MULH-family instructions require the full 2X-bit result; MUL only requires the low X bits. Most designs share the upper datapath and select the low half. + +ASSUMPTION: XH-1 cores implement the M extension. If only the I extension is required, the MUL/DIV unit collapses dramatically. The remainder of this document assumes M is in scope. +ASSUMPTION: XH-1 is RV64. If RV32 only, all 64-bit-specific considerations below can be relaxed. + +## Existing Approaches + +### Multipliers +1. Iterative shift-and-add multiplier +2. Array (braid) multiplier +3. Wallace tree multiplier +4. Dadda tree multiplier +5. Booth-encoded tree multiplier (radix-4, modified Booth) +6. Higher-radix Booth (radix-8/16) with 4:2/5:2 compressor trees +7. Pipelined versions of any of the above, with stage counts from 1 to N + +### Dividers +1. Restoring divider +2. Non-restoring divider +3. SRT divider (s radix-2, radix-4, radix-8, radix-16) +4. Newton–Raphson reciprocal + multiply (software or hardware) +5. Goldschmidt divider +6. Digit-recurrence with prescaling (for faster convergence) +7. Lookup-table (LUT) assisted dividers (small LUT, big LUT) + +PROPOSAL: Classify each option along four axes — latency (cycles), throughput (1/n per cycle), area (gate equivalent or µm² estimate), and design complexity (verification effort, corner cases). + +## Alternative Designs + +### Option A: Pipelined iterative multiplier + iterative non-restoring divider +- One shared 64-bit datapath, Booth-2 radix-4, latency ≈ 3–4 cycles for MUL, fully pipelined at 1 result/cycle. +- Divider: non-restoring, ≈ 32–64 cycles for RV64 DIV, variable. +- Area: smallest among the options; one of the smallest in published small-core implementations. +- Verification: standard; iterative state machines are well understood. + +### Option B: Fully combinational 64×64 array multiplier + SRT-4 divider +- MUL latency: 1 cycle, but very high area and long critical path. +- DIV: SRT-4, ≈ 16 cycles typical. +- Area: largest. +- Verification: simple timing closure problem, but 128 cores × this area is likely prohibitive. + +### Option C: Pipelined Wallace/Booth tree (3-stage) + SRT-4 divider +- MUL: 3-cycle pipelined, 1 result/cycle sustained. +- DIV: 16–20 cycles, fully pipelined divider. +- Area: moderate to large; tree is irregular. +- Verification: irregular partial-product reduction is harder to verify than array. + +### Option D: Fused multiply–accumulate (MAC) with Booth-3 (radix-8) and 4:2 compressors +- MUL: 3-cycle, also accepts an accumulate operand each cycle. +- Provides a useful primitive for dot-product kernels and soft-fp libraries. +- Divider: separate iterative non-restoring path. +- Area: similar to C, but the accumulator latch and bypass network add cost. + +### Option E: Truncated 32×32→64 only, with MULHW/DIVW synthesized in microcode +- Illegal in general: software synthesis of MULH is too slow to be acceptable. +- Listed for completeness and to discard. + +## Comparison + +| Option | MUL latency | MUL throughput | DIV latency (RV64) | Area (relative) | Design complexity | +|--------|-------------|----------------|---------------------|------------------|--------------------| +| A — Iterative Booth + NR div | 3–4 cyc | 1/cyc pipelined | 32–64 cyc | 0.6–0.8× | Low | +| B — Array + SRT-4 | 1 cyc | 1/cyc | 16 cyc | 2.5–3.0× | Medium (timing) | +| C — Wallace/Booth + SRT-4 | 3 cyc | 1/cyc | 16–20 cyc | 1.0–1.2× | High (tree) | +| D — MAC (Booth-3, 4:2) + NR div | 3 cyc | 1/cyc (with acc) | 32–64 cyc | 1.1–1.3× | High | +| E — Truncated, microcoded | n/a | n/a | n/a | smallest | n/a (incomplete) | + +Numbers above are ORDER-OF-MAGNITUDE ESTIMATES based on published small-core implementations; INSUFFICIENT EVIDENCE exists to claim a specific gate count or µm² without synthesis at a known target node. + +## Advantages +- Option A: smallest area per core → most 128-core replication headroom; trivially fits any pipeline depth; well-understood verification. +- Option B: best single-cycle latency; useful for OoO cores with tight issue windows. +- Option C: good balance of latency and area; widely used in commercial OoO cores. +- Option D: enables efficient dot-product and SIMD-style software; helps a soft FP stack. + +## Disadvantages +- Option A: lower peak MUL throughput than pipelined trees; multi-cycle DIV may stall the issue queue on back-to-back DIVs. +- Option B: area is prohibitive at 128 cores; long critical path forces a slow core clock or deep pipelining (defeating the latency benefit). +- Option C: Wallace/Dadda partial-product reduction has irregular carry-save structures that are harder to formally verify and harder to fix in ECO. +- Option D: accumulator adds bypass and forwarding complexity into the pipeline and register-file writeback path; benefits only workloads that can be rewritten to use the MAC. +- Option E: not viable for an M-extension-compliant core. + +## XH-1 Considerations +ASSUMPTION: XH-1 is a tiled, replicated design where per-core area is a first-class constraint because the whole 128-core array must fit in the package, power, and yield envelope. A small per-core MUL/DIV is therefore a high-value design point. + +PROPOSAL: Treat MUL/DIV as one of the "shared-tile resource" candidates. Specifically: +- If 128 cores × 1 MUL/DIV per core exceeds the area budget, fall back to a per-cluster shared unit (e.g., 1 MUL/DIV per 4 or 8 cores) behind a dedicated interconnect port. +- The crossbar/coherence documents in the repository should be checked before committing to a per-core vs. shared decision. +- A shared unit complicates the scoreboard: the issue queue must track remote MUL/DIV latency, which can be 4–8× the per-core MUL latency. + +## 128-Core Scalability +- Per-core area: Option A scales best; Option B is the worst case. +- Frequency: Option A and Option C both close timing at typical small-core targets; Option B is the only option likely to force a slower core clock. +- Yield: small per-core datapath → higher core yield, fewer fatal defects per die. +- Coherence traffic: a long-latency MUL/DIV keeps the core stalled but does not generate coherence traffic; an off-core MUL/DIV (shared) does, because the issuing core may continue past the result and the unit must return through the interconnect. +- Verification: per-core unit × 128 = 128 instances × one verification suite. This is a strong argument for the simplest microarchitecture that meets the latency target. + +## Performance Considerations +- For most non-numeric workloads, MUL/DIV throughput does not bound IPC; latency does not either, because the operand is rarely on the critical path. A 3–4 cycle MUL and a 32–64 cycle DIV is acceptable. +- For numeric, crypto, and hash kernels, MUL throughput is critical; the pipelined-tree options (C, D) win by roughly 2–3× peak throughput. +- DIV is rarely on the hot path; a 32–64 cycle iterative divider is almost always sufficient. + +RECOMMENDATION (conditional): If XH-1 targets general-purpose + occasional numeric, prefer Option A. If XH-1 targets numeric/CV/crypto explicitly, prefer Option C. + +## Area Considerations +- A 64×64→128 array multiplier at a modern node is roughly the size of the register file's writeback port plus the ALU; this is the dominant per-core cost in Option B. +- A pipelined 3-stage tree (Option C) shrinks the per-stage critical path at the cost of three sets of partial-product reduction and accumulation latches. +- An iterative multiplier (Option A) is dominated by a single 33–64-bit adder and a shift register, and is the smallest. +- PROPOSAL: Capture the per-option area in the implementation document once the target node is fixed. Until then, treat the "relative" column in the comparison table as the working estimate. + +## Power and Energy Considerations +- Combinational array multipliers (Option B) toggle the entire partial-product array every cycle; energy per MUL is highest. +- Pipelined tree multipliers (Option C) distribute the switching across pipeline registers, lowering per-cycle peak power but with similar energy per operation. +- Iterative shift-and-add (Option A) is the lowest per-operation energy because the active datapath per cycle is small (one adder stage). +- 128-core replication: peak power is the product of per-core dynamic power and number of active cores. Option A keeps per-core power lowest, which is the most important axis for a 128-core envelope. + +## Implementation Considerations +- Sign handling: MULHSU requires signed×unsigned with full sign extension. The most common bug source in MUL/DIV is signed/unsigned mode selection; design the control path so that mode bits are sourced from the decoder, never from a sticky register. +- Zero-detection: DIV by zero must complete in bounded time and write a defined result. The IEEE / RISC-V rule is that DIV/REM by zero return −1 / the dividend. The unit must NOT trap on divide-by-zero; that is a software choice. +- Overflow: for DIV, the only overflow case is INT_MIN / −1. RISC-V mandates that this returns INT_MIN. The divider must detect this explicitly; iterative non-restoring dividers do, but SRT designs must be checked. +- Latency variability: DIV latency is data-dependent in non-restoring designs (it is fixed in SRT). If the pipeline assumes a fixed MUL/DIV latency, prefer an SRT divider or a constant-iteration iterative divider. +- Reset and scan: a 128-core replication multiplies the scan chain length. PROPOSAL: gate scan on the MUL/DIV unit to limit shift power, at the cost of reduced fault coverage. This is a verification trade-off. + +## Verification Considerations +- The MUL/DIV unit has the highest ratio of corner cases to lines of RTL of any execution unit. Famous verification pitfalls: + - Signed multiplication: 0 × INT_MIN, INT_MIN × INT_MIN. + - MULH overflow into the high half with sign extension. + - MULHSU sign/unsigned mixing. + - DIV by zero, REM by zero. + - DIV overflow (INT_MIN / −1, INT_MIN % −1). + - All 16 combinations of signed/unsigned × 4 ops × 2 widths on RV64. +- PROPOSAL: Maintain a directed-test corpus at the unit level that exhausts the (sign × op × corner-operand) matrix, and a constrained-random suite at the core level. +- ASSUMPTION: XH-1 follows the riscv-formal convention of writing the MUL/DIV shadow model in a functional language. If so, the shadow model is the most expensive deliverable in the verification flow and is the strongest argument for picking the simplest microarchitecture. + +## Software Considerations +- The compiler's latency model for MUL/DIV must match the hardware's; otherwise the scheduler will insert unnecessary stalls or miss scheduling opportunities. If Option A is chosen, the compiler should treat MUL as a 3–4 cycle latency op and DIV as a 32–64 cycle latency op. +- If MUL/DIV is moved off-core (shared), the toolchain needs an intrinsic or scheduling model that reflects the round-trip latency through the interconnect. This is non-trivial; most toolchains do not model non-uniform functional unit latency across cores. +- The Linux kernel's alternatives patching and the C library's soft-float paths sometimes use MUL/DIV in hot paths; verify that the chosen unit's latency is acceptable for the kernel configurations that the project intends to boot. + +## Recommendation +PROPOSAL: For XH-1's first-pass implementation, adopt **Option A** (iterative Booth-2 multiplier pipelined to 1 result/cycle, non-restoring divider at 1 bit/cycle) for the following reasons: +1. Per-core area is minimized, which is the dominant axis in a 128-core replication. +2. The verification footprint is the smallest among the viable options. +3. The latency/throughput profile is acceptable for the majority of non-numeric workloads. +4. The microarchitecture is the easiest to ECO and to port across nodes. + +PROPOSAL: If benchmarks (TBD) show that MUL/DIV throughput is on the critical path, migrate to **Option C** (pipelined 3-stage Booth/Wallace tree + SRT-4 divider) as a follow-on revision. Reserve **Option B** (fully combinational) for a hypothetical single-core "XH-1 Big" variant where replication is not the binding constraint. + +ASSUMPTION: This recommendation is contingent on the pipeline depth, target clock, and per-core area budget not being fixed by other documents in the repository. If, for example, the pipeline document commits to a 5+ GHz target, Option A may not close timing and Option C becomes the floor rather than the upgrade. + +## Confidence +- Low-to-medium on the relative area numbers; they are order-of-magnitude estimates from published small-core work, not XH-1 synthesis. +- High on the qualitative trade-offs (area, verification complexity, replication cost). +- Low on the latency numbers, which depend on the target node and the adder architecture that is not yet specified in `alu.md` or `datapath.md`. + +## Open Questions +- What is the target node and the target core clock? This gates the latency-vs-area decision. +- Is XH-1 RV32 or RV64? If RV32, the 32×32→64 multiplier is dramatically smaller. +- Is the M extension mandatory for the first tape-out, or can it be gated behind a config bit? +- Is there a plan for a shared per-cluster MAC/vector unit that subsumes the MUL/DIV? If so, the per-core MUL/DIV may be reduced to a minimal subset. +- What is the pipeline depth (see `pipeline.md`, currently SOON)? This determines whether a 1-cycle, 3-cycle, or N-cycle MUL fits cleanly into the execute stage. +- What is the issue model — in-order (single-issue, dual-issue, etc.)? The structural hazard model for MUL/DIV depends on it. +- Does the project plan to use `riscv-formal` for verification, or a UVM/constrained-random flow? The two flows have different optimal unit-level architectures. +- What are the integer workload targets — server, embedded, HPC, ML inference? Each implies a different MUL/DIV throughput requirement. + +## Sources +- The RISC-V "Unprivileged ISA" specification, M-extension chapter — INSUFFICIENT EVIDENCE to cite a specific revision without checking the repository. The relevant normative requirements are: MUL/MULH/MULHU/MULHSU produce full 2X-bit results, DIV/DIVU/REM/REMU round toward zero, DIV/REM by zero write back defined results, and INT_MIN / −1 is defined. +- Canonical microarchitecture textbooks (Hennessy & Patterson; Parhami; Ercegovac & Lang) — referenced only at the level of well-known technique names (Booth, Wallace, Dadda, SRT, non-restoring division). No specific page or claim is attributed. +- INSUFFICIENT EVIDENCE: any XH-1-specific synthesis, layout, or benchmark data. The comparison table is built from public-domain technique properties and must be re-validated with synthesis at the target node before tape-out. diff --git a/research/.xh1/runs/20260825T180900Z/research_03-core-design_mul-div-unit.md.round1.md.json b/research/.xh1/runs/20260825T180900Z/research_03-core-design_mul-div-unit.md.round1.md.json new file mode 100644 index 0000000..64fe9bb --- /dev/null +++ b/research/.xh1/runs/20260825T180900Z/research_03-core-design_mul-div-unit.md.round1.md.json @@ -0,0 +1,263 @@ + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +{"id":"gen-1787681341-c8aqJM9lDIy4RiIbyGD9","object":"chat.completion","created":1787681341,"model":"minimax/minimax-m3:free","provider":"GMICloud","system_fingerprint":null,"service_tier":null,"choices":[{"index":0,"logprobs":null,"finish_reason":"stop","native_finish_reason":"stop","message":{"role":"assistant","content":"# Multiplier–Divider Unit for the XH-1 Core\n\n## Status\nDRAFT — Engineering research, no silicon data.\n\n## Abstract\nThis document investigates the design of an integer multiply/divide (MUL/DIV) unit suitable for one tile of the XH-1 128-core RISC-V processor. It frames the problem space (RV32M/RV64M, latency vs. throughput, area and energy), surveys the canonical microarchitectural options (iterative shift-and-add, array multipliers, Booth/Wallace/Dadda trees, reciprocal/sRT dividers, radix-2/4/8 dividers, and merged MAC-fused units), and maps them against the constraints implied by the other XH-1 core documents (small per-core budget, 128-core replication, expected integration with the pipeline, register file, and load–store unit). The document is written so that downstream choices about pipeline depth, core clock target, and shared vs. private FP/MAC can be made coherently with the MUL/DIV decision.\n\n## Research Question\nWhat microarchitecture for the integer multiply/divide unit best fits one XH-1 core, given that the core will be replicated 128 times, must support the M extension on RV32 and/or RV64, and must coexist with a pipeline, register file, and load–store unit whose budgets are not yet fixed?\n\nSub-questions:\n1. Should the unit be fully combinatorial, multi-cycle iterative, or pipelined?\n2. Which radix and which encoding (Booth-2, Booth-3, modified Booth-4) is appropriate for the expected operand width?\n3. Should MUL and DIV share silicon, or be separate datapaths?\n4. How does the choice interact with 128-core replication (area amortization, frequency, yield)?\n5. What are the verification implications of each choice?\n\n## Background\nThe RISC-V M extension defines four signed/unsigned multiply variants producing 2X-bit results (MUL, MULH, MULHU, MULHSU), four matching multiply-high variants, and signed/unsigned divide and remainder (DIV, DIVU, REM, REMU). On RV64 there is an additional MULW/DIVW/REM family that operates on 32-bit values and sign-extends.\n\nKey microarchitectural properties that drive the design:\n- Latency-tolerance: in a deep pipeline, multi-cycle iterative units are acceptable as long as the structural hazard is bounded and the scoreboard/issue logic handles in-flight multiplies.\n- Throughput: the M extension opcodes are infrequent in many workloads but dominate in linear-algebra kernels, crypto, and hash functions.\n- Operand width: an RV64 core needs both 64-bit and 32-bit paths. A 64×64→128 multiplier is roughly 4× the area of a 32×32→64 multiplier.\n- Result width: MULH-family instructions require the full 2X-bit result; MUL only requires the low X bits. Most designs share the upper datapath and select the low half.\n\nASSUMPTION: XH-1 cores implement the M extension. If only the I extension is required, the MUL/DIV unit collapses dramatically. The remainder of this document assumes M is in scope.\nASSUMPTION: XH-1 is RV64. If RV32 only, all 64-bit-specific considerations below can be relaxed.\n\n## Existing Approaches\n\n### Multipliers\n1. Iterative shift-and-add multiplier\n2. Array (braid) multiplier\n3. Wallace tree multiplier\n4. Dadda tree multiplier\n5. Booth-encoded tree multiplier (radix-4, modified Booth)\n6. Higher-radix Booth (radix-8/16) with 4:2/5:2 compressor trees\n7. Pipelined versions of any of the above, with stage counts from 1 to N\n\n### Dividers\n1. Restoring divider\n2. Non-restoring divider\n3. SRT divider (s radix-2, radix-4, radix-8, radix-16)\n4. Newton–Raphson reciprocal + multiply (software or hardware)\n5. Goldschmidt divider\n6. Digit-recurrence with prescaling (for faster convergence)\n7. Lookup-table (LUT) assisted dividers (small LUT, big LUT)\n\nPROPOSAL: Classify each option along four axes — latency (cycles), throughput (1/n per cycle), area (gate equivalent or µm² estimate), and design complexity (verification effort, corner cases).\n\n## Alternative Designs\n\n### Option A: Pipelined iterative multiplier + iterative non-restoring divider\n- One shared 64-bit datapath, Booth-2 radix-4, latency ≈ 3–4 cycles for MUL, fully pipelined at 1 result/cycle.\n- Divider: non-restoring, ≈ 32–64 cycles for RV64 DIV, variable.\n- Area: smallest among the options; one of the smallest in published small-core implementations.\n- Verification: standard; iterative state machines are well understood.\n\n### Option B: Fully combinational 64×64 array multiplier + SRT-4 divider\n- MUL latency: 1 cycle, but very high area and long critical path.\n- DIV: SRT-4, ≈ 16 cycles typical.\n- Area: largest.\n- Verification: simple timing closure problem, but 128 cores × this area is likely prohibitive.\n\n### Option C: Pipelined Wallace/Booth tree (3-stage) + SRT-4 divider\n- MUL: 3-cycle pipelined, 1 result/cycle sustained.\n- DIV: 16–20 cycles, fully pipelined divider.\n- Area: moderate to large; tree is irregular.\n- Verification: irregular partial-product reduction is harder to verify than array.\n\n### Option D: Fused multiply–accumulate (MAC) with Booth-3 (radix-8) and 4:2 compressors\n- MUL: 3-cycle, also accepts an accumulate operand each cycle.\n- Provides a useful primitive for dot-product kernels and soft-fp libraries.\n- Divider: separate iterative non-restoring path.\n- Area: similar to C, but the accumulator latch and bypass network add cost.\n\n### Option E: Truncated 32×32→64 only, with MULHW/DIVW synthesized in microcode\n- Illegal in general: software synthesis of MULH is too slow to be acceptable.\n- Listed for completeness and to discard.\n\n## Comparison\n\n| Option | MUL latency | MUL throughput | DIV latency (RV64) | Area (relative) | Design complexity |\n|--------|-------------|----------------|---------------------|------------------|--------------------|\n| A — Iterative Booth + NR div | 3–4 cyc | 1/cyc pipelined | 32–64 cyc | 0.6–0.8× | Low |\n| B — Array + SRT-4 | 1 cyc | 1/cyc | 16 cyc | 2.5–3.0× | Medium (timing) |\n| C — Wallace/Booth + SRT-4 | 3 cyc | 1/cyc | 16–20 cyc | 1.0–1.2× | High (tree) |\n| D — MAC (Booth-3, 4:2) + NR div | 3 cyc | 1/cyc (with acc) | 32–64 cyc | 1.1–1.3× | High |\n| E — Truncated, microcoded | n/a | n/a | n/a | smallest | n/a (incomplete) |\n\nNumbers above are ORDER-OF-MAGNITUDE ESTIMATES based on published small-core implementations; INSUFFICIENT EVIDENCE exists to claim a specific gate count or µm² without synthesis at a known target node.\n\n## Advantages\n- Option A: smallest area per core → most 128-core replication headroom; trivially fits any pipeline depth; well-understood verification.\n- Option B: best single-cycle latency; useful for OoO cores with tight issue windows.\n- Option C: good balance of latency and area; widely used in commercial OoO cores.\n- Option D: enables efficient dot-product and SIMD-style software; helps a soft FP stack.\n\n## Disadvantages\n- Option A: lower peak MUL throughput than pipelined trees; multi-cycle DIV may stall the issue queue on back-to-back DIVs.\n- Option B: area is prohibitive at 128 cores; long critical path forces a slow core clock or deep pipelining (defeating the latency benefit).\n- Option C: Wallace/Dadda partial-product reduction has irregular carry-save structures that are harder to formally verify and harder to fix in ECO.\n- Option D: accumulator adds bypass and forwarding complexity into the pipeline and register-file writeback path; benefits only workloads that can be rewritten to use the MAC.\n- Option E: not viable for an M-extension-compliant core.\n\n## XH-1 Considerations\nASSUMPTION: XH-1 is a tiled, replicated design where per-core area is a first-class constraint because the whole 128-core array must fit in the package, power, and yield envelope. A small per-core MUL/DIV is therefore a high-value design point.\n\nPROPOSAL: Treat MUL/DIV as one of the \"shared-tile resource\" candidates. Specifically:\n- If 128 cores × 1 MUL/DIV per core exceeds the area budget, fall back to a per-cluster shared unit (e.g., 1 MUL/DIV per 4 or 8 cores) behind a dedicated interconnect port.\n- The crossbar/coherence documents in the repository should be checked before committing to a per-core vs. shared decision.\n- A shared unit complicates the scoreboard: the issue queue must track remote MUL/DIV latency, which can be 4–8× the per-core MUL latency.\n\n## 128-Core Scalability\n- Per-core area: Option A scales best; Option B is the worst case.\n- Frequency: Option A and Option C both close timing at typical small-core targets; Option B is the only option likely to force a slower core clock.\n- Yield: small per-core datapath → higher core yield, fewer fatal defects per die.\n- Coherence traffic: a long-latency MUL/DIV keeps the core stalled but does not generate coherence traffic; an off-core MUL/DIV (shared) does, because the issuing core may continue past the result and the unit must return through the interconnect.\n- Verification: per-core unit × 128 = 128 instances × one verification suite. This is a strong argument for the simplest microarchitecture that meets the latency target.\n\n## Performance Considerations\n- For most non-numeric workloads, MUL/DIV throughput does not bound IPC; latency does not either, because the operand is rarely on the critical path. A 3–4 cycle MUL and a 32–64 cycle DIV is acceptable.\n- For numeric, crypto, and hash kernels, MUL throughput is critical; the pipelined-tree options (C, D) win by roughly 2–3× peak throughput.\n- DIV is rarely on the hot path; a 32–64 cycle iterative divider is almost always sufficient.\n\nRECOMMENDATION (conditional): If XH-1 targets general-purpose + occasional numeric, prefer Option A. If XH-1 targets numeric/CV/crypto explicitly, prefer Option C.\n\n## Area Considerations\n- A 64×64→128 array multiplier at a modern node is roughly the size of the register file's writeback port plus the ALU; this is the dominant per-core cost in Option B.\n- A pipelined 3-stage tree (Option C) shrinks the per-stage critical path at the cost of three sets of partial-product reduction and accumulation latches.\n- An iterative multiplier (Option A) is dominated by a single 33–64-bit adder and a shift register, and is the smallest.\n- PROPOSAL: Capture the per-option area in the implementation document once the target node is fixed. Until then, treat the \"relative\" column in the comparison table as the working estimate.\n\n## Power and Energy Considerations\n- Combinational array multipliers (Option B) toggle the entire partial-product array every cycle; energy per MUL is highest.\n- Pipelined tree multipliers (Option C) distribute the switching across pipeline registers, lowering per-cycle peak power but with similar energy per operation.\n- Iterative shift-and-add (Option A) is the lowest per-operation energy because the active datapath per cycle is small (one adder stage).\n- 128-core replication: peak power is the product of per-core dynamic power and number of active cores. Option A keeps per-core power lowest, which is the most important axis for a 128-core envelope.\n\n## Implementation Considerations\n- Sign handling: MULHSU requires signed×unsigned with full sign extension. The most common bug source in MUL/DIV is signed/unsigned mode selection; design the control path so that mode bits are sourced from the decoder, never from a sticky register.\n- Zero-detection: DIV by zero must complete in bounded time and write a defined result. The IEEE / RISC-V rule is that DIV/REM by zero return −1 / the dividend. The unit must NOT trap on divide-by-zero; that is a software choice.\n- Overflow: for DIV, the only overflow case is INT_MIN / −1. RISC-V mandates that this returns INT_MIN. The divider must detect this explicitly; iterative non-restoring dividers do, but SRT designs must be checked.\n- Latency variability: DIV latency is data-dependent in non-restoring designs (it is fixed in SRT). If the pipeline assumes a fixed MUL/DIV latency, prefer an SRT divider or a constant-iteration iterative divider.\n- Reset and scan: a 128-core replication multiplies the scan chain length. PROPOSAL: gate scan on the MUL/DIV unit to limit shift power, at the cost of reduced fault coverage. This is a verification trade-off.\n\n## Verification Considerations\n- The MUL/DIV unit has the highest ratio of corner cases to lines of RTL of any execution unit. Famous verification pitfalls:\n - Signed multiplication: 0 × INT_MIN, INT_MIN × INT_MIN.\n - MULH overflow into the high half with sign extension.\n - MULHSU sign/unsigned mixing.\n - DIV by zero, REM by zero.\n - DIV overflow (INT_MIN / −1, INT_MIN % −1).\n - All 16 combinations of signed/unsigned × 4 ops × 2 widths on RV64.\n- PROPOSAL: Maintain a directed-test corpus at the unit level that exhausts the (sign × op × corner-operand) matrix, and a constrained-random suite at the core level.\n- ASSUMPTION: XH-1 follows the riscv-formal convention of writing the MUL/DIV shadow model in a functional language. If so, the shadow model is the most expensive deliverable in the verification flow and is the strongest argument for picking the simplest microarchitecture.\n\n## Software Considerations\n- The compiler's latency model for MUL/DIV must match the hardware's; otherwise the scheduler will insert unnecessary stalls or miss scheduling opportunities. If Option A is chosen, the compiler should treat MUL as a 3–4 cycle latency op and DIV as a 32–64 cycle latency op.\n- If MUL/DIV is moved off-core (shared), the toolchain needs an intrinsic or scheduling model that reflects the round-trip latency through the interconnect. This is non-trivial; most toolchains do not model non-uniform functional unit latency across cores.\n- The Linux kernel's alternatives patching and the C library's soft-float paths sometimes use MUL/DIV in hot paths; verify that the chosen unit's latency is acceptable for the kernel configurations that the project intends to boot.\n\n## Recommendation\nPROPOSAL: For XH-1's first-pass implementation, adopt **Option A** (iterative Booth-2 multiplier pipelined to 1 result/cycle, non-restoring divider at 1 bit/cycle) for the following reasons:\n1. Per-core area is minimized, which is the dominant axis in a 128-core replication.\n2. The verification footprint is the smallest among the viable options.\n3. The latency/throughput profile is acceptable for the majority of non-numeric workloads.\n4. The microarchitecture is the easiest to ECO and to port across nodes.\n\nPROPOSAL: If benchmarks (TBD) show that MUL/DIV throughput is on the critical path, migrate to **Option C** (pipelined 3-stage Booth/Wallace tree + SRT-4 divider) as a follow-on revision. Reserve **Option B** (fully combinational) for a hypothetical single-core \"XH-1 Big\" variant where replication is not the binding constraint.\n\nASSUMPTION: This recommendation is contingent on the pipeline depth, target clock, and per-core area budget not being fixed by other documents in the repository. If, for example, the pipeline document commits to a 5+ GHz target, Option A may not close timing and Option C becomes the floor rather than the upgrade.\n\n## Confidence\n- Low-to-medium on the relative area numbers; they are order-of-magnitude estimates from published small-core work, not XH-1 synthesis.\n- High on the qualitative trade-offs (area, verification complexity, replication cost).\n- Low on the latency numbers, which depend on the target node and the adder architecture that is not yet specified in `alu.md` or `datapath.md`.\n\n## Open Questions\n- What is the target node and the target core clock? This gates the latency-vs-area decision.\n- Is XH-1 RV32 or RV64? If RV32, the 32×32→64 multiplier is dramatically smaller.\n- Is the M extension mandatory for the first tape-out, or can it be gated behind a config bit?\n- Is there a plan for a shared per-cluster MAC/vector unit that subsumes the MUL/DIV? If so, the per-core MUL/DIV may be reduced to a minimal subset.\n- What is the pipeline depth (see `pipeline.md`, currently SOON)? This determines whether a 1-cycle, 3-cycle, or N-cycle MUL fits cleanly into the execute stage.\n- What is the issue model — in-order (single-issue, dual-issue, etc.)? The structural hazard model for MUL/DIV depends on it.\n- Does the project plan to use `riscv-formal` for verification, or a UVM/constrained-random flow? The two flows have different optimal unit-level architectures.\n- What are the integer workload targets — server, embedded, HPC, ML inference? Each implies a different MUL/DIV throughput requirement.\n\n## Sources\n- The RISC-V \"Unprivileged ISA\" specification, M-extension chapter — INSUFFICIENT EVIDENCE to cite a specific revision without checking the repository. The relevant normative requirements are: MUL/MULH/MULHU/MULHSU produce full 2X-bit results, DIV/DIVU/REM/REMU round toward zero, DIV/REM by zero write back defined results, and INT_MIN / −1 is defined.\n- Canonical microarchitecture textbooks (Hennessy & Patterson; Parhami; Ercegovac & Lang) — referenced only at the level of well-known technique names (Booth, Wallace, Dadda, SRT, non-restoring division). No specific page or claim is attributed.\n- INSUFFICIENT EVIDENCE: any XH-1-specific synthesis, layout, or benchmark data. The comparison table is built from public-domain technique properties and must be re-validated with synthesis at the target node before tape-out.","refusal":null,"reasoning":null}}],"usage":{"prompt_tokens":706,"completion_tokens":4143,"total_tokens":4849,"cost":0,"is_byok":false,"prompt_tokens_details":{"cached_tokens":135,"cache_write_tokens":0,"audio_tokens":0,"video_tokens":0},"cost_details":{"upstream_inference_cost":0,"upstream_inference_prompt_cost":0,"upstream_inference_completions_cost":0},"completion_tokens_details":{"reasoning_tokens":0,"image_tokens":0,"audio_tokens":0}}} \ No newline at end of file diff --git a/research/.xh1/runs/20260825T181153Z/attempts b/research/.xh1/runs/20260825T181153Z/attempts new file mode 100644 index 0000000..01862ec --- /dev/null +++ b/research/.xh1/runs/20260825T181153Z/attempts @@ -0,0 +1,2 @@ +2026-08-25T18:13:01Z research/03-core-design/mul-div-unit.md 1 research completed +2026-08-25T18:13:16Z research/03-core-design/mul-div-unit.md 1 review api-failure diff --git a/research/.xh1/runs/20260825T181153Z/completed b/research/.xh1/runs/20260825T181153Z/completed new file mode 100644 index 0000000..e69de29 diff --git a/research/.xh1/runs/20260825T181153Z/failed b/research/.xh1/runs/20260825T181153Z/failed new file mode 100644 index 0000000..e427ae2 --- /dev/null +++ b/research/.xh1/runs/20260825T181153Z/failed @@ -0,0 +1 @@ +research/03-core-design/mul-div-unit.md diff --git a/research/.xh1/runs/20260825T181153Z/research_03-core-design_mul-div-unit.md.current.md b/research/.xh1/runs/20260825T181153Z/research_03-core-design_mul-div-unit.md.current.md new file mode 100644 index 0000000..a6e5db0 --- /dev/null +++ b/research/.xh1/runs/20260825T181153Z/research_03-core-design_mul-div-unit.md.current.md @@ -0,0 +1,248 @@ +# XH-1 Multiply/Divide Unit Research + +## Topic + +`research/03-core-design/mul-div-unit.md` — design of the integer multiply/divide (MUL/DIV) execution unit for the XH-1 core, with particular attention to how this unit is replicated across 128 cores. + +## Status + +**Stub.** The current document contains only the placeholder `SOON`. This research document establishes the design-space analysis required to populate that stub. No committed micro-architecture exists yet. + +## Abstract + +The integer multiply/divide unit (MDU) implements the RISC-V `M` extension (and optionally `B`, `K`, or vector variants) within each XH-1 core. This document surveys MDU micro-architectures — array multipliers, Wallace/Dadda trees, Booth-recoded multipliers, iterative shift-add multipliers, radix-2/radix-4/radix-8 dividers, and higher-radix/SRT dividers — and evaluates them against the XH-1 design constraints: per-core area and power budgets, latency targets, throughput requirements, verification complexity, and the multiplicative cost of replicating the unit across 128 cores. The analysis concludes that a parameterized iterative shift-add MDU with optional early-exit and a modest dedicated array multiplier is the most defensible starting point, but flags the assumption-laden nature of the recommendation. + +## Research Question + +> What MDU micro-architecture best satisfies the XH-1 per-core area, power, latency, and throughput targets while remaining tractable to verify, to replicate 128 times, and to expose to the RISC-V ISA and toolchain? + +Sub-questions: + +1. Should multiplication and division share datapath hardware or be independent units? +2. What latency is required to avoid becoming a structural hazard in the XH-1 pipeline? +3. What is the minimum-radix divider that meets the per-core throughput target? +4. Does the cost of a fast array multiplier (e.g., radix-4 Booth with Wallace tree) justify its latency benefit over an iterative shift-add design replicated 128 times? +5. Which RISC-V extensions must be supported in v1 (`M` only vs. `M+B` vs. `M+B+K`)? + +## Background + +### RISC-V `M` Extension Semantics (FACT) + +The RISC-V `M` extension defines the following operations on the `XLEN`-bit registers (RV64 assumed for XH-1 unless stated otherwise): + +- `MUL`, `MULH`, `MULHSU`, `MULHU` — 64×64 → lower-64 / upper-64 signed/unsigned products. +- `DIV`, `DIVU`, `REM`, `REMU` — signed/unsigned division and remainder, 64÷64 → 64 quotient and 64 remainder. +- `MULW`, `DIVW`, `DIVUW`, `REMW`, `REMWU` — 32-bit operations that sign-extend the 32-bit result to 64 bits. + +Corner cases the MDU must handle per the ISA: + +- Division by zero returns `0xFFFFFFFFFFFFFFFF` (quotient) and the dividend (remainder) for unsigned, and `-1` / dividend for signed. +- The most-negative signed integer (`0x8000…`) divided by `-1` must return the dividend as quotient and `0` as remainder; this traps on x86 but is non-trapping in RISC-V. +- Overflow behavior is *only* defined for signed division/remainder by `-1`; no other overflow is architecturally visible. + +### Latency vs. Throughput (FACT) + +- A 64×64 → 128-bit full multiplication requires at minimum 128 partial-product rows for a non-Booth radix-2 design, ~64 rows for radix-4 Booth, ~43 rows for radix-8 Booth. +- Bit-serial shift-add multiplication completes in 64 cycles. +- Restoring division completes in 64 cycles; non-restoring in 32–64 cycles depending on the normalization of the remainder; higher-radix SRT dividers in O(log radix(N)) cycles but with substantially larger area. + +### RISC-V `B` and `K` Extensions (FACT, partial) + +- `B` (Bitmanip) under RV64 1.0.0 includes `MULH`, `MULHU`, `MULHSU` already in `M`, and adds rev-8.0/1.0 bitmanip groups such as `CLMUL`, `CLMULH`, `CLMULR`, `MIN[U]`, `MAX[U]`, `ANDN`, `ORN`, `XNOR`, `SEXT.B/H/W`, and the `Zba`/`Zbb`/`Zbs` sub-extensions. +- `K` (cryptography) adds `Zkn` (AES/SHA) and is not in scope for the base MDU. +- `Zmmul` provides multiply-only and is a common power-optimized subset of `M`. + +## Existing Approaches + +| Approach | Latency (typ., 64-bit) | Area (relative) | Throughput | Notes | +|---|---|---|---|---| +| Iterative shift-add multiplier | 64 cycles | 1× | 1 / 64 cycles | Smallest area; low throughput | +| Radix-4 Booth + carry-save | ~17 cycles | ~4–6× | 1 / ~17 | Common in embedded OoO | +| Radix-4 Booth + Wallace tree + final CPA | ~3–5 cycles | ~8–12× | 1 / cycle (pipelined) | Used in high-perf cores | +| Pipelined array multiplier (k stages) | k cycles | large | 1 / cycle | Throughput-bound, area-heavy | +| Radix-2 non-restoring divider | 32–64 cycles | 1× (shared with MUL possible) | 1 / 32–64 | Low area | +| Radix-4 SRT divider | ~16 cycles | ~3–5× | 1 / ~16 | Moderate | +| Radix-8/16 SRT divider | ~8–10 cycles | large | 1 / ~10 | High-end OoO | +| Combinational 64×64 array (no pipelining) | ~10+ gate delays | very large | 1 / many cycles | Impractical | +| Shared MUL/DIV datapath (Liang/Montusiewicz style) | variable | ~1.5× single-function | variable | Saves ~30–40% area at the cost of CPI | + +(Quantitative magnitudes in this table are estimates from public micro-architecture literature; values vary considerably with cell library, target frequency, and pipeline integration. See Sources.) + +## Alternative Designs + +### A. Pure Iterative Shift-Add MDU + +A single 128-bit accumulator-based datapath. Each cycle shifts the multiplicand and adds if the corresponding multiplier bit is set, or performs a non-restoring division step. The same hardware services both MUL and DIV. + +### B. Pipelined Iterative MDU + +Same datapath as A, but cut into N pipeline stages (commonly 2 or 4) to allow clock-frequency scaling. Issues per-cycle throughput remains 1 op / N cycles for a single operand width. + +### C. Dedicated Array Multiplier + Iterative Divider + +A small radix-4 Booth/Wallace multiplier (lower 64 bits in 3–4 cycles) and a separate radix-2 or radix-4 non-restoring divider. Two independent units that can operate concurrently. + +### D. Shared Pipelined Radix-4 MDU + +A single radix-4 Booth partial-product generator feeding a shared CSA tree, with a reconfigurable final-stage adder and a small division state machine. Latency: ~5–8 cycles for MUL, ~16 for DIV. + +### E. Pipelined High-Performance MDU (Wallace/Wallace + SRT-4) + +Multi-cycle but pipelined radix-4 multiplier and SRT-4 divider. Both deliver 1 op/cycle throughput once steady-state. Area dominates the integer datapath. + +### F. `Zmmul`-Only MUL, External DIV Trap + +Implement MUL in hardware (small array or iterative), trap DIV/REM to a software handler. The privileged trap path is fast on XH-1 only if the software model tolerates it; SPECint and most server workloads execute division frequently enough that this is usually unacceptable. + +### G. Compressed / Shared Across Cores + +A single physical MDU shared by N cores via a NoC-side arbiter. Removes per-core area cost but introduces structural contention; this is generally considered an *adverse* design for a 128-core tiled machine. + +## Comparison + +| Design | Area/cell count (est.) | Cycles/op MUL | Cycles/op DIV | Throughput (steady) | Verif. complexity | Replicated ×128 cost | Power/cores at 1 GHz (est.) | +|---|---|---|---|---|---|---|---| +| A. Iterative | 1.0× | 64 | 32–64 | 1/64 | Low | low | low | +| B. Pipelined iter. (2-stage) | ~1.1× | 32 | 16–32 | 1/32 | Low–Medium | low | low–medium | +| C. Array MUL + iter DIV | ~6–8× | 3–5 | 32–64 | 1/cycle (MUL) | Medium | high | medium–high | +| D. Shared radix-4 | ~4–6× | 5–8 | 16 | 1/cycle (MUL) | Medium–High | medium | medium | +| E. Pipelined Wallace + SRT-4 | ~12–20× | 1/cycle | 1/cycle | 1/cycle | High | very high | high | +| F. Zmmul + trap | ~1–2× | variable | trap | low | Lowest | lowest | lowest | +| G. Shared across cores | 1/N per core | +NoC latency | +NoC latency | contended | High (coherency) | lowest in area, highest in latency | lowest in static power, high in dynamic on miss | + +(All non-FACT values are PROPOSAL/ESTIMATE pending the XH-1 cell library and frequency target.) + +## Advantages + +- **A (iterative):** lowest area and verification cost; trivially replicated ×128; deterministic timing; well-suited to in-order pipelines with a single MDU reservation station. +- **C (array + iter):** low MUL latency benefits compiled code with tight multiply chains (CRC, hashing, address computation); the iterative divider keeps area under control. +- **D (shared radix-4):** good balance; single unit means fewer ports on the issue queue, simpler back-pressure, easier to verify than E. +- **E (pipelined):** removes MDU as a critical path for high-frequency operation; appropriate only if XH-1 is a high-frequency out-of-order core. +- **F (Zmmul + trap):** smallest possible die; useful if division is genuinely rare on the target workload. +- **G (shared across cores):** minimal replicated area; only viable if the NoC has spare bandwidth and the workload tolerates MUL/DIV latency (uncommon). + +## Disadvantages + +- **A:** 64-cycle MUL is a structural hazard in nearly all non-trivial programs. Most RISC-V cores ship a faster multiplier. +- **B:** still poor MUL throughput; the pipelining helps frequency but not CPI for multiply-heavy code. +- **C:** the array multiplier is the largest single addition; replicating 128 instances dominates the per-core MDU area. The iterative divider becomes the long pole. +- **D:** the radix-4 booth encoder + Wallace reduction is the single most verification-intensive block in the integer datapath; high combinatorial depth impacts timing closure. +- **E:** the worst replication cost; if the XH-1 core is in-order (not established by repository context), this is overkill. +- **F:** division traps in scientific, crypto, and some HPC kernels; would require a fast software path that defeats the area saving. +- **G:** a single shared MDU becomes a hot spot; under load, the queuing delay exceeds a 64-cycle per-core iterative design. Also violates the principle that each core should be self-sufficient for any single instruction. + +## XH-1 Considerations + +- **ISA scope (ASSUMPTION):** XH-1 implements RV64IMAC or RV64GC. Until the decoder document (`research/03-core-design/decoder.md`) defines the supported extension set, MDU scope must include both `M` and the bitmanip multiply subset if `B` is present. +- **Pipeline style (UNKNOWN):** if the XH-1 core is in-order, an iterative or low-radix MDU is sufficient. If out-of-order, a pipelined unit is more defensible. The repository does not yet establish this. +- **Per-core MDU ports (ASSUMPTION):** at most one MDU op per cycle per core, regardless of out-of-order issue width. +- **FPU integration (OPEN):** RISC-V `F`/`D` extensions can be serviced by a separate FPU MDU (for fused multiply-add) or by forwarding to the integer MDU. Coupling decisions affect the MDU port count. +- **NoC and coherence (ASSUMPTION):** the 128 cores share a coherent memory subsystem; the MDU reads two source registers and writes one (or two for `MULH`), so register-file read/write port count is the binding constraint, not the NoC. + +## 128-Core Scalability + +- The replication cost of any MDU is *exactly* the per-core area × 128 plus shared overhead. A design choice that adds 0.05 mm² per core adds **6.4 mm²** across the chip. The C/D/E rows of the comparison table are therefore the most consequential decision in the MDU document. +- A shared-across-cores MDU (G) trades replicated area for **structural contention** at 128 cores. With a single shared unit, even a steady-state mix of 10% multiply/divide instructions causes queueing proportional to (1 / (1 − utilization)). At 50% utilization, mean wait time already exceeds the latency of an iterative per-core design. (Estimate, Little's law, INSUFFICIENT EVIDENCE for absolute numbers without a workload profile.) +- A *clustered* variant — one MDU per cluster of 4 or 8 cores — reduces the replication cost while bounding the worst-case contention. This is RECOMMENDATION-pending and is treated as an OPEN QUESTION below. +- The MDU contributes directly to per-core power density; if XH-1 is thermally constrained at 128 cores, an iterative design (A/B) is materially preferable. + +## Performance Considerations + +- A 64-cycle MUL means a `for (i=0;i What MDU micro-architecture best satisfies the XH-1 per-core area, power, latency, and throughput targets while remaining tractable to verify, to replicate 128 times, and to expose to the RISC-V ISA and toolchain? + +Sub-questions: + +1. Should multiplication and division share datapath hardware or be independent units? +2. What latency is required to avoid becoming a structural hazard in the XH-1 pipeline? +3. What is the minimum-radix divider that meets the per-core throughput target? +4. Does the cost of a fast array multiplier (e.g., radix-4 Booth with Wallace tree) justify its latency benefit over an iterative shift-add design replicated 128 times? +5. Which RISC-V extensions must be supported in v1 (`M` only vs. `M+B` vs. `M+B+K`)? + +## Background + +### RISC-V `M` Extension Semantics (FACT) + +The RISC-V `M` extension defines the following operations on the `XLEN`-bit registers (RV64 assumed for XH-1 unless stated otherwise): + +- `MUL`, `MULH`, `MULHSU`, `MULHU` — 64×64 → lower-64 / upper-64 signed/unsigned products. +- `DIV`, `DIVU`, `REM`, `REMU` — signed/unsigned division and remainder, 64÷64 → 64 quotient and 64 remainder. +- `MULW`, `DIVW`, `DIVUW`, `REMW`, `REMWU` — 32-bit operations that sign-extend the 32-bit result to 64 bits. + +Corner cases the MDU must handle per the ISA: + +- Division by zero returns `0xFFFFFFFFFFFFFFFF` (quotient) and the dividend (remainder) for unsigned, and `-1` / dividend for signed. +- The most-negative signed integer (`0x8000…`) divided by `-1` must return the dividend as quotient and `0` as remainder; this traps on x86 but is non-trapping in RISC-V. +- Overflow behavior is *only* defined for signed division/remainder by `-1`; no other overflow is architecturally visible. + +### Latency vs. Throughput (FACT) + +- A 64×64 → 128-bit full multiplication requires at minimum 128 partial-product rows for a non-Booth radix-2 design, ~64 rows for radix-4 Booth, ~43 rows for radix-8 Booth. +- Bit-serial shift-add multiplication completes in 64 cycles. +- Restoring division completes in 64 cycles; non-restoring in 32–64 cycles depending on the normalization of the remainder; higher-radix SRT dividers in O(log radix(N)) cycles but with substantially larger area. + +### RISC-V `B` and `K` Extensions (FACT, partial) + +- `B` (Bitmanip) under RV64 1.0.0 includes `MULH`, `MULHU`, `MULHSU` already in `M`, and adds rev-8.0/1.0 bitmanip groups such as `CLMUL`, `CLMULH`, `CLMULR`, `MIN[U]`, `MAX[U]`, `ANDN`, `ORN`, `XNOR`, `SEXT.B/H/W`, and the `Zba`/`Zbb`/`Zbs` sub-extensions. +- `K` (cryptography) adds `Zkn` (AES/SHA) and is not in scope for the base MDU. +- `Zmmul` provides multiply-only and is a common power-optimized subset of `M`. + +## Existing Approaches + +| Approach | Latency (typ., 64-bit) | Area (relative) | Throughput | Notes | +|---|---|---|---|---| +| Iterative shift-add multiplier | 64 cycles | 1× | 1 / 64 cycles | Smallest area; low throughput | +| Radix-4 Booth + carry-save | ~17 cycles | ~4–6× | 1 / ~17 | Common in embedded OoO | +| Radix-4 Booth + Wallace tree + final CPA | ~3–5 cycles | ~8–12× | 1 / cycle (pipelined) | Used in high-perf cores | +| Pipelined array multiplier (k stages) | k cycles | large | 1 / cycle | Throughput-bound, area-heavy | +| Radix-2 non-restoring divider | 32–64 cycles | 1× (shared with MUL possible) | 1 / 32–64 | Low area | +| Radix-4 SRT divider | ~16 cycles | ~3–5× | 1 / ~16 | Moderate | +| Radix-8/16 SRT divider | ~8–10 cycles | large | 1 / ~10 | High-end OoO | +| Combinational 64×64 array (no pipelining) | ~10+ gate delays | very large | 1 / many cycles | Impractical | +| Shared MUL/DIV datapath (Liang/Montusiewicz style) | variable | ~1.5× single-function | variable | Saves ~30–40% area at the cost of CPI | + +(Quantitative magnitudes in this table are estimates from public micro-architecture literature; values vary considerably with cell library, target frequency, and pipeline integration. See Sources.) + +## Alternative Designs + +### A. Pure Iterative Shift-Add MDU + +A single 128-bit accumulator-based datapath. Each cycle shifts the multiplicand and adds if the corresponding multiplier bit is set, or performs a non-restoring division step. The same hardware services both MUL and DIV. + +### B. Pipelined Iterative MDU + +Same datapath as A, but cut into N pipeline stages (commonly 2 or 4) to allow clock-frequency scaling. Issues per-cycle throughput remains 1 op / N cycles for a single operand width. + +### C. Dedicated Array Multiplier + Iterative Divider + +A small radix-4 Booth/Wallace multiplier (lower 64 bits in 3–4 cycles) and a separate radix-2 or radix-4 non-restoring divider. Two independent units that can operate concurrently. + +### D. Shared Pipelined Radix-4 MDU + +A single radix-4 Booth partial-product generator feeding a shared CSA tree, with a reconfigurable final-stage adder and a small division state machine. Latency: ~5–8 cycles for MUL, ~16 for DIV. + +### E. Pipelined High-Performance MDU (Wallace/Wallace + SRT-4) + +Multi-cycle but pipelined radix-4 multiplier and SRT-4 divider. Both deliver 1 op/cycle throughput once steady-state. Area dominates the integer datapath. + +### F. `Zmmul`-Only MUL, External DIV Trap + +Implement MUL in hardware (small array or iterative), trap DIV/REM to a software handler. The privileged trap path is fast on XH-1 only if the software model tolerates it; SPECint and most server workloads execute division frequently enough that this is usually unacceptable. + +### G. Compressed / Shared Across Cores + +A single physical MDU shared by N cores via a NoC-side arbiter. Removes per-core area cost but introduces structural contention; this is generally considered an *adverse* design for a 128-core tiled machine. + +## Comparison + +| Design | Area/cell count (est.) | Cycles/op MUL | Cycles/op DIV | Throughput (steady) | Verif. complexity | Replicated ×128 cost | Power/cores at 1 GHz (est.) | +|---|---|---|---|---|---|---|---| +| A. Iterative | 1.0× | 64 | 32–64 | 1/64 | Low | low | low | +| B. Pipelined iter. (2-stage) | ~1.1× | 32 | 16–32 | 1/32 | Low–Medium | low | low–medium | +| C. Array MUL + iter DIV | ~6–8× | 3–5 | 32–64 | 1/cycle (MUL) | Medium | high | medium–high | +| D. Shared radix-4 | ~4–6× | 5–8 | 16 | 1/cycle (MUL) | Medium–High | medium | medium | +| E. Pipelined Wallace + SRT-4 | ~12–20× | 1/cycle | 1/cycle | 1/cycle | High | very high | high | +| F. Zmmul + trap | ~1–2× | variable | trap | low | Lowest | lowest | lowest | +| G. Shared across cores | 1/N per core | +NoC latency | +NoC latency | contended | High (coherency) | lowest in area, highest in latency | lowest in static power, high in dynamic on miss | + +(All non-FACT values are PROPOSAL/ESTIMATE pending the XH-1 cell library and frequency target.) + +## Advantages + +- **A (iterative):** lowest area and verification cost; trivially replicated ×128; deterministic timing; well-suited to in-order pipelines with a single MDU reservation station. +- **C (array + iter):** low MUL latency benefits compiled code with tight multiply chains (CRC, hashing, address computation); the iterative divider keeps area under control. +- **D (shared radix-4):** good balance; single unit means fewer ports on the issue queue, simpler back-pressure, easier to verify than E. +- **E (pipelined):** removes MDU as a critical path for high-frequency operation; appropriate only if XH-1 is a high-frequency out-of-order core. +- **F (Zmmul + trap):** smallest possible die; useful if division is genuinely rare on the target workload. +- **G (shared across cores):** minimal replicated area; only viable if the NoC has spare bandwidth and the workload tolerates MUL/DIV latency (uncommon). + +## Disadvantages + +- **A:** 64-cycle MUL is a structural hazard in nearly all non-trivial programs. Most RISC-V cores ship a faster multiplier. +- **B:** still poor MUL throughput; the pipelining helps frequency but not CPI for multiply-heavy code. +- **C:** the array multiplier is the largest single addition; replicating 128 instances dominates the per-core MDU area. The iterative divider becomes the long pole. +- **D:** the radix-4 booth encoder + Wallace reduction is the single most verification-intensive block in the integer datapath; high combinatorial depth impacts timing closure. +- **E:** the worst replication cost; if the XH-1 core is in-order (not established by repository context), this is overkill. +- **F:** division traps in scientific, crypto, and some HPC kernels; would require a fast software path that defeats the area saving. +- **G:** a single shared MDU becomes a hot spot; under load, the queuing delay exceeds a 64-cycle per-core iterative design. Also violates the principle that each core should be self-sufficient for any single instruction. + +## XH-1 Considerations + +- **ISA scope (ASSUMPTION):** XH-1 implements RV64IMAC or RV64GC. Until the decoder document (`research/03-core-design/decoder.md`) defines the supported extension set, MDU scope must include both `M` and the bitmanip multiply subset if `B` is present. +- **Pipeline style (UNKNOWN):** if the XH-1 core is in-order, an iterative or low-radix MDU is sufficient. If out-of-order, a pipelined unit is more defensible. The repository does not yet establish this. +- **Per-core MDU ports (ASSUMPTION):** at most one MDU op per cycle per core, regardless of out-of-order issue width. +- **FPU integration (OPEN):** RISC-V `F`/`D` extensions can be serviced by a separate FPU MDU (for fused multiply-add) or by forwarding to the integer MDU. Coupling decisions affect the MDU port count. +- **NoC and coherence (ASSUMPTION):** the 128 cores share a coherent memory subsystem; the MDU reads two source registers and writes one (or two for `MULH`), so register-file read/write port count is the binding constraint, not the NoC. + +## 128-Core Scalability + +- The replication cost of any MDU is *exactly* the per-core area × 128 plus shared overhead. A design choice that adds 0.05 mm² per core adds **6.4 mm²** across the chip. The C/D/E rows of the comparison table are therefore the most consequential decision in the MDU document. +- A shared-across-cores MDU (G) trades replicated area for **structural contention** at 128 cores. With a single shared unit, even a steady-state mix of 10% multiply/divide instructions causes queueing proportional to (1 / (1 − utilization)). At 50% utilization, mean wait time already exceeds the latency of an iterative per-core design. (Estimate, Little's law, INSUFFICIENT EVIDENCE for absolute numbers without a workload profile.) +- A *clustered* variant — one MDU per cluster of 4 or 8 cores — reduces the replication cost while bounding the worst-case contention. This is RECOMMENDATION-pending and is treated as an OPEN QUESTION below. +- The MDU contributes directly to per-core power density; if XH-1 is thermally constrained at 128 cores, an iterative design (A/B) is materially preferable. + +## Performance Considerations + +- A 64-cycle MUL means a `for (i=0;i What MDU micro-architecture best satisfies the XH-1 per-core area, power, latency, and throughput targets while remaining tractable to verify, to replicate 128 times, and to expose to the RISC-V ISA and toolchain?\n\nSub-questions:\n\n1. Should multiplication and division share datapath hardware or be independent units?\n2. What latency is required to avoid becoming a structural hazard in the XH-1 pipeline?\n3. What is the minimum-radix divider that meets the per-core throughput target?\n4. Does the cost of a fast array multiplier (e.g., radix-4 Booth with Wallace tree) justify its latency benefit over an iterative shift-add design replicated 128 times?\n5. Which RISC-V extensions must be supported in v1 (`M` only vs. `M+B` vs. `M+B+K`)?\n\n## Background\n\n### RISC-V `M` Extension Semantics (FACT)\n\nThe RISC-V `M` extension defines the following operations on the `XLEN`-bit registers (RV64 assumed for XH-1 unless stated otherwise):\n\n- `MUL`, `MULH`, `MULHSU`, `MULHU` — 64×64 → lower-64 / upper-64 signed/unsigned products.\n- `DIV`, `DIVU`, `REM`, `REMU` — signed/unsigned division and remainder, 64÷64 → 64 quotient and 64 remainder.\n- `MULW`, `DIVW`, `DIVUW`, `REMW`, `REMWU` — 32-bit operations that sign-extend the 32-bit result to 64 bits.\n\nCorner cases the MDU must handle per the ISA:\n\n- Division by zero returns `0xFFFFFFFFFFFFFFFF` (quotient) and the dividend (remainder) for unsigned, and `-1` / dividend for signed.\n- The most-negative signed integer (`0x8000…`) divided by `-1` must return the dividend as quotient and `0` as remainder; this traps on x86 but is non-trapping in RISC-V.\n- Overflow behavior is *only* defined for signed division/remainder by `-1`; no other overflow is architecturally visible.\n\n### Latency vs. Throughput (FACT)\n\n- A 64×64 → 128-bit full multiplication requires at minimum 128 partial-product rows for a non-Booth radix-2 design, ~64 rows for radix-4 Booth, ~43 rows for radix-8 Booth.\n- Bit-serial shift-add multiplication completes in 64 cycles.\n- Restoring division completes in 64 cycles; non-restoring in 32–64 cycles depending on the normalization of the remainder; higher-radix SRT dividers in O(log radix(N)) cycles but with substantially larger area.\n\n### RISC-V `B` and `K` Extensions (FACT, partial)\n\n- `B` (Bitmanip) under RV64 1.0.0 includes `MULH`, `MULHU`, `MULHSU` already in `M`, and adds rev-8.0/1.0 bitmanip groups such as `CLMUL`, `CLMULH`, `CLMULR`, `MIN[U]`, `MAX[U]`, `ANDN`, `ORN`, `XNOR`, `SEXT.B/H/W`, and the `Zba`/`Zbb`/`Zbs` sub-extensions.\n- `K` (cryptography) adds `Zkn` (AES/SHA) and is not in scope for the base MDU.\n- `Zmmul` provides multiply-only and is a common power-optimized subset of `M`.\n\n## Existing Approaches\n\n| Approach | Latency (typ., 64-bit) | Area (relative) | Throughput | Notes |\n|---|---|---|---|---|\n| Iterative shift-add multiplier | 64 cycles | 1× | 1 / 64 cycles | Smallest area; low throughput |\n| Radix-4 Booth + carry-save | ~17 cycles | ~4–6× | 1 / ~17 | Common in embedded OoO |\n| Radix-4 Booth + Wallace tree + final CPA | ~3–5 cycles | ~8–12× | 1 / cycle (pipelined) | Used in high-perf cores |\n| Pipelined array multiplier (k stages) | k cycles | large | 1 / cycle | Throughput-bound, area-heavy |\n| Radix-2 non-restoring divider | 32–64 cycles | 1× (shared with MUL possible) | 1 / 32–64 | Low area |\n| Radix-4 SRT divider | ~16 cycles | ~3–5× | 1 / ~16 | Moderate |\n| Radix-8/16 SRT divider | ~8–10 cycles | large | 1 / ~10 | High-end OoO |\n| Combinational 64×64 array (no pipelining) | ~10+ gate delays | very large | 1 / many cycles | Impractical |\n| Shared MUL/DIV datapath (Liang/Montusiewicz style) | variable | ~1.5× single-function | variable | Saves ~30–40% area at the cost of CPI |\n\n(Quantitative magnitudes in this table are estimates from public micro-architecture literature; values vary considerably with cell library, target frequency, and pipeline integration. See Sources.)\n\n## Alternative Designs\n\n### A. Pure Iterative Shift-Add MDU\n\nA single 128-bit accumulator-based datapath. Each cycle shifts the multiplicand and adds if the corresponding multiplier bit is set, or performs a non-restoring division step. The same hardware services both MUL and DIV.\n\n### B. Pipelined Iterative MDU\n\nSame datapath as A, but cut into N pipeline stages (commonly 2 or 4) to allow clock-frequency scaling. Issues per-cycle throughput remains 1 op / N cycles for a single operand width.\n\n### C. Dedicated Array Multiplier + Iterative Divider\n\nA small radix-4 Booth/Wallace multiplier (lower 64 bits in 3–4 cycles) and a separate radix-2 or radix-4 non-restoring divider. Two independent units that can operate concurrently.\n\n### D. Shared Pipelined Radix-4 MDU\n\nA single radix-4 Booth partial-product generator feeding a shared CSA tree, with a reconfigurable final-stage adder and a small division state machine. Latency: ~5–8 cycles for MUL, ~16 for DIV.\n\n### E. Pipelined High-Performance MDU (Wallace/Wallace + SRT-4)\n\nMulti-cycle but pipelined radix-4 multiplier and SRT-4 divider. Both deliver 1 op/cycle throughput once steady-state. Area dominates the integer datapath.\n\n### F. `Zmmul`-Only MUL, External DIV Trap\n\nImplement MUL in hardware (small array or iterative), trap DIV/REM to a software handler. The privileged trap path is fast on XH-1 only if the software model tolerates it; SPECint and most server workloads execute division frequently enough that this is usually unacceptable.\n\n### G. Compressed / Shared Across Cores\n\nA single physical MDU shared by N cores via a NoC-side arbiter. Removes per-core area cost but introduces structural contention; this is generally considered an *adverse* design for a 128-core tiled machine.\n\n## Comparison\n\n| Design | Area/cell count (est.) | Cycles/op MUL | Cycles/op DIV | Throughput (steady) | Verif. complexity | Replicated ×128 cost | Power/cores at 1 GHz (est.) |\n|---|---|---|---|---|---|---|---|\n| A. Iterative | 1.0× | 64 | 32–64 | 1/64 | Low | low | low |\n| B. Pipelined iter. (2-stage) | ~1.1× | 32 | 16–32 | 1/32 | Low–Medium | low | low–medium |\n| C. Array MUL + iter DIV | ~6–8× | 3–5 | 32–64 | 1/cycle (MUL) | Medium | high | medium–high |\n| D. Shared radix-4 | ~4–6× | 5–8 | 16 | 1/cycle (MUL) | Medium–High | medium | medium |\n| E. Pipelined Wallace + SRT-4 | ~12–20× | 1/cycle | 1/cycle | 1/cycle | High | very high | high |\n| F. Zmmul + trap | ~1–2× | variable | trap | low | Lowest | lowest | lowest |\n| G. Shared across cores | 1/N per core | +NoC latency | +NoC latency | contended | High (coherency) | lowest in area, highest in latency | lowest in static power, high in dynamic on miss |\n\n(All non-FACT values are PROPOSAL/ESTIMATE pending the XH-1 cell library and frequency target.)\n\n## Advantages\n\n- **A (iterative):** lowest area and verification cost; trivially replicated ×128; deterministic timing; well-suited to in-order pipelines with a single MDU reservation station.\n- **C (array + iter):** low MUL latency benefits compiled code with tight multiply chains (CRC, hashing, address computation); the iterative divider keeps area under control.\n- **D (shared radix-4):** good balance; single unit means fewer ports on the issue queue, simpler back-pressure, easier to verify than E.\n- **E (pipelined):** removes MDU as a critical path for high-frequency operation; appropriate only if XH-1 is a high-frequency out-of-order core.\n- **F (Zmmul + trap):** smallest possible die; useful if division is genuinely rare on the target workload.\n- **G (shared across cores):** minimal replicated area; only viable if the NoC has spare bandwidth and the workload tolerates MUL/DIV latency (uncommon).\n\n## Disadvantages\n\n- **A:** 64-cycle MUL is a structural hazard in nearly all non-trivial programs. Most RISC-V cores ship a faster multiplier.\n- **B:** still poor MUL throughput; the pipelining helps frequency but not CPI for multiply-heavy code.\n- **C:** the array multiplier is the largest single addition; replicating 128 instances dominates the per-core MDU area. The iterative divider becomes the long pole.\n- **D:** the radix-4 booth encoder + Wallace reduction is the single most verification-intensive block in the integer datapath; high combinatorial depth impacts timing closure.\n- **E:** the worst replication cost; if the XH-1 core is in-order (not established by repository context), this is overkill.\n- **F:** division traps in scientific, crypto, and some HPC kernels; would require a fast software path that defeats the area saving.\n- **G:** a single shared MDU becomes a hot spot; under load, the queuing delay exceeds a 64-cycle per-core iterative design. Also violates the principle that each core should be self-sufficient for any single instruction.\n\n## XH-1 Considerations\n\n- **ISA scope (ASSUMPTION):** XH-1 implements RV64IMAC or RV64GC. Until the decoder document (`research/03-core-design/decoder.md`) defines the supported extension set, MDU scope must include both `M` and the bitmanip multiply subset if `B` is present.\n- **Pipeline style (UNKNOWN):** if the XH-1 core is in-order, an iterative or low-radix MDU is sufficient. If out-of-order, a pipelined unit is more defensible. The repository does not yet establish this.\n- **Per-core MDU ports (ASSUMPTION):** at most one MDU op per cycle per core, regardless of out-of-order issue width.\n- **FPU integration (OPEN):** RISC-V `F`/`D` extensions can be serviced by a separate FPU MDU (for fused multiply-add) or by forwarding to the integer MDU. Coupling decisions affect the MDU port count.\n- **NoC and coherence (ASSUMPTION):** the 128 cores share a coherent memory subsystem; the MDU reads two source registers and writes one (or two for `MULH`), so register-file read/write port count is the binding constraint, not the NoC.\n\n## 128-Core Scalability\n\n- The replication cost of any MDU is *exactly* the per-core area × 128 plus shared overhead. A design choice that adds 0.05 mm² per core adds **6.4 mm²** across the chip. The C/D/E rows of the comparison table are therefore the most consequential decision in the MDU document.\n- A shared-across-cores MDU (G) trades replicated area for **structural contention** at 128 cores. With a single shared unit, even a steady-state mix of 10% multiply/divide instructions causes queueing proportional to (1 / (1 − utilization)). At 50% utilization, mean wait time already exceeds the latency of an iterative per-core design. (Estimate, Little's law, INSUFFICIENT EVIDENCE for absolute numbers without a workload profile.)\n- A *clustered* variant — one MDU per cluster of 4 or 8 cores — reduces the replication cost while bounding the worst-case contention. This is RECOMMENDATION-pending and is treated as an OPEN QUESTION below.\n- The MDU contributes directly to per-core power density; if XH-1 is thermally constrained at 128 cores, an iterative design (A/B) is materially preferable.\n\n## Performance Considerations\n\n- A 64-cycle MUL means a `for (i=0;i