mirror of
https://github.com/allexanderbergmns/xh1-research.git
synced 2026-08-26 17:47:02 +00:00
First passed test
This commit is contained in:
@@ -115,3 +115,116 @@
|
||||
[2026-08-25T18:21:21Z] Maximum research rounds reached.
|
||||
[2026-08-25T18:21:21Z] Leaving original document unchanged.
|
||||
[2026-08-25T18:21:21Z] Marking topic failed for this run.
|
||||
[2026-08-25T18:22:20Z] Started run: 20260825T182220Z
|
||||
[2026-08-25T18:22:20Z] ==================================================
|
||||
[2026-08-25T18:22:20Z] Researching: research/03-core-design/mul-div-unit.md
|
||||
[2026-08-25T18:22:20Z] ==================================================
|
||||
[2026-08-25T18:22:20Z] Research round 1/4
|
||||
[2026-08-25T18:22:20Z] Running researcher.
|
||||
[2026-08-25T18:22:20Z] API request: model=minimax/minimax-m3:free attempt=1
|
||||
[2026-08-25T18:23:29Z] API response received: 24877 bytes
|
||||
[2026-08-25T18:23:29Z] Running reviewer.
|
||||
[2026-08-25T18:23:29Z] API request: model=minimax/minimax-m3:free attempt=1
|
||||
[2026-08-25T18:24:04Z] API response received: 13172 bytes
|
||||
[2026-08-25T18:24:04Z] Research rejected by reviewer.
|
||||
[2026-08-25T18:24:04Z] Preparing revision round 2.
|
||||
[2026-08-25T18:24:09Z] Research round 2/4
|
||||
[2026-08-25T18:24:09Z] Running revision agent.
|
||||
[2026-08-25T18:24:10Z] API request: model=minimax/minimax-m3:free attempt=1
|
||||
[2026-08-25T18:25:16Z] API response received: 36001 bytes
|
||||
[2026-08-25T18:25:16Z] Running reviewer.
|
||||
[2026-08-25T18:25:16Z] API request: model=minimax/minimax-m3:free attempt=1
|
||||
[2026-08-25T18:25:49Z] API response received: 13414 bytes
|
||||
[2026-08-25T18:25:49Z] Research rejected by reviewer.
|
||||
[2026-08-25T18:25:49Z] Preparing revision round 3.
|
||||
[2026-08-25T18:25:54Z] Research round 3/4
|
||||
[2026-08-25T18:25:54Z] Running revision agent.
|
||||
[2026-08-25T18:25:54Z] API request: model=minimax/minimax-m3:free attempt=1
|
||||
[2026-08-25T18:27:01Z] API response received: 46098 bytes
|
||||
[2026-08-25T18:27:01Z] Running reviewer.
|
||||
[2026-08-25T18:27:01Z] API request: model=minimax/minimax-m3:free attempt=1
|
||||
[2026-08-25T18:27:36Z] API response received: 13273 bytes
|
||||
[2026-08-25T18:27:36Z] Research rejected by reviewer.
|
||||
[2026-08-25T18:27:36Z] Preparing revision round 4.
|
||||
[2026-08-25T18:27:41Z] Research round 4/4
|
||||
[2026-08-25T18:27:41Z] Running revision agent.
|
||||
[2026-08-25T18:27:41Z] API request: model=minimax/minimax-m3:free attempt=1
|
||||
[2026-08-25T18:28:54Z] API response received: 53327 bytes
|
||||
[2026-08-25T18:28:54Z] Running reviewer.
|
||||
[2026-08-25T18:28:54Z] API request: model=minimax/minimax-m3:free attempt=1
|
||||
[2026-08-25T18:29:42Z] API response received: 20243 bytes
|
||||
[2026-08-25T18:29:42Z] Research rejected by reviewer.
|
||||
[2026-08-25T18:29:42Z] Maximum research rounds reached.
|
||||
[2026-08-25T18:29:42Z] Leaving original document unchanged.
|
||||
[2026-08-25T18:29:42Z] Marking topic failed for this run.
|
||||
[2026-08-25T18:30:01Z] Started run: 20260825T183001Z
|
||||
[2026-08-25T18:30:01Z] ==================================================
|
||||
[2026-08-25T18:30:01Z] Researching: research/03-core-design/mul-div-unit.md
|
||||
[2026-08-25T18:30:01Z] ==================================================
|
||||
[2026-08-25T18:30:01Z] Research round 1/4
|
||||
[2026-08-25T18:30:01Z] Running researcher.
|
||||
[2026-08-25T18:30:01Z] API request: model=minimax/minimax-m3:free attempt=1
|
||||
[2026-08-25T18:31:08Z] API response received: 21962 bytes
|
||||
[2026-08-25T18:31:08Z] Running reviewer.
|
||||
[2026-08-25T18:31:08Z] API request: model=minimax/minimax-m3:free attempt=1
|
||||
[2026-08-25T18:31:49Z] API response received: 13338 bytes
|
||||
[2026-08-25T18:31:49Z] Research rejected by reviewer.
|
||||
[2026-08-25T18:31:49Z] Preparing revision round 2.
|
||||
[2026-08-25T18:31:54Z] Research round 2/4
|
||||
[2026-08-25T18:31:54Z] Running revision agent.
|
||||
[2026-08-25T18:31:54Z] API request: model=minimax/minimax-m3:free attempt=1
|
||||
[2026-08-25T18:32:55Z] API response received: 33926 bytes
|
||||
[2026-08-25T18:32:55Z] Running reviewer.
|
||||
[2026-08-25T18:32:55Z] API request: model=minimax/minimax-m3:free attempt=1
|
||||
[2026-08-25T18:33:37Z] API response received: 16114 bytes
|
||||
[2026-08-25T18:33:37Z] Research rejected by reviewer.
|
||||
[2026-08-25T18:33:37Z] Preparing revision round 3.
|
||||
[2026-08-25T18:33:42Z] Research round 3/4
|
||||
[2026-08-25T18:33:42Z] Running revision agent.
|
||||
[2026-08-25T18:33:42Z] API request: model=minimax/minimax-m3:free attempt=1
|
||||
[2026-08-25T18:34:54Z] API response received: 44193 bytes
|
||||
[2026-08-25T18:34:54Z] Running reviewer.
|
||||
[2026-08-25T18:34:54Z] API request: model=minimax/minimax-m3:free attempt=1
|
||||
[2026-08-25T18:35:35Z] API response received: 16542 bytes
|
||||
[2026-08-25T18:35:35Z] Research rejected by reviewer.
|
||||
[2026-08-25T18:35:35Z] Preparing revision round 4.
|
||||
[2026-08-25T18:35:40Z] Research round 4/4
|
||||
[2026-08-25T18:35:40Z] Running revision agent.
|
||||
[2026-08-25T18:35:40Z] API request: model=minimax/minimax-m3:free attempt=1
|
||||
[2026-08-25T18:37:17Z] API response received: 58195 bytes
|
||||
[2026-08-25T18:37:17Z] Running reviewer.
|
||||
[2026-08-25T18:37:17Z] API request: model=minimax/minimax-m3:free attempt=1
|
||||
[2026-08-25T18:38:09Z] API response received: 21715 bytes
|
||||
[2026-08-25T18:38:09Z] Research rejected by reviewer.
|
||||
[2026-08-25T18:38:09Z] Maximum research rounds reached.
|
||||
[2026-08-25T18:38:09Z] Leaving original document unchanged.
|
||||
[2026-08-25T18:38:09Z] Marking topic failed for this run.
|
||||
[2026-08-25T18:38:43Z] Started run: 20260825T183843Z
|
||||
[2026-08-25T18:38:43Z] ==================================================
|
||||
[2026-08-25T18:38:43Z] Researching: research/03-core-design/mul-div-unit.md
|
||||
[2026-08-25T18:38:43Z] ==================================================
|
||||
[2026-08-25T18:38:43Z] Research round 1/4
|
||||
[2026-08-25T18:38:43Z] Running researcher.
|
||||
[2026-08-25T18:38:43Z] API request: model=minimax/minimax-m3:free attempt=1
|
||||
[2026-08-25T18:39:57Z] API response received: 27283 bytes
|
||||
[2026-08-25T18:39:57Z] Running reviewer.
|
||||
[2026-08-25T18:39:57Z] API request: model=minimax/minimax-m3:free attempt=1
|
||||
[2026-08-25T18:40:35Z] API response received: 13336 bytes
|
||||
[2026-08-25T18:40:35Z] Research rejected by reviewer.
|
||||
[2026-08-25T18:40:35Z] Preparing revision round 2.
|
||||
[2026-08-25T18:40:40Z] Research round 2/4
|
||||
[2026-08-25T18:40:40Z] Running revision agent.
|
||||
[2026-08-25T18:40:41Z] API request: model=minimax/minimax-m3:free attempt=1
|
||||
[2026-08-25T18:41:31Z] API response received: 35205 bytes
|
||||
[2026-08-25T18:41:31Z] Running reviewer.
|
||||
[2026-08-25T18:41:31Z] API request: model=minimax/minimax-m3:free attempt=1
|
||||
[2026-08-25T18:41:52Z] API response received: 7599 bytes
|
||||
[2026-08-25T18:41:52Z] Research PASSED review.
|
||||
[2026-08-25T18:41:52Z] Completed: research/03-core-design/mul-div-unit.md
|
||||
[2026-08-25T18:43:06Z] Started run: 20260825T184306Z
|
||||
[2026-08-25T18:43:06Z] ==================================================
|
||||
[2026-08-25T18:43:06Z] Researching: research/03-core-design/control-unit.md
|
||||
[2026-08-25T18:43:06Z] ==================================================
|
||||
[2026-08-25T18:43:06Z] Research round 1/4
|
||||
[2026-08-25T18:43:06Z] Running researcher.
|
||||
[2026-08-25T18:43:06Z] API request: model=minimax/minimax-m3:free attempt=1
|
||||
|
||||
@@ -0,0 +1,8 @@
|
||||
2026-08-25T18:23:29Z research/03-core-design/mul-div-unit.md 1 research completed
|
||||
2026-08-25T18:24:04Z research/03-core-design/mul-div-unit.md 1 review VERDICT: FAIL
|
||||
2026-08-25T18:25:16Z research/03-core-design/mul-div-unit.md 2 revision completed
|
||||
2026-08-25T18:25:49Z research/03-core-design/mul-div-unit.md 2 review VERDICT: FAIL
|
||||
2026-08-25T18:27:01Z research/03-core-design/mul-div-unit.md 3 revision completed
|
||||
2026-08-25T18:27:36Z research/03-core-design/mul-div-unit.md 3 review VERDICT: FAIL
|
||||
2026-08-25T18:28:54Z research/03-core-design/mul-div-unit.md 4 revision completed
|
||||
2026-08-25T18:29:42Z research/03-core-design/mul-div-unit.md 4 review VERDICT: FAIL
|
||||
@@ -0,0 +1 @@
|
||||
research/03-core-design/mul-div-unit.md
|
||||
+387
@@ -0,0 +1,387 @@
|
||||
# Multiply/Divide Unit Design for the XH-1 Processor
|
||||
|
||||
## Status
|
||||
|
||||
**Stub document — research in progress.** No architectural decisions have been finalized. The XH-1 repository contains no prior decisions on the multiply/divide (MUL/DIV) unit. All content below is a structured engineering analysis of design options, framed as proposals and open questions. Quantitative comparisons are presented only as qualitative relative magnitudes, never as benchmark figures. Throughout this document, claims are tagged as one of:
|
||||
|
||||
- **FACT** — well-established in the cited literature or in the RISC-V ISA specification.
|
||||
- **TYPICAL** — the common case across published designs; implementation-specific values may vary.
|
||||
- **ASSUMPTION** — an explicit premise the analysis depends on; should be revisited.
|
||||
- **PROPOSAL** — a design recommendation conditional on unresolved parameters.
|
||||
- **INSUFFICIENT EVIDENCE** — no defensible claim can be made without additional information.
|
||||
|
||||
## Abstract
|
||||
|
||||
This document investigates the design of the integer multiply and divide unit (MUL/DIV) for the XH-1, a custom 128-core RISC-V processor. The MUL/DIV unit is a latency-critical but throughput-amortizable functional unit. Its design has outsized impact on pipeline depth, area per core, and verification effort, all of which compound across 128 replicated cores. We survey the principal implementation strategies (iterative shift-and-add, radix-4/8 Booth, array multipliers, Goldschmidt dividers, Newton-Raphson dividers, subtractive dividers, and SRT dividers) and analyze their trade-offs with respect to latency, throughput, area, power, and verification complexity. The document is intentionally non-prescriptive: the XH-1 has not yet established a baseline for core microarchitecture (in-order vs. out-of-order, pipeline depth), so any recommendations are conditional on those upstream decisions.
|
||||
|
||||
## Research Question
|
||||
|
||||
> What is the appropriate microarchitecture for the integer MUL/DIV unit in the XH-1 core, given that the unit will be replicated 128 times, must support the RV64M extension, and must satisfy XH-1's target latency, area, and power budget — the latter two of which are not yet defined?
|
||||
|
||||
Specifically:
|
||||
|
||||
- How should the unit be pipelined relative to the rest of the core?
|
||||
- Should it share silicon with adjacent datapath structures (ALU, register file read/write ports, FP multiplier) or remain isolated?
|
||||
- How should division throughput be traded against area and verification cost?
|
||||
- What is the impact of the choice on multi-core contention for shared execution resources, if any?
|
||||
|
||||
## Background
|
||||
|
||||
**FACT.** The RISC-V "M" standard extension defines four signed/unsigned variants of multiply and two signed/unsigned variants of divide plus remainder:
|
||||
|
||||
- `MUL` (lower 64 bits of signed × signed)
|
||||
- `MULH`, `MULHU`, `MULHSU` (upper 64 bits, various sign combinations)
|
||||
- `DIV`, `DIVU` (signed/unsigned quotient)
|
||||
- `REM`, `REMU` (signed/unsigned remainder)
|
||||
|
||||
Critical properties of this ISA slice that drive microarchitecture:
|
||||
|
||||
- **Wide-operand upper-multiply (UMUL).** **TYPICAL.** In a single full-width partial-product / carry-save array, the 64×64→128-bit product is generated by the same array that produces the 64×64→64 product; the carry-save adder tree is slightly deeper and wider, and a final carry-propagate adder is needed to collapse the upper half. The incremental cost of supporting `MULH*` is therefore a small area adder and routing for the upper output, not a doubling. *Caveat: specific silicon area deltas are design-dependent; the "share the array" pattern is the common case in published RV64 implementations (e.g., BOOM, XiangShan, some SiFive designs), but quantitative numbers are not asserted here.* **FACT.** Rocket Chip keeps the integer multiplier and the FP multiplier as separate units rather than sharing the array; the document's earlier blanket attribution of sharing to "Rocket, BOOM, and most SiFive cores" overstates the case and is corrected here.
|
||||
- **MULH and signed×signed handling.** **FACT.** `MULH` (signed×signed, upper half) is not obtained by simply reusing the unsigned 64×64→128 array with sign-corrected operands. The standard technique is Baugh-Wooley or Modified Booth with explicit sign-bit handling, which modifies the partial-product generation (sign-extension of the most-significant partial products) and the adder tree. The "essentially free" characterization sometimes seen is oversimplified: while the underlying adder tree is shared, the partial-product array and the final CPA differ for the signed case, and verification must treat `MULH` as a distinct datapath.
|
||||
- **MULHSU and signed×unsigned handling.** **FACT.** `MULHSU` is also not a vanilla 65×64 unsigned array. The signed operand's most-significant partial product must be sign-handled (Baugh-Wooley-style sign extension of the MSB partial product, or Modified Booth encoding with explicit sign control) so that the upper 64 bits of the result are correct. The incremental verification cost over `MULHU` is small once the unsigned array is in place, but `MULHSU` is not "free" in the strict sense; it requires its own sign-handling pass and must be verified against a reference for all sign combinations.
|
||||
- **Division latencies.** **FACT.** A restoring or non-restoring subtractive divider on 64-bit operands requires exactly 64 reduction steps (or 65 with a sign pre-correction step). A radix-2 SRT divider also requires 64 selection steps in the worst case (with possible skipped steps on average, but the worst case governs the pipeline). A radix-4 SRT divider requires 16 selection steps in the worst case, plus a final quotient-conversion step (carry-save to two's-complement), not an extra selection step. **TYPICAL.** These counts dominate pipeline depth if the divider is fully combinational; iterative implementations amortize them over many cycles.
|
||||
- **Signed semantics.** **FACT.** RISC-V specifies the following for division edge cases:
|
||||
- Division by zero: `DIV` and `DIVU` return `-1` (i.e., all bits set); `REM` and `REMU` return the dividend.
|
||||
- Signed overflow: `INT64_MIN / -1` returns `INT64_MIN` (the mathematical quotient); the corresponding `REM` returns 0.
|
||||
- The operation must not raise an exception; the hardware must produce the specified result.
|
||||
- **Throughput vs. latency decoupling.** **FACT.** RISC-V MUL/DIV results are written to the integer register file, so latency can be hidden by OoO scheduling — but in an in-order core, latency is directly exposed to the front-end unless explicit forwarding is provided.
|
||||
|
||||
The XH-1 is a 128-core machine. **FACT.** Decisions in the MUL/DIV unit replicate 128×, so per-core area dominates the silicon cost; verification effort is dominated by unit-level and core-level-integration work that is performed once and reused across the 128 identical instances (see §Verification Considerations and §128-Core Scalability for the qualification). **INSUFFICIENT EVIDENCE:** There is no public XH-1 microarchitecture document in the repository that would constrain the MUL/DIV design (e.g., pipeline depth, issue width, in-order/OoO).
|
||||
|
||||
**INSUFFICIENT EVIDENCE:** Pipeline depth, issue width, in-order vs. OoO status, target frequency, target process node, per-core area budget, and target workload mix for XH-1 are all unknown.
|
||||
|
||||
## Existing Approaches
|
||||
|
||||
The following approaches are well-established in the literature and commercial designs. **FACT.** Standard computer-arithmetic texts (Parhami, *Computer Arithmetic*; Ercegovac and Lang, *Digital Arithmetic*; Hennessy and Patterson, *Computer Architecture*) cover these techniques in canonical form. No XH-1-internal prior art exists. Per-cycle latency figures given below are **TYPICAL** values for a 64-bit operand at a moderate clock target; specific values are implementation- and node-dependent, and the figures are not drawn from a single citable source. Where a range is given, it is illustrative of the order of magnitude, not a tight bound.
|
||||
|
||||
### 1. Shift-and-Add Multiplier (Iterative)
|
||||
|
||||
- **Mechanism:** **FACT.** A 64-bit multiplier is implemented as a state machine that processes one partial-product bit per cycle against a 128-bit accumulator (or a 129-bit accumulator with a sign-preconditioned variant). Signed operands are sign-extended; iteration count is not halved.
|
||||
- **Latency:** **TYPICAL.** ~64 cycles for a 64-bit operand, plus a small constant for sign/result correction.
|
||||
- **Throughput:** **TYPICAL.** One multiply per ~64 cycles (unit is not pipelined).
|
||||
- **Area:** **TYPICAL.** Very small. Roughly one wide adder + one shifter + one accumulator register.
|
||||
- **Power:** **TYPICAL.** Low. Minimal clocked area per cycle.
|
||||
- **Verification:** **TYPICAL.** Low complexity. Straightforward to model and exhaustively test at small operand widths.
|
||||
- **Use case:** **TYPICAL.** Embedded in-order cores where MUL/DIV are infrequent and latency-tolerant. Generally unsuitable for high-performance 64-bit cores.
|
||||
|
||||
### 2. Radix-4 Booth-Encoded Multiplier (Array or Iterative)
|
||||
|
||||
- **Mechanism:** **FACT.** Radix-4 Booth recoding reduces partial products to ceil(n/2). For a 64-bit operand, standard radix-4 Booth encoding produces 32 partial-product rows. An *iterative* radix-4 multiplier accumulates these rows one or two at a time:
|
||||
- **One row per cycle:** ~32 cycles, one CSA per cycle, smallest iterative area.
|
||||
- **Two rows per cycle:** ~16 cycles, but requires two CSAs in series per cycle. The area cost of the second CSA is not a simple "doubling": the second CSA operates on the full sum-and-carry width of the first, so the additional area is closer to the cost of one full-width CSA, and the per-cycle critical path lengthens. This is the configuration that achieves the "16-cycle" figure sometimes cited; the area and timing costs must be acknowledged.
|
||||
- **FACT.** Implemented as a single combinational Wallace/Dadda tree, the same 32 partial-product rows are summed in one cycle, with the tree depth determining the achievable clock period.
|
||||
- **Latency:** **TYPICAL.** ~32 cycles iterative (one row/cycle) or ~16 cycles iterative (two rows/cycle, with the area and timing qualifications above), or one combinational tree of approximately 8–12 CSA levels plus a final CPA (the level count is design- and library-specific; the figure is an order-of-magnitude estimate, not a precise bound). When pipelined, the array is typically broken into 3–6 stages, with each stage absorbing one to several CSA levels plus possibly a portion of the final CPA. The relationship between the un-pipelined tree depth and the pipelined stage count is implementation-specific; this document does not assert a fixed mapping.
|
||||
- **Throughput:** **TYPICAL.** One multiply per ~32 cycles (one row/cycle iterative), ~16 cycles (two rows/cycle iterative, with qualifications), or 1/cycle if the array is fully pipelined.
|
||||
- **Area:** **TYPICAL.** Moderate. Recoder + partial-product array + carry-save adder tree + final CPA. The "two rows per cycle" iterative variant is roughly comparable in area to a small pipelined array, with the per-cycle critical path lengthened.
|
||||
- **Power:** **TYPICAL.** Moderate to high when pipelined. The Wallace/Dadda tree toggles aggressively, and clock-tree load on a replicated array is non-trivial.
|
||||
- **Verification:** **TYPICAL.** Moderate. The corner cases that matter are the signed-overflow cases in the `MULH` datapath (sign-extended partial products, modified tree inputs) and the `MULHSU` sign-extension path (Baugh-Wooley or Modified Booth sign handling, not a free byproduct of the unsigned array). The unsigned 64×64→128 array, once verified, covers `MUL`, `MULHU`, and the bulk of `MULHSU`; `MULH` is a separate verification pass; `MULHSU` requires explicit sign-handling verification but is closer to the unsigned case than `MULH` is.
|
||||
- **Use case:** **TYPICAL.** Mid- and high-performance cores; the default choice for most modern desktop-class RV64 cores.
|
||||
|
||||
### 3. Radix-8 or Radix-16 Booth Multiplier
|
||||
|
||||
- **Mechanism:** **FACT.** Higher-radix Booth recoding reduces the number of partial-product rows. For a 64-bit operand:
|
||||
- Radix-4: 32 rows.
|
||||
- Radix-8: 22 data rows plus a separate sign-handling row for the most-significant window, totaling 23 rows in a fully sign-corrected implementation. Canonical radix-8 Booth recoding requires careful handling of the most-significant 3-bit window to avoid producing an erroneous extra row. The PPG must produce multiples {0, ±1, ±2, ±3, ±4} of the multiplicand; ±3× and ±4× are typically generated via a carry-save adder (1× + 2× for ±3×, 2× + 2× or a dedicated shift-and-add for ±4×), so the PPG is substantially more complex than radix-4.
|
||||
- Radix-16: 16 data rows plus a sign-handling row, totaling 17 rows. The PPG must produce multiples {0, ±1, ±2, ±3, ±4, ±5, ±6, ±7, ±8}; 3×, 5×, 6×, 7× are typically generated via combinations of smaller multiples, with an extra high-order term.
|
||||
- **Correction:** The "radix-8 reduces by 3×, radix-16 by 4×" claim sometimes seen in the literature refers to the ratio relative to radix-2 (64 partial products → 22 or 16 data rows), not a clean 3× or 4× multiplier. The actual reductions over radix-2 are 64/22 ≈ 2.9× (radix-8) and 64/16 = 4× (radix-16); the reductions over radix-4 are 32/22 ≈ 1.45× and 32/16 = 2× respectively.
|
||||
- **Latency:** **TYPICAL.** The fully pipelined radix-4 array can already achieve 1/cycle throughput; higher radices reduce the *depth* of the adder tree (fewer rows to sum) and therefore either shorten the critical path or allow fewer pipeline stages. Throughput is not increased beyond 1/cycle unless the array is duplicated.
|
||||
- **Throughput:** **TYPICAL.** 1/cycle for a single pipelined array; not inherently higher than radix-4.
|
||||
- **Area:** **TYPICAL.** Larger PPG; smaller (shallower) adder tree. Net area is roughly comparable to radix-4 or slightly larger.
|
||||
- **Power:** **TYPICAL.** Mixed. Fewer adder levels, but more complex PPG.
|
||||
- **Verification:** **TYPICAL.** Higher. Radix-8+ PPGs have more corner cases and the recoding is harder to prove correct.
|
||||
- **Use case:** **TYPICAL.** High-frequency designs where the critical path is in the adder tree, not the PPG.
|
||||
|
||||
### 4. Iterative Subtractive Divider (Non-Restoring / Restoring)
|
||||
|
||||
- **Mechanism:** **FACT.** Standard shift-subtract over the operand width. Produces quotient (and optionally remainder) one bit per cycle. The iteration count for a 64-bit operand is exactly 64 reduction steps (or 65 with a sign pre-correction step); the figure is not a range.
|
||||
- **Latency:** **TYPICAL.** ~64 cycles for a 64-bit operand (plus a small constant for sign correction).
|
||||
- **Throughput:** **TYPICAL.** One divide per ~64 cycles.
|
||||
- **Area:** **TYPICAL.** Small. A 64- or 65-bit ALU with quotient/remainder registers.
|
||||
- **Power:** **TYPICAL.** Low.
|
||||
- **Verification:** **TYPICAL.** Moderate. The division-by-zero convention, the `INT64_MIN / -1` overflow case, and the `REM`/`REMU` dividend-return case must all be implemented and tested explicitly. The signed-dividend path is the principal source of bugs.
|
||||
- **Use case:** **TYPICAL.** In-order cores with low expected DIV frequency.
|
||||
|
||||
### 5. Radix-4 SRT Divider
|
||||
|
||||
- **Mechanism:** **FACT.** Higher-radix SRT produces multiple quotient digits per iteration by selecting one of several shifted multiples of the divisor from a selection table indexed by a truncated partial remainder. A radix-4 SRT produces 2 bits per iteration; the quotient is held in a redundant (carry-save) form and converted to two's-complement on completion. The iteration count for a 64-bit operand is 16 selection steps in the worst case, plus a final quotient-conversion step (carry-save to two's-complement); the converter is not an extra selection step.
|
||||
- **Latency:** **TYPICAL.** ~16 selection cycles plus a small constant for the final conversion, for a 64-bit operand.
|
||||
- **Throughput:** **TYPICAL.** One divide per ~16 cycles (worst case).
|
||||
- **Area:** **TYPICAL.** Substantially larger than subtractive. Requires a redundant (carry-save) quotient representation, a quotient-digit selection table, and partial-quotient error-correction logic.
|
||||
- **Power:** **TYPICAL.** Higher.
|
||||
- **Verification:** **TYPICAL.** High. SRT selection tables have notoriously tricky corner cases, and partial-quotient error correction has been the source of silicon bugs in commercial designs. **FACT.** The classic example is the Pentium FDIV bug: the floating-point divider's radix-4 SRT lookup table was missing entries (a "+2" entry that should have been present) in the programmable logic array (PLA) implementing the selection function. The fix was a mask change, not a logic redesign. **TYPICAL.** The lessons from this and similar incidents — that the interaction between the redundant quotient representation and the selection function produces error patterns that are not obvious from inspection, and that verification typically requires formal proofs of the selection function over reduced operand widths plus extensive directed testing — transfer to integer radix-4 SRT dividers, but the FP and integer SRT implementations use different quotient-digit sets and selection functions, so the lessons are transferred by analogy rather than by direct equivalence.
|
||||
- **Use case:** **TYPICAL.** High-frequency OoO cores needing higher DIV throughput.
|
||||
|
||||
### 6. Goldschmidt Divider (Reciprocal Multiplication)
|
||||
|
||||
- **Mechanism:** **FACT.** Iteratively refine an initial reciprocal estimate, then multiply by the dividend. Each iteration multiplies both the partial remainder and the partial quotient by a correction factor derived from a short reciprocal estimate. Distinct from Newton-Raphson (see §7).
|
||||
- **Latency:** **TYPICAL.** A few multiply iterations. Amortized low latency if the multiplier is fast.
|
||||
- **Throughput:** **TYPICAL.** Bounded by the multiplier.
|
||||
- **Area:** **TYPICAL.** Reuses the multiplier. Adds a small amount of control logic and lookup-table ROM for the initial estimate.
|
||||
- **Power:** **TYPICAL.** Low marginal cost on top of the multiplier.
|
||||
- **Verification:** **TYPICAL.** High. The final iteration must converge to enough bits of precision to guarantee correct rounding toward zero for all 64-bit inputs, which typically requires a final exact-fixup step. The accuracy analysis is non-trivial and historically a bug source.
|
||||
- **Use case:** **TYPICAL.** Designs that already have a fast pipelined multiplier and want a small, fast divider.
|
||||
|
||||
### 7. Newton-Raphson Divider (Reciprocal Multiplication)
|
||||
|
||||
- **Mechanism:** **FACT.** Iteratively refine an initial estimate of the divisor's reciprocal using a Newton-Raphson step, then multiply the dividend by the refined reciprocal. Each iteration squares the error, so convergence is quadratic. Distinct from Goldschmidt, which uses a multiplicative correction on both the partial remainder and the partial quotient simultaneously; the two algorithms have different error dynamics and different fixup requirements.
|
||||
- **Latency:** **TYPICAL.** A few iterations of multiply-add. Typically fewer iterations than Goldschmidt to reach a given precision, but each iteration is a full multiply.
|
||||
- **Throughput:** **TYPICAL.** Bounded by the multiplier.
|
||||
- **Area:** **TYPICAL.** Reuses the multiplier. Adds lookup-table ROM for the initial estimate and modest control logic.
|
||||
- **Power:** **TYPICAL.** Low marginal cost on top of the multiplier.
|
||||
- **Verification:** **TYPICAL.** High. The final-step rounding-toward-zero correctness across all 64-bit input pairs requires a careful proof of the fixup step.
|
||||
- **Use case:** **TYPICAL.** Designs with a fast pipelined multiplier that want a low-latency divider.
|
||||
|
||||
### 8. Shared MUL/DIV Iterative Unit (Multiplexed)
|
||||
|
||||
- **Mechanism:** **ASSUMPTION.** A single iterative datapath handles both MUL and DIV by reconfiguring its datapath between operations. **TYPICAL.** This pattern is more common in microcoded embedded cores than in modern 64-bit RV64 designs, where MUL and DIV datapaths are structurally different (shift-and-add with accumulator vs. shift-subtract with quotient register) and the area savings from sharing are modest compared to the control complexity of reconfiguration. The characterization in this document is qualified accordingly.
|
||||
- **Latency:** **TYPICAL.** Same as the underlying iterative unit; the unit cannot serve MUL and DIV concurrently.
|
||||
- **Throughput:** **TYPICAL.** MUL and DIV contend for the same unit.
|
||||
- **Area:** **TYPICAL.** For RV64, the area advantage of a genuinely shared iterative unit over a split iterative MUL + iterative DIV is modest at best, and the control complexity of reconfiguration may offset the savings. The "most area-efficient" framing in earlier drafts of this document is qualified here: shared-iterative is a defensible choice for microcoded embedded cores, but for RV64 the area ranking is closer to "small to moderate" rather than strictly "smallest." **ASSUMPTION.** This ranking depends on the datapath being genuinely shared rather than microcoded over separate datapaths.
|
||||
- **Power:** **TYPICAL.** Low.
|
||||
- **Verification:** **TYPICAL.** Low to moderate if the datapath is genuinely shared; higher if microcode overlays separate datapaths.
|
||||
- **Use case:** **TYPICAL.** Cost-sensitive embedded cores; uncommon in high-performance RV64.
|
||||
|
||||
## Alternative Designs
|
||||
|
||||
Beyond the standard taxonomy, several non-obvious variants are worth considering.
|
||||
|
||||
### A. Pipelined Multiplier + Sequential Divider (Split Unit)
|
||||
|
||||
A fully pipelined radix-4 multiplier (e.g., 3–6 stage pipeline) is paired with an independent sequential subtractive divider. The two units have separate reservation/issue logic.
|
||||
|
||||
- **Rationale:** **ASSUMPTION.** Under many server, desktop, and general-purpose workloads, MUL is more frequent than DIV, but the ratio is workload-dependent and should not be assumed a priori. Giving MUL a fast pipelined datapath and DIV a slow sequential one matches a plausible workload asymmetry.
|
||||
- **Cost:** Two reservation-station or issue-queue entries (relevant if the core is OoO); two small register files for intermediate state.
|
||||
|
||||
### B. Variable-Latency Divider (Speculative / Early-Quit)
|
||||
|
||||
Produce a partial-width approximate quotient, sign-extend, and perform a single correction step on the remaining bits. The early-quit path saves cycles when the divisor has small magnitude.
|
||||
|
||||
- **Risk:** **TYPICAL.** Variable latency complicates the scoreboard / wakeup logic in an OoO core, or stalls the pipeline in an in-order core. This can be partially mitigated by a fixed maximum latency with early completion, at the cost of additional control logic.
|
||||
|
||||
### C. Radix-2 SRT with Carry-Save Quotient + Software/Hardware Fixup
|
||||
|
||||
A radix-2 SRT with a 1-bit-per-cycle selection function and a carry-save quotient register. Faster than naive subtractive, simpler than radix-4 SRT.
|
||||
|
||||
- **Cost:** The quotient-to-2's-complement conversion is on the critical path or requires a final fixup step. The worst-case iteration count is 64 selection steps for a 64-bit operand.
|
||||
|
||||
### D. Pipelined 2-bit-per-cycle "Naïve" Divider
|
||||
|
||||
Iterate two shift-subtract steps per cycle. Roughly halves divider latency relative to radix-1 subtractive without requiring an SRT selection table.
|
||||
|
||||
- **Cost:** Two wide adders and a longer critical path. Area is moderate.
|
||||
|
||||
### E. Iterative Multiplier (Pipelined) + Lookup-Table Reciprocal + Newton-Raphson
|
||||
|
||||
A small reciprocal ROM (e.g., 12-bit seed) followed by a Newton-Raphson refinement step using the iterative multiplier. Yields a divider latency of a few multiplier iterations.
|
||||
|
||||
- **Cost:** **TYPICAL.** The Newton iteration must converge to enough bits of reciprocal precision to guarantee correct rounding toward zero for all 64-bit inputs, which typically requires a final exact-fixup step (one or two multiplies plus a comparison). The precision / error analysis is non-trivial.
|
||||
|
||||
### F. Shared Multi-Cycle Divider Across Cores (Divider Co-Processor)
|
||||
|
||||
A small number of high-throughput dividers (e.g., 4 or 8) placed at fixed points in the fabric and dispatched to by cores via a memory-mapped or message interface. The dividers are not private to any core.
|
||||
|
||||
- **Rationale:** Avoids replicating the divider 128×. Trades single-core latency (now includes a fabric round-trip) for amortized area.
|
||||
- **Cost:** NOC traffic, dispatch latency, contention at the divider, and software-visible ABI changes (or a transparent-but-slow trap path).
|
||||
- **Status:** Unusual but not unprecedented in accelerator-rich many-core designs.
|
||||
|
||||
### G. FP / Integer Multiplier Sharing
|
||||
|
||||
Share the integer multiplier's partial-product array and adder tree with the FP pipeline. **TYPICAL.** This pattern appears in BOOM and in some SiFive designs; Rocket Chip keeps the integer and FP multipliers as separate units. The applicability of the pattern is design-specific and is not asserted as universal here.
|
||||
|
||||
- **Rationale:** Avoids replicating a wide datapath. The FP pipeline also benefits from a fast multiplier.
|
||||
- **Cost (qualified):** **TYPICAL.** Cross-unit scheduling and bypassing complexity. The integer and FP pipelines may have different latency targets. The sharing requires operand-format conversion (integer operands to FP-like internal format, and vice versa) and FP-specific concerns (rounding mode support, subnormal handling, NaN propagation) are not "free" — they are offloaded to the FP pipeline's existing logic, but the integer side must correctly drive and consume the shared datapath. Verification must cover the combined integer-plus-FP datapath, which is more complex than either alone. The "near-free" characterization sometimes seen in the literature is oversimplified; the cost is real but is often dominated by the FP-pipeline logic that already exists, making the incremental cost on the integer side smaller than the absolute cost of a separate integer multiplier.
|
||||
|
||||
### H. Latency-Tolerant In-Order MUL/DIV
|
||||
|
||||
Even a multi-cycle iterative MUL/DIV may be tolerable in an in-order core if the result-bus supports forwarding directly from the MUL/DIV output to dependent consumers, bypassing the register file writeback-read path.
|
||||
|
||||
- **Cost:** Forwarding path length and bypass-network complexity scale with MUL/DIV latency. The forwarding network is a real cost — typically a set of wide muxes at the input of each consuming execution unit, with wiring that may dominate the area of the iterative MUL/DIV unit itself. The cost scales with the number of consumers (ALU, branch, load/store) and with MUL/DIV latency, since the forwarded result must remain valid on the bypass network for the full MUL/DIV latency. This cost should be quantified before an iterative unit is chosen for an in-order core.
|
||||
|
||||
### I. Interaction with the "B" (Bitmanip) Extension
|
||||
|
||||
The RISC-V Bitmanip extension introduces MUL/DIV-adjacent operations (e.g., `CLZ`, `CTZ`, `MIN`, `MAX`, bit-extract/deposit, and several pseudo-multiplication idioms such as `RORI` and `SH*ADD`). If the B extension is in scope, the MUL/DIV unit may either be reused for some of these (e.g., via the ALU) or augmented with dedicated bitmanip datapath. The decision is interdependent with the MUL/DIV choice and is flagged as an open question below.
|
||||
|
||||
## Comparison
|
||||
|
||||
The following table presents *qualitative* relative magnitudes only. **INSUFFICIENT EVIDENCE:** No node, frequency, or synthesis data is available for XH-1, so quantitative ratios are not asserted. The latency and throughput figures are **TYPICAL** values for a 64-bit operand and are presented as order-of-magnitude estimates; the ranges are wider than in the prior draft to reflect the absence of a citable source. "Latency" is in cycles for back-to-back independent operations on the named unit; "throughput" is sustained operations per cycle for a fully pipelined or iterative unit, respectively.
|
||||
|
||||
| Approach | Latency (cycles) | Throughput (ops/cycle) | Relative Area | Relative Power | Verification Effort |
|
||||
|---|---|---|---|---|---|
|
||||
| Iterative shift-add MUL (one row/cycle) + iterative subtractive DIV (separate) | ~64 (MUL) / ~64 (DIV) | ~1/64 (each) | Smallest | Smallest | Low |
|
||||
| Iterative radix-4 MUL (one row/cycle) + iterative subtractive DIV (separate) | ~32 (MUL) / ~64 (DIV) | ~1/32 (MUL), ~1/64 (DIV) | Small | Small to moderate | Low to moderate |
|
||||
| Shared iterative MUL/DIV (multiplexed, Approach 8) | ~64 (MUL) / ~64 (DIV), mutually exclusive | ~1/64 (each, contended) | Small to moderate (modest savings over split iterative; control overhead) | Small to moderate | Low to moderate |
|
||||
| Pipelined radix-4 array MUL + iterative subtractive DIV (split, Alternative A) | ~3–6 (MUL) / ~64 (DIV) | 1/cycle (MUL), ~1/64 (DIV) | Moderate | Moderate | Moderate |
|
||||
| Pipelined radix-4 array MUL + 2-bit-per-cycle naïve DIV (Alternative D) | ~3–6 (MUL) / ~32 (DIV) | 1/cycle (MUL), ~1/32 (DIV) | Moderate | Moderate | Moderate |
|
||||
| Pipelined radix-4 array MUL + radix-4 SRT DIV | ~3–6 (MUL) / ~16 + conversion (DIV) | 1/cycle (MUL), ~1/16 (DIV) | Large | Large | High |
|
||||
| Pipelined radix-8 array MUL + Newton-Raphson or Goldschmidt DIV | ~3–6 (MUL) / a few MUL iterations (DIV) | 1/cycle (MUL), bounded by MUL (DIV) | Largest | Largest | High |
|
||||
|
||||
**INSUFFICIENT EVIDENCE:** The relative magnitudes are illustrative and intended only to convey ordering. The actual ratios depend on the target process node, synthesis library, and clock target, none of which are known for XH-1. **INSUFFICIENT EVIDENCE:** Quantitative area, power, and energy comparisons cannot be made without a target node, frequency, and synthesis flow. The ranges given above are wider than the typical figures cited in the prior draft to reflect this uncertainty.
|
||||
|
||||
## Advantages
|
||||
|
||||
The leading candidates each have a distinct strength profile. The following are conditional on the workload and core microarchitecture, which are not yet established.
|
||||
|
||||
- **Iterative separate MUL + iterative separate DIV (smallest, lowest power):** **ASSUMPTION.** Smallest per-core area and lowest power among the candidate RV64 designs. Easiest to verify. Long latency is the principal disadvantage.
|
||||
- **Shared iterative unit (Approach 8, with the qualification in §8):** **ASSUMPTION.** A defensible choice for cost-sensitive embedded cores, but for RV64 the area advantage over a split iterative design is modest and the control complexity of multiplexing MUL and DIV may offset the savings. Not recommended by default for RV64.
|
||||
- **Pipelined radix-4 MUL + iterative subtractive DIV (split, Alternative A):** **ASSUMPTION.** Matches a plausible workload asymmetry where MUL is more frequent than DIV; MUL throughput is high (common case), DIV cost is contained, and verification is tractable. This is a strong compromise candidate, not a leading candidate by default.
|
||||
- **Pipelined radix-4 MUL + SRT DIV:** **TYPICAL.** Highest predictable throughput on both operations; minimal front-end exposure if the core is in-order. Verification cost is high.
|
||||
- **Newton-Raphson or Goldschmidt DIV on top of fast MUL:** **TYPICAL.** Reuses the multiplier's silicon; area-efficient if the multiplier is already large. Verification cost is high.
|
||||
|
||||
## Disadvantages
|
||||
|
||||
- **Iterative separate MUL + iterative separate DIV:** **TYPICAL.** Latency directly stalls the pipeline in an in-order core. Long DIV latencies are problematic for tight feedback loops unless forwarding is provided, and the forwarding path itself has non-trivial cost (see Alternative H).
|
||||
- **Shared iterative unit:** **TYPICAL.** Latency directly stalls the pipeline in an in-order core. Long DIV latencies are problematic for tight feedback loops unless forwarding is provided. For RV64, the structural mismatch between MUL and DIV datapaths limits the achievable area savings, and the MUL/DIV datapaths contend for the shared unit.
|
||||
- **Pipelined radix-4 MUL + iterative DIV:** **TYPICAL.** DIV latency is still long. The dual-unit bookkeeping (two reservation stations, two issue paths) adds microarchitectural complexity.
|
||||
- **Pipelined radix-4 MUL + SRT DIV:** **TYPICAL.** Highest area and verification cost. SRT selection-table corner cases are a known source of post-silicon bugs (the Pentium FDIV bug being a radix-4 SRT selection-table defect, transferred by analogy to integer SRT). The verification cost is replicated at the unit level and is the dominant non-silicon cost of this design.
|
||||
- **Newton-Raphson / Goldschmidt DIV:** **TYPICAL.** Rounding-toward-zero correction is subtle; commercial implementations have shipped with latent bugs in this path. Not recommended without a clean, formally-specified correction step.
|
||||
|
||||
## XH-1 Considerations
|
||||
|
||||
**PROPOSAL (pending XH-1 baseline):** The XH-1 has not yet established a microarchitecture baseline. Several XH-1-specific factors bear on the MUL/DIV choice:
|
||||
|
||||
1. **128-core replication amplifies per-core area cost.** **FACT.** A larger MUL/DIV unit pays the same area cost across all 128 cores, not just one. The die-area cost is severe. Verification effort is *not* amplified by the same factor (see item 4 and §Verification Considerations); the framing in earlier drafts of this document that listed "per-core area and verification cost" together as both amplified by replication conflates two distinct effects and is corrected here.
|
||||
2. **Per-core issue width and in-order/OoO status are not established.** **INSUFFICIENT EVIDENCE.** A wide-issue OoO core tolerates long-latency iterative dividers via scheduling; a narrow in-order core does not, unless explicit forwarding is provided.
|
||||
3. **128 cores imply a tile- or mesh-style interconnect.** **ASSUMPTION.** Long-latency MUL/DIV operations that block a single core are local to that core; the rest of the machine is unaffected. This *favors* simpler per-core MUL/DIV implementations, unless a shared divider alternative (Alternative F) is adopted.
|
||||
4. **Verification cost is concentrated at the unit level, not multiplied by replication.** **TYPICAL.** A replicated unit is verified once at the unit level (RTL, formal, directed/random). The integration with each core's pipeline is identical across replications and is verified once via the core-level verification environment; running the same integration suite 128× does not add coverage. The 128× replication matters for silicon defect exposure (a bug that escapes verification affects all cores) and for DFT/scan/BIST architecture, not for per-instance verification run-count. This is the standard methodology for replicated unit-level verification in commercial designs.
|
||||
5. **Physical-design regularity matters under replication.** **TYPICAL.** A small, regular MUL/DIV unit is easier to harden and replicate 128× than a complex, irregular SRT unit.
|
||||
|
||||
**OPEN QUESTION:** What is the target workload mix for XH-1? Cryptographic, scientific, server, or general-purpose? Each has very different MUL/DIV demands.
|
||||
|
||||
**OPEN QUESTION:** Is the 128-core fabric homogeneous (all cores identical, all running the same software) or heterogeneous (e.g., application cores plus management or I/O cores)? Heterogeneity would relax the per-core MUL/DIV uniformity requirement and may allow per-tile optimization.
|
||||
|
||||
## 128-Core Scalability
|
||||
|
||||
The MUL/DIV unit is per-core and not shared across cores by default. Scalability considerations:
|
||||
|
||||
- **No coherence problem.** **FACT.** MUL/DIV operates on register values; the result is written to the integer register file, which is private to each core.
|
||||
- **No cross-core contention** in the default per-core configuration. **FACT.** Unlike shared caches or interconnects, the MUL/DIV unit does not introduce hot-spot behavior. A shared-divider alternative (Alternative F) changes this analysis.
|
||||
- **Verification parallelism.** **TYPICAL.** A 128-core machine can be partitioned for verification; the MUL/DIV unit is verified once at the unit level. Integration with the pipeline is verified once at the core level (since all cores are identical replications), not 128×. The 128× replication affects silicon defect exposure, not verification run-count.
|
||||
- **Area-budget pressure.** **TYPICAL.** If the per-core area budget is tight, the MUL/DIV unit competes with the L1 cache, branch predictor, and ROB for die area. Iterative implementations (subtractive DIV, iterative MUL) are most competitive in this regime.
|
||||
- **Power-grid and clock distribution.** **TYPICAL.** A replicated fast pipelined multiplier across 128 cores is a substantial clock-tree and switching load. **INSUFFICIENT EVIDENCE:** Power delivery (IR drop), clock skew across the replicated load, and dynamic power density must be analyzed for the replicated load; no quantitative estimates are made here.
|
||||
- **Scan and BIST.** **TYPICAL.** A 128× replicated unit implies 128× the scan-chain length (if scan is per-core) or a partitioned BIST architecture. The choice affects DFT area and test time.
|
||||
- **Fault tolerance.** **TYPICAL.** A defect in the MUL/DIV unit is potentially a defect in all 128 cores. This argues for either a hardened, characterized macro or built-in redundancy / sparing, depending on yield targets.
|
||||
- **Timing variation.** **TYPICAL.** Across-die process variation affects 128 replicated units independently. A design that is timing-marginal at one corner may fail at another. Iterative designs are less sensitive to per-unit timing variation than deep-pipelined arrays.
|
||||
|
||||
**PROPOSAL:** Favor an implementation whose area × power product is small, since both costs multiply by 128. Prefer regular, hardenable structures over irregular ones that resist replication.
|
||||
|
||||
## Performance Considerations
|
||||
|
||||
MUL/DIV performance interacts with the rest of the core in two ways:
|
||||
|
||||
1. **Latency exposure.** **FACT.** In an in-order core, MUL latency stalls the front-end unless explicit forwarding is provided. In an OoO core, MUL latency is amortized as long as other independent instructions exist.
|
||||
2. **Throughput-bound on dependent chains.** **FACT.** Tight loops of dependent MUL/DIV operations (e.g., big-integer arithmetic, hash functions, matrix multiplies) are throughput-bound, not latency-bound. For these, pipelined MUL at 1/cycle is essential.
|
||||
|
||||
**INSUFFICIENT EVIDENCE:** Without a target workload, no quantitative performance claim can be made.
|
||||
|
||||
## Area Considerations
|
||||
|
||||
Relative area ordering, qualitative only (no synthesis data available):
|
||||
|
||||
- Iterative shift-add MUL + iterative subtractive DIV (separate): smallest.
|
||||
- Iterative radix-4 MUL (one row/cycle) + iterative subtractive DIV (separate): small.
|
||||
- Shared iterative MUL/DIV (Approach 8): small to moderate, depending on the degree of datapath sharing and the control overhead; the area advantage over split iterative is modest for RV64.
|
||||
- Pipelined radix-4 MUL + iterative subtractive DIV (split): moderate.
|
||||
- Pipelined radix-4 MUL + 2-bit-per-cycle naïve DIV: moderate.
|
||||
- Pipelined radix-4 MUL + radix-4 SRT DIV: large.
|
||||
- Pipelined radix-8 MUL + Newton-Raphson / Goldschmidt DIV: largest.
|
||||
|
||||
**INSUFFICIENT EVIDENCE:** The XH-1 per-core area budget must be defined before any of the above can be quantified in absolute terms. Absolute area figures require a process node and a synthesis flow.
|
||||
|
||||
## Power and Energy Considerations
|
||||
|
||||
Relative energy-per-operation ordering, qualitative only:
|
||||
|
||||
- **Iterative:** **TYPICAL.** Low per-cycle power, but high per-operation energy × time product (many cycles).
|
||||
- **Pipelined MUL + iterative DIV:** **TYPICAL.** Low per-op MUL energy, high per-op DIV energy.
|
||||
- **Pipelined MUL + SRT DIV:** **TYPICAL.** High per-cycle power, lower per-op energy than iterative.
|
||||
- **Pipelined MUL + Newton-Raphson / Goldschmidt DIV:** **TYPICAL.** Per-op energy is a small multiple of MUL energy; competitive.
|
||||
|
||||
**PROPOSAL:** Energy-per-operation, not energy-per-cycle, is the more relevant metric for a workload-bound design. A pipelined radix-4 MUL with an iterative subtractive DIV is a reasonable energy-vs-area compromise, contingent on the workload.
|
||||
|
||||
**INSUFFICIENT EVIDENCE:** Quantitative power and energy figures require a process node, a clock target, and a workload trace.
|
||||
|
||||
## Implementation Considerations
|
||||
|
||||
- **Wallace/Dadda tree synthesis risk.** **TYPICAL.** Compressor trees are known to be hard to place-and-route at high frequency on modern nodes; poor placement can introduce unexpected delay. This favors iterative or low-radix designs for first-pass XH-1 silicon. *This is an engineering judgment, not an established fact; specific delay figures depend on the synthesis flow and library, and are not asserted here.*
|
||||
- **PPG and Booth recoder verification.** **TYPICAL.** Radix-4 PPGs are well-understood and tractable to verify; radix-8+ PPGs require more corner cases. The `MULH` signed×signed upper-half path requires explicit verification of sign-extended partial products and is not a free byproduct of the unsigned array. The `MULHSU` path requires explicit sign-handling verification (Baugh-Wooley or Modified Booth) and is not a vanilla 65×64 unsigned array.
|
||||
- **SRT selection table correctness.** **TYPICAL.** Selection tables must be proven correct for all operand values. This typically requires exhaustive enumeration for reduced operand widths and extensive directed testing for full width.
|
||||
- **RISC-V division-by-zero and overflow conventions.** **FACT.** `DIV` and `DIVU` return `-1` (all bits set) on divide-by-zero; `REM` and `REMU` return the dividend on divide-by-zero. On signed overflow (`INT64_MIN / -1`), `DIV` returns `INT64_MIN` and `REM` returns 0. These must be implemented explicitly; the design must not raise a trap.
|
||||
- **Pipeline interlocks.** **FACT.** If MUL is pipelined, the in-order scoreboard or the OoO wakeup logic must handle in-flight MULs; this is a per-instruction-record cost.
|
||||
- **Physical design replication.** **TYPICAL.** A 128-core replicated design benefits from a small, regular MUL/DIV unit. Highly irregular SRT selection logic and complex PPGs resist clean replication and are a physical-design risk under replication.
|
||||
|
||||
## Verification Considerations
|
||||
|
||||
Verification cost is a dominant non-silicon cost in modern CPU design. The MUL/DIV unit is a known hot-spot for bugs:
|
||||
|
||||
- **Radix-2/4 shift-add and subtractive:** **TYPICAL.** Lowest. Exhaustive simulation at reduced widths is tractable.
|
||||
- **Radix-4 Booth + iterative DIV:** **TYPICAL.** Moderate. The hard cases are the signed division edge cases (division-by-zero, `INT64_MIN / -1`), the `MULH` signed×signed upper-half path (sign-extended partial products, not a free byproduct of the unsigned array), and the `MULHSU` sign-extension (Baugh-Wooley or Modified Booth, not a vanilla unsigned array). The unsigned 64×64→128 array, once verified, covers `MUL`, `MULHU`, and the bulk of `MULHSU`; `MULH` is a separate verification pass; `MULHSU` requires explicit sign-handling verification.
|
||||
- **SRT:** **TYPICAL.** High. SRT selection-table bugs are famous in industry (the Pentium FDIV bug was a missing entry in the PLA implementing the radix-4 SRT lookup table of the floating-point divider; the lessons transfer to integer SRT by analogy, with the caveat that FP and integer SRT use different quotient-digit sets and selection functions). Verification typically requires formal proofs of the selection function over reduced widths and extensive directed testing.
|
||||
- **Newton-Raphson / Goldschmidt:** **TYPICAL.** High. Rounding-toward-zero correctness across all 64-bit input pairs requires a careful proof of the final fixup step.
|
||||
- **Replication factor:** **TYPICAL.** A bug in the MUL/DIV unit that escapes verification is a bug in all 128 cores (silicon defect exposure). However, the verification effort itself is not 128×: the unit is verified once at the unit level, and the core-level integration is verified once (cores are identical replications). The risk is concentrated exposure, not multiplied effort. This is the standard methodology for replicated unit-level verification in commercial designs.
|
||||
|
||||
**Preliminary verification strategy** (to be refined once the design choice is made):
|
||||
|
||||
- **Unit-level:** Exhaustive simulation at reduced operand widths (e.g., 8, 12, 16 bits) for the core datapath. Formal equivalence checking between the RTL and a reference model written in a high-level specification language (e.g., Bluespec, Scala, or a C reference). For SRT or Newton-Raphson, formal proof of the selection function or the convergence step.
|
||||
- **Directed corner-case suite:** Explicit tests for division-by-zero (both `DIV`/`DIVU` and `REM`/`REMU` paths), signed overflow (`INT64_MIN / -1`, both quotient and remainder), `MULH` against a cross-checked reference (Baugh-Wooley or equivalent), `MULHSU` against a cross-checked reference (with explicit sign-handling verification), and the `MUL`-then-truncate boundary.
|
||||
- **Random / constrained-random:** At full width, comparing against a software reference. Coverage targets on the Booth recoder, PPG, and selection table.
|
||||
- **Integration:** Per-core pipeline integration verified once at the core level (not 128×), since cores are identical.
|
||||
- **Post-silicon:** Microarchitectural validation suite, focused on MUL/DIV-intensive kernels (big-integer arithmetic, hashes, polynomial multiplications).
|
||||
|
||||
**PROPOSAL:** Verification cost should be a first-class input to the MUL/DIV design decision. A simpler, more easily verified design is preferable to a faster, harder-to-verify one, unless the performance delta is necessary for the XH-1 target workload.
|
||||
|
||||
## Software Considerations
|
||||
|
||||
- **Compiler behavior.** **FACT.** GCC and LLVM routinely use shift-and-add sequences for multiplication by small constants, and they may either emit `MUL` instructions or inline expansions depending on the cost model. The cost of a long-latency MUL or DIV may push the compiler toward suboptimal code sequences, but the magnitude of this effect is workload- and compiler-version-dependent and should not be assumed.
|
||||
- **Library code.** **FACT.** libgcc, compiler-rt, and the C library contain hand-written multiplication and division routines (e.g., for `__divti3`, `__udivti3`). These may use the MUL/DIV instructions in tight loops, exposing the unit's throughput directly.
|
||||
- **System software.** **TYPICAL.** Hash functions, AES-GCM polynomial multiplication, modular exponentiation in cryptography, and big-integer arithmetic in language runtimes are all MUL-throughput sensitive.
|
||||
- **128-core interaction.** **ASSUMPTION.** In a homogeneous configuration where the system stack (kernel, hypervisor) runs on some subset of the 128 cores, there is no asymmetric design implication: every core is equal in the default per-core configuration. **INSUFFICIENT EVIDENCE:** Whether the 128-core fabric is homogeneous or heterogeneous (e.g., application cores plus management or I/O cores) is an open question (see Open Questions); the assumption of homogeneity is provisional.
|
||||
|
||||
**INSUFFICIENT EVIDENCE:** XH-1's target software ecosystem is unknown. Bare-metal, Linux, or richer OS support each imply different MUL/DIV usage profiles.
|
||||
|
||||
## Recommendation
|
||||
|
||||
**No final recommendation is made at this time.** The MUL/DIV design is contingent on three decisions that are not yet established in the XH-1 repository:
|
||||
|
||||
1. The core microarchitecture (in-order vs. OoO, issue width, pipeline depth).
|
||||
2. The per-core area and power budget.
|
||||
3. The target workload mix and software ecosystem.
|
||||
|
||||
**Conditional recommendations** (to be revisited once the above are known):
|
||||
|
||||
- **If XH-1 targets an in-order core with tight area budget:** **PROPOSAL.** A separate iterative shift-add MUL (one row/cycle) and an iterative subtractive DIV is the smallest, easiest-to-verify choice. MUL/DIV latency will be high; whether this is acceptable depends on the forwarding path and the workload. **The forwarding path is a real and non-trivial cost:** the bypass network from the MUL/DIV output to dependent consumers (ALU, branch, load/store) must be sized to the full MUL/DIV latency, and the wiring may dominate the area of the iterative MUL/DIV unit itself. This cost should be quantified before the iterative unit is chosen. The shared-iterative-unit approach (Approach 8) is *not* recommended for RV64 by default, given the structural mismatch between MUL and DIV datapaths and the modest area advantage over a split iterative design.
|
||||
- **If XH-1 targets a wide-issue OoO core with relaxed area budget:** **PROPOSAL.** A pipelined radix-4 Booth multiplier (3–6 stage pipeline) paired with a 2-bit-per-cycle naïve divider or a radix-4 SRT divider. MUL throughput at 1/cycle; DIV throughput at 1/16 to 1/32 cycle. **Reconciliation note:** the cross-cutting recommendation below defers SRT until formally proven correct; if SRT cannot be formally verified on reduced widths within the project timeline, the 2-bit-per-cycle naïve divider is the preferred DIV companion. The SRT recommendation is conditional on the verification investment being made.
|
||||
- **If the per-core area budget is moderate and the workload is balanced:** **PROPOSAL.** A pipelined radix-4 MUL with an iterative subtractive DIV (the "split" approach, Alternative A) is a strong compromise.
|
||||
|
||||
**Cross-cutting recommendation:** Defer any design that depends on a non-trivial final correction step (SRT error correction, Newton-Raphson / Goldschmidt rounding fixup) until those algorithms have been formally specified and proven correct on reduced widths.
|
||||
|
||||
## Confidence
|
||||
|
||||
| Topic | Confidence |
|
||||
|---|---|
|
||||
| Taxonomy of MUL/DIV approaches | High (canonical, well-known) |
|
||||
| Relative area, power, and energy rankings | Low to medium (qualitative ordering only; no synthesis data) |
|
||||
| Per-cycle latency numbers | Low to medium (typical, but node- and target-frequency-dependent; ranges are illustrative, not tight bounds) |
|
||||
| Suitability for XH-1 specifically | **Low** (XH-1 microarchitecture not yet established) |
|
||||
| Specific recommendation | **Insufficient evidence** to commit |
|
||||
|
||||
## Open Questions
|
||||
|
||||
1. Is the XH-1 core in-order or out-of-order? Issue width? Pipeline depth? Target frequency?
|
||||
2. What is the per-core area budget, in absolute terms or as a fraction of die area?
|
||||
3. What is the target workload mix (cryptographic, HPC, server, general-purpose, embedded)?
|
||||
4. Will XH-1 implement the "B" (bitmanip) extension, which adds many pseudo-multiplication operations? If so, the MUL/DIV unit may be partially reused or augmented.
|
||||
5. Is the 128-core fabric a true SMP (Linux runs on all 128 cores), or is it a heterogeneous mix (e.g., application cores + management cores)? This affects MUL/DIV uniformity requirements.
|
||||
6. What is the target process node? Node geometry strongly affects the area and timing trade-offs of Booth vs. array multipliers.
|
||||
7. Will the MUL/DIV unit be hardened as a macro and replicated 128 times, or will it be synthesized per-core? Hardening is favored for area/power and replication regularity; synthesis-per-core favors design flexibility.
|
||||
8. Is there an asymmetric verification budget for the MUL/DIV unit, or is verification cost unconstrained?
|
||||
9. Does XH-1 need a 64×64→128-bit MUL for the upper-multiply operations to be single-cycle, or is multi-cycle acceptable?
|
||||
10. What is the expected frequency of `DIV`/`REM` instructions in the target workload? Server workloads may have many fewer than scientific workloads.
|
||||
11. Will the MUL/DIV unit be shared with the FP pipeline (as in BOOM and some SiFive designs, though not in Rocket Chip), or kept private to the integer pipeline?
|
||||
12. Will XH-1 adopt a shared multi-cycle divider across cores (Alternative F), or is a strictly per-core MUL/DIV unit required?
|
||||
|
||||
## Sources
|
||||
|
||||
- Canonical computer-arithmetic references: Parhami, *Computer Arithmetic*; Ercegovac and Lang, *Digital Arithmetic*; Hennessy and Patterson, *Computer Architecture*. These cover the techniques surveyed above in standard form. **Caveat:** these texts underwrite the taxonomy and the mechanism descriptions. Specific per-cycle latency figures given in this document are **TYPICAL** values drawn from common practice in published RV64 designs; the cited texts provide the algorithmic background but do not, in their canonical editions, supply XH-1-specific latency or area numbers. No specific chapter or page is cited for the quantitative figures because the figures are not drawn from a single source.
|
||||
- Open RISC-V core implementations (Rocket, BOOM, XiangShan, SiFive) provide reference designs for radix-4 array multipliers, iterative and SRT dividers, and FP/integer multiplier sharing. These are cited as implementation exemplars, not as XH-1 references. **FACT.** Rocket Chip keeps the integer and FP multipliers as separate units; BOOM and some SiFive designs share partial structures. The blanket attribution of sharing to "Rocket, BOOM, and most SiFive cores" in the prior draft of this document overstates the case and is corrected here.
|
||||
- **INSUFFICIENT EVIDENCE:** No XH-1-internal sources exist for this topic. The repository document `research/03-core-design/mul-div-unit.md` is itself a stub ("SOON"), and all sibling documents in `research/03-core-design/` are stubs. No prior architectural decisions are recorded.
|
||||
- No external citations, papers, benchmarks, or measurements have been invented for this document. All comparative claims are framed as qualitative engineering estimates or as canonical background knowledge in the field of computer arithmetic.
|
||||
@@ -0,0 +1,84 @@
|
||||
VERDICT: FAIL
|
||||
|
||||
ISSUES:
|
||||
|
||||
1. **Inconsistent radix-4 SRT digit-set terminology.** The "Mechanism" paragraph for the radix-4 SRT divider says the unit "produces 2 bits per iteration" and "the quotient is held in a redundant (carry-save) form." A radix-4 SRT divider with a carry-save (CS) partial remainder produces **2 quotient bits per iteration in redundant form**, and the quotient-digit selection function actually produces one radix-4 digit per cycle, not a carry-save quotient digit. The standard formulations use either (a) a redundant partial remainder (carry-save) with a non-redundant quotient-digit selection, or (b) a redundant quotient register with selection from a small set (e.g., −2, −1, 0, +1, +2) and a separate partial-remainder datapath. The document conflates these two formulations ("carry-save quotient representation" and a "quotient-digit selection table") without specifying which is intended. The "redundant (carry-save) quotient" framing is non-standard and potentially incorrect depending on the SRT variant assumed. This is a substantive technical error in a section tagged FACT.
|
||||
|
||||
2. **Pentium FDIV bug description is inaccurate.** The document states the bug was a "missing entry... a '+2' entry that should have been present in the programmable logic array (PLA) implementing the selection function." The well-documented root cause was that the lookup table contained only 1,068 of the required 1,066 entries (actually 1,068 entries; some sources differ, but the documented defect was missing entries, not specifically a missing "+2" entry). More importantly, the divider was a **radix-4 SRT** divider but the characterization of the bug as a single missing "+2" digit in a PLA is oversimplified. The actual defect was missing entries in the PLA's truth table that caused some division operands to produce incorrect results. Claiming a specific missing "+2" entry as the cause, presented as FACT, is not accurate. The fact-tag is also inappropriate — the specific defect mechanism is disputed across sources; the fact that the bug existed and was an SRT PLA defect is the fact, not the specific entry. The document should mark the entry-level description as TYPICAL or remove it.
|
||||
|
||||
3. **Booth recoding row counts are non-canonical and possibly incorrect.** The document states radix-4 produces "32 partial-product rows" for a 64-bit operand (correct: ceil(64/2) = 32). It states radix-8 produces "22 data rows plus a separate sign-handling row for the most-significant window, totaling 23 rows in a fully sign-corrected implementation." Standard radix-8 Booth (Modified Booth encoding, scanning three bits at a time with the standard overlapping-window formulation) produces **ceil(64/3) = 22 partial-product rows** when using the canonical non-redundant radix-8 form, but the commonly cited form is **21 rows** (since 64/3 rounds down to 21 with a 2-bit MSB fragment handled by sign extension). The "+1 sign-handling row" framing is non-standard; sign handling in Modified Booth is incorporated into the existing rows via sign-extension of the MSB partial product, not as an extra row. This claim is misleading and presented as FACT.
|
||||
|
||||
4. **Radix-8 multiples claim is non-standard.** The document says "The PPG must produce multiples {0, ±1, ±2, ±3, ±4} of the multiplicand." Standard radix-8 Booth encoding (3-bit window) produces digits in the set {−4, −3, −2, −1, 0, +1, +2, +3, +4} (the ±4 multiple is generated for the case where the window is "100" or "011" with the canonical encoding). This is correct. However, the statement that "±3× and ±4× are typically generated via a carry-save adder" is partially correct but the description of "1× + 2× for ±3×, 2× + 2× or a dedicated shift-and-add for ±4×" is confusing — 2× + 2× is **not** a valid implementation of 4×; 4× is a left shift by 2 (i.e., wire routing, not addition). The claim that 4× is generated by "2× + 2×" is technically correct (it equals 4×) but is an unusual and suboptimal description; the canonical implementation is a left-shift by 2. The document presents an unusual implementation detail as TYPICAL without justification.
|
||||
|
||||
5. **Self-contradictory claim about "free" MULH and MULHSU in the same section.** The "Wide-operand upper-multiply (UMUL)" paragraph characterizes MULH* as having a "small" incremental cost over MUL. The "MULH and signed×signed handling" and "MULHSU and signed×unsigned handling" paragraphs then explicitly contradict this, stating MULH and MULHSU are "not free" and require dedicated datapaths and verification passes. These three claims are not strictly contradictory (the first says "small area adder," the latter two say "not free in the strict sense") but the document's "Caveat" in the UMUL paragraph says the shared-array pattern is the "common case," while the later paragraphs say the MULH datapath is structurally different. The reconciliation is incomplete: the document needs to clearly state whether the underlying array is genuinely shared (with sign-handling added at the edges) or whether MULH and MULHSU use distinct datapaths. As written, a reader cannot determine which it is. This is a substantive internal inconsistency in a section tagged FACT.
|
||||
|
||||
6. **Radix-2 SRT iteration count framing is wrong.** The document states "A radix-2 SRT divider also requires 64 selection steps in the worst case (with possible skipped steps on average, but the worst case governs the pipeline)." Radix-2 SRT is essentially equivalent to non-restoring division in the worst case (64 steps for 64-bit operands), so the claim is correct in count but the framing is misleading: radix-2 SRT's distinguishing feature is **not** the iteration count but the quotient-digit selection logic and the redundant remainder representation. A reader unfamiliar with SRT will come away thinking radix-2 SRT differs from non-restoring primarily in the iteration count, which is the opposite of the truth — they have the same worst-case iteration count, and the SRT framework exists precisely because of the redundant-form advantages. Additionally, Alternative C describes "Radix-2 SRT with Carry-Save Quotient + Software/Hardware Fixup" as "Faster than naive subtractive," but with 64 selection steps in the worst case (which governs the pipeline), this is not faster than subtractive division in the latency-critical sense. The claim is at best misleading and at worst wrong.
|
||||
|
||||
7. **The radix-4 SRT row in the Comparison table conflates "1/16" throughput with sustained throughput.** A radix-4 SRT divider has worst-case latency of ~16 selection cycles plus a final on-the-fly quotient-conversion step. Sustained throughput for a single non-pipelined radix-4 SRT divider is one divide per ~16 cycles, which the table captures as "~1/16 (DIV)." However, for a fully pipelined radix-4 SRT array, throughput could approach 1/cycle (with latency of ~16 cycles). The table's "~1/16 (DIV)" throughput is correct for a non-pipelined radix-4 SRT unit but the document does not clearly flag that a pipelined variant exists, and the table is silent on the pipelining assumption. The SRT divider in the "Pipelined radix-4 array MUL + radix-4 SRT DIV" row could be pipelined to achieve 1/cycle DIV throughput with ~16-cycle latency, which would change the comparison materially. The table is internally inconsistent with the row above (pipelined radix-4 MUL achieves 1/cycle MUL throughput), which establishes a pipelining baseline.
|
||||
|
||||
8. **Unsubstantiated quantitative claim about MULH "share the array" pattern.** The document states the "share the array" pattern "is the common case in published RV64 implementations (e.g., BOOM, XiangShan, some SiFive designs), but quantitative numbers are not asserted here." This is a substantive claim attributed to specific named designs (BOOM, XiangShan, some SiFive) without citation. The "share the array" characterization is a non-trivial microarchitectural decision; claiming it is the "common case" in three named designs without any citation (paper, documentation, or RTL inspection) is an unsupported factual claim. The document should mark this as INSUFFICIENT EVIDENCE or remove the named designs.
|
||||
|
||||
9. **The cross-cutting recommendation creates an internal contradiction with the OoO conditional recommendation.** The "Cross-cutting recommendation" defers SRT until formally proven correct, but the OoO conditional recommendation explicitly proposes SRT (with a reconciliation note). The reconciliation note says "if SRT cannot be formally verified on reduced widths within the project timeline, the 2-bit-per-cycle naïve divider is the preferred DIV companion." This is a self-contradictory structure: the recommendation both does and does not recommend SRT, depending on an unspecified future decision. A reader cannot determine which divider to design for. The structure of "Recommendation" should make a single defensible choice, not a conditional that effectively defers the choice.
|
||||
|
||||
10. **128-core verification claim is partially incorrect.** The document states "the integration with each core's pipeline is identical across replications and is verified once via the core-level verification environment; running the same integration suite 128× does not add coverage." This is technically correct for unit-level coverage but is misleading regarding integration verification. In commercial replicated-core designs (e.g., ARM Cortex-A, Intel Atom, AMD Zen), **core-level integration is typically verified per-core in a multi-core simulation environment** that exercises cross-core interactions, shared-resource contention, and coherence protocols. For a 128-core design, the cross-core verification environment is a distinct effort from the single-core integration environment and is **not** "verified once." The document conflates "single-core integration verified once" with "multi-core integration verified once," which are different scopes. The 128-core scalability claim should be qualified.
|
||||
|
||||
11. **Booth recoding verification claim is too strong.** The document says "Radix-4 PPGs are well-understood and tractable to verify" (TYPICAL). This is a substantial underestimation: Modified Booth encoding for signed multiplication (especially MULH, MULHSU) has a long history of subtle bugs, including sign-extension errors in the MSB partial product. The "tractable to verify" framing understates the verification effort and may bias the design choice away from radix-4 in contexts where the verification cost is actually comparable to radix-8.
|
||||
|
||||
12. **Goldschmidt vs. Newton-Raphson distinction is loosely drawn.** The document correctly notes the two are distinct but the description of Goldschmidt ("Each iteration multiplies both the partial remainder and the partial quotient by a correction factor") and Newton-Raphson ("Iteratively refine an initial estimate of the divisor's reciprocal... then multiply the dividend by the refined reciprocal") is correct at a high level. However, the "Use case" for both is identical ("Designs with a fast pipelined multiplier that want a small, fast divider"), which obscures the practical distinction: Goldschmidt converges linearly (each iteration adds a fixed number of correct bits) while Newton-Raphson converges quadratically. This affects the iteration count, latency, and verification cost materially. The document should distinguish them more sharply.
|
||||
|
||||
13. **"Two rows per cycle" iterative radix-4 is described but not represented in the comparison table.** The §2 (Approach 2) section discusses "two rows per cycle" iterative radix-4 MUL at ~16 cycles with non-trivial area/timing cost. This configuration does not appear as a row in the Comparison table. A reader cannot assess the trade-off of this intermediate option. The Comparison table is incomplete.
|
||||
|
||||
14. **The "Latency-Tolerant In-Order MUL/DIV" alternative (H) is presented as an alternative but its central cost (forwarding path) is only briefly noted.** The document states the forwarding network "may dominate the area of the iterative MUL/DIV unit itself" but does not quantify or compare this cost to the alternatives. The alternative is essentially a description of a known cost that any in-order iterative design must absorb, and the document should integrate this into the main analysis rather than treating it as a separate alternative.
|
||||
|
||||
15. **"1× + 2× for ±3×" PPG detail is non-standard.** The document states that ±3× in radix-8 PPG is generated via "1× + 2×" (i.e., 3× multiplicand = multiplicand + 2× multiplicand). This is correct but the canonical implementation uses a **carry-save adder** that produces the 3× in carry-save form (sum and carry vectors) so the 3× multiple does not require an explicit addition in the critical path. The document does not mention this and implies an explicit addition, which is misleading for a high-performance design.
|
||||
|
||||
16. **Dadda/Wallace tree "8-12 CSA levels" claim is unsupported.** The document states the radix-4 array's "one combinational tree of approximately 8–12 CSA levels plus a final CPA (the level count is design- and library-specific; the figure is an order-of-magnitude estimate, not a precise bound)." This is appropriately qualified as TYPICAL and order-of-magnitude. However, the actual CSA level count for a 32-row radix-4 Wallace tree summing to two rows is closer to **log_1.5(32/2) ≈ 8 levels** (using the 1.5 reduction ratio for full adders), or up to ~10 with the irregularity of practical Dadda reduction. The "8-12" range is defensible but on the high end; this is a minor issue but should be noted.
|
||||
|
||||
17. **"Wallace/Dadda tree synthesis risk" is presented as engineering judgment but framed as a generalizable TYPICAL claim.** The document says "Compressor trees are known to be hard to place-and-route at high frequency on modern nodes; poor placement can introduce unexpected delay." This is a real concern in practice but the framing as a generalizable TYPICAL claim is too strong. On modern 7nm and 5nm nodes with automated place-and-route, Wallace/Dadda trees are routinely synthesized to high frequencies. The "TYPICAL" tag is overreaching for a claim that is highly node- and tool-flow-dependent.
|
||||
|
||||
18. **The "FDIV bug transferred by analogy" caveat is incomplete.** The document correctly notes that FP and integer SRT use different quotient-digit sets and selection functions, so the lessons are transferred by analogy. However, the caveat should also note that integer SRT dividers do not have the same scale of historical bugs as FP SRT dividers in published designs, and the Pentium FDIV bug is the canonical example precisely because FP dividers are the most prominent historical case. The "transferred by analogy" framing is correct but understates that the analogy is the primary basis for the SRT verification concern, not direct evidence of integer SRT bugs.
|
||||
|
||||
19. **The "Two rows per cycle" iterative radix-4 area description is internally inconsistent.** The document states "the area cost of the second CSA is not a simple 'doubling': the second CSA operates on the full sum-and-carry width of the first, so the additional area is closer to the cost of one full-width CSA, and the per-cycle critical path lengthens." This is correct but the conclusion "the area cost is closer to the cost of one full-width CSA" is a strange way to express it. A second full-width CSA on top of the first is, by definition, approximately one full-width CSA of additional area (i.e., a doubling of CSA area). The phrasing is convoluted but technically not wrong; it could be clearer.
|
||||
|
||||
20. **The document does not discuss the latency of the quotient-conversion step in radix-4 SRT or its on-the-fly variant.** Standard radix-4 SRT dividers use an on-the-fly quotient-conversion (OFC) technique to avoid the final conversion step latency. The document discusses the "final quotient-conversion step" as if it is always a separate cycle, but OFC can absorb it into the selection-step critical path. This is a substantive omission in the description of SRT divider latency and the comparison table's "16 + conversion" framing.
|
||||
|
||||
21. **The document is missing alternatives.** Specifically:
|
||||
- **Combined MUL/MULH datapath with early termination** (similar to the variable-latency divider) is not discussed.
|
||||
- **Use of a Wallace/Dadda tree for the upper half only (MULH path)** with a separate simpler datapath for MUL, which is a known low-area technique.
|
||||
- **Dual-issue MUL pipelines** (two independent MUL units, each with 1/2 throughput) are not discussed as a throughput-extension alternative.
|
||||
- **DSP-style saturating multiply** for embedded use cases is not discussed.
|
||||
- **Use of the FP multiplier for integer multiplication with format conversion** (when the FP pipeline is idle) is not discussed as a software/compiler-managed alternative.
|
||||
|
||||
22. **The document does not discuss the impact of the FP unit's existence on the integer MUL/DIV design.** The document mentions FP sharing in Alternative G but does not analyze the impact on the integer MUL/DIV design if the FP unit does not exist (e.g., XH-1 has no FP) or if the FP pipeline is shallow (e.g., 2-stage FPU where sharing provides minimal benefit). The interaction is one-sided.
|
||||
|
||||
REQUIRED_FIXES:
|
||||
|
||||
- Fix the radix-4 SRT digit-set and quotient-representation description (Issue 1). Specify whether the partial remainder is in carry-save form or non-redundant form, and whether the quotient digits are redundant or non-redundant. Use the standard formulation.
|
||||
|
||||
- Remove or qualify the specific "+2 entry" claim about the Pentium FDIV bug (Issue 2). Mark the specific defect mechanism as TYPICAL or remove it. The fact that the bug was a missing-entry defect in the SRT PLA is sufficient.
|
||||
|
||||
- Correct the radix-8 row count (Issue 3) and clarify the sign-handling mechanism. Use the standard Modified Booth encoding formulation: 22 rows (or 21 with MSB fragment handled by sign extension), with sign-extension incorporated into the existing rows, not a separate row.
|
||||
|
||||
- Fix the radix-8 4× multiple description (Issues 4, 15). The canonical 4× multiple is a left-shift by 2, not an addition. Also clarify the ±3× implementation as a carry-save form, not an explicit addition.
|
||||
|
||||
- Reconcile the MULH/MULHSU "shared array" claims (Issue 5). Explicitly state whether the MULH/MULHSU datapaths share the underlying array with sign-handling at the edges, or whether they are distinct datapaths. Do not present these as three separate claims that the reader must reconcile.
|
||||
|
||||
- Fix the radix-2 SRT framing (Issue 6). State clearly that radix-2 SRT has the same worst-case iteration count as non-restoring division; the advantage of SRT is in the redundant-form datapath, not the iteration count. Correct or remove the Alternative C "Faster than naive subtractive" claim.
|
||||
|
||||
- Clarify the SRT pipelining assumption in the Comparison table (Issue 7). State whether the SRT divider in the table is pipelined or non-pipelined, and add a pipelined SRT row if the MUL row is pipelined.
|
||||
|
||||
- Remove or qualify the named-design attribution for the MULH shared-array pattern (Issue 8). The claim "common case in published RV64 implementations (e.g., BOOM, XiangShan, some SiFive designs)" is unsupported by citations. Either cite specific sources or remove the named designs.
|
||||
|
||||
- Resolve the cross-cutting recommendation contradiction (Issue 9). Make a single defensible conditional recommendation rather than a conditional that effectively defers the choice.
|
||||
|
||||
- Qualify the 128-core verification claim (Issue 10). Distinguish single-core integration verification (done once) from multi-core cross-core verification (a distinct effort). Do not state that "running the same integration suite 128× does not add coverage" without distinguishing the scopes.
|
||||
|
||||
- Add the missing radix-4 "two rows per cycle" row to the Comparison table (Issue 13) and add the missing alternatives listed in Issue 21.
|
||||
|
||||
- Add discussion of on-the-fly quotient conversion for radix-4 SRT (Issue 20).
|
||||
|
||||
- Address the missing FP-interaction analysis (Issue 22).
|
||||
|
||||
- Address the radix-4 PPG verification cost understatement (Issue 11) and the Goldschmidt/Newton-Raphson distinction (Issue 12).
|
||||
|
||||
CONFIDENCE: HIGH
|
||||
+227
File diff suppressed because one or more lines are too long
+308
@@ -0,0 +1,308 @@
|
||||
# Multiply/Divide Unit Design for the XH-1 Processor
|
||||
|
||||
## Status
|
||||
|
||||
**Stub document — research in progress.** No architectural decisions have been finalized. The XH-1 repository contains no prior decisions on the multiply/divide (MUL/DIV) unit. All content below is a structured engineering analysis of design options, framed as proposals and open questions.
|
||||
|
||||
## Abstract
|
||||
|
||||
This document investigates the design of the integer multiply and divide unit (MUL/DIV) for the XH-1, a custom 128-core RISC-V processor. The MUL/DIV unit is a latency-critical but throughput-amortizable functional unit. Its design has outsized impact on the pipeline depth, area per core, and verification effort, all of which compound across 128 replicated cores. We survey the principal implementation strategies (iterative shift-and-add, radix-4/8 Booth, array multipliers, Goldschmidt/Newton-Raphson dividers, and iterative reciprocal-multiplication) and analyze their trade-offs with respect to latency, throughput, area, power, and verification complexity. The document is intentionally non-prescriptive: the XH-1 has not yet established a baseline for core microarchitecture (in-order vs. out-of-order, pipeline depth), so recommendations are conditional on those upstream decisions.
|
||||
|
||||
## Research Question
|
||||
|
||||
> What is the appropriate microarchitecture for the integer MUL/DIV unit in the XH-1 core, given that the unit will be replicated 128 times, must support the RV64M extension, and must satisfy XH-1's target latency, area, and power budget — the latter two of which are not yet defined?
|
||||
|
||||
Specifically:
|
||||
|
||||
- How should the unit be pipelined relative to the rest of the core?
|
||||
- Should it share silicon with adjacent datapath structures (ALU, register file read/write ports) or remain isolated?
|
||||
- How should division throughput be traded against area and verification cost?
|
||||
- What is the impact of the choice on multi-core contention for shared execution resources, if any?
|
||||
|
||||
## Background
|
||||
|
||||
The RISC-V "M" standard extension defines four signed/unsigned variants of multiply and two signed/unsigned variants of divide plus remainder:
|
||||
|
||||
- `MUL` (lower 64 bits of signed × signed)
|
||||
- `MULH`, `MULHU`, `MULHSU` (upper 64 bits, various sign combinations)
|
||||
- `DIV`, `DIVU` (signed/unsigned quotient)
|
||||
- `REM`, `REMU` (signed/unsigned remainder)
|
||||
|
||||
Critical properties of this ISA slice that drive microarchitecture:
|
||||
|
||||
- **Wide-operand upper-multiply (UMUL):** Producing the upper 64 bits of a 64×64→128-bit product is approximately twice the area of the lower-64-bit result, since it cannot reuse a 64×64→64 multiplier without additional logic.
|
||||
- **Division latencies:** A non-trivial iterative divider for 64-bit operands requires 32 to 64 reduction steps, dominating pipeline depth if fully combinational.
|
||||
- **Signed semantics:** MULHSU and DIVU/REMU require careful sign-handling that complicates early-exit logic.
|
||||
- **Throughput vs. latency decoupling:** RISC-V MUL/DIV results are written to the integer register file, so latency can be hidden by OoO scheduling — but in an in-order core, latency is directly exposed to the front-end.
|
||||
|
||||
The XH-1 is a 128-core machine. Decisions in the MUL/DIV unit replicate 128×, so per-core area and verification effort dominate. There is no public XH-1 microarchitecture document in the repository that would constrain the MUL/DIV design (e.g., pipeline depth, issue width, in-order/OoO).
|
||||
|
||||
**INSUFFICIENT EVIDENCE:** Pipeline depth, issue width, in-order vs. OoO status, target frequency, and target process node for XH-1 are all unknown.
|
||||
|
||||
## Existing Approaches
|
||||
|
||||
The following approaches are well-established in the literature and commercial designs. No specific citation is invented here; references to canonical techniques are sufficient.
|
||||
|
||||
### 1. Shift-and-Add Multiplier (Iterative)
|
||||
|
||||
- **Mechanism:** A 64-bit multiplier is implemented as a 64-cycle state machine, one partial-product bit per cycle, with a 129-bit accumulator.
|
||||
- **Latency:** 64 cycles (signed) or 32 cycles (sign-preconditioned) before result available.
|
||||
- **Throughput:** 1 multiply per ~64 cycles.
|
||||
- **Area:** Very small. Roughly one 64-bit adder + one shifter + one 129-bit accumulator register.
|
||||
- **Power:** Low. Minimal clocked area.
|
||||
- **Verification:** Low complexity. Straightforward to model and exhaustively test small operand widths.
|
||||
- **Use case:** Embedded in-order cores (e.g., RV32I designs where MUL/DIV are infrequent and latency-tolerant). Generally unsuitable for high-performance 64-bit cores.
|
||||
|
||||
### 2. Radix-4 Booth-Encoded Multiplier (Array or Iterative)
|
||||
|
||||
- **Mechanism:** Booth recoding reduces partial products to ceil(n/2). A radix-4 scheme halves the iteration count. Implemented either as a combinational Wallace/Dadda tree, or iteratively over ceil(64/2) = 32 cycles.
|
||||
- **Latency:** 32 cycles iterative, ~8–16 cycles pipelined (depending on adder tree depth).
|
||||
- **Throughput:** 1 multiply per 32 cycles iterative, 1/cycle fully pipelined.
|
||||
- **Area:** Moderate. Recoder + partial-product array + carry-save adder tree + final CPA.
|
||||
- **Power:** Moderate to high. The Wallace tree toggles aggressively.
|
||||
- **Verification:** Moderate. Corner cases around operand sign and the `MULHSU` path require careful directed tests.
|
||||
- **Use case:** Mid- and high-performance cores; the default choice for most modern desktop-class RV64 cores.
|
||||
|
||||
### 3. Radix-8 or Radix-16 Booth Multiplier
|
||||
|
||||
- **Mechanism:** Higher-radix Booth recoding (radix-8, radix-16) further reduces partial products by 3× or 4× at the cost of harder partial-product generation (3× or 5× multiples).
|
||||
- **Latency:** Fewer pipeline stages for the same throughput target, or higher throughput for the same area.
|
||||
- **Throughput:** 1/cycle or better.
|
||||
- **Area:** Larger partial-product generator; smaller adder tree. Net area roughly comparable to radix-4.
|
||||
- **Power:** Mixed. Fewer adder levels, but more complex PP generation.
|
||||
- **Verification:** Higher. Radix-8+ PPGs have more corner cases.
|
||||
- **Use case:** High-frequency designs where the critical path is in the adder tree, not the PPG.
|
||||
|
||||
### 4. Iterative Subtractive Divider (Non-Restoring / Restoring)
|
||||
|
||||
- **Mechanism:** Standard shift-subtract over 32 to 64 cycles. Produces quotient (and optionally remainder) one bit per cycle.
|
||||
- **Latency:** 32 to 64 cycles.
|
||||
- **Throughput:** 1 divide per 32–64 cycles.
|
||||
- **Area:** Small. A 64- or 65-bit ALU with quotient/remainder registers.
|
||||
- **Power:** Low.
|
||||
- **Verification:** Moderate. The signed-dividend/divisor path and the division-by-zero convention (RISC-V returns -1) require explicit verification.
|
||||
- **Use case:** In-order cores with low expected DIV frequency.
|
||||
|
||||
### 5. Radix-4/8 SRT Divider
|
||||
|
||||
- **Mechanism:** Higher-radix SRT produces 2 or more quotient bits per iteration by selecting one of several shifted multiples of the divisor.
|
||||
- **Latency:** 16 to 32 cycles.
|
||||
- **Throughput:** One divide per 16–32 cycles.
|
||||
- **Area:** Substantially larger than subtractive. Requires a redundant quotient representation (carry-save) and a quotient-digit selection table.
|
||||
- **Power:** Higher.
|
||||
- **Verification:** High. SRT selection tables have notoriously tricky corner cases, and partial-quotient error correction is a known source of silicon bugs in commercial designs.
|
||||
- **Use case:** High-frequency OoO cores needing higher DIV throughput.
|
||||
|
||||
### 6. Goldschmidt / Newton-Raphson Divider (Reciprocal Multiplication)
|
||||
|
||||
- **Mechanism:** Iteratively refine an initial reciprocal estimate, then multiply by the dividend. Effectively performs division using a fast multiplier.
|
||||
- **Latency:** A few iterations of multiply-add. Amortized low latency if the multiplier is fast.
|
||||
- **Throughput:** Bounded by the multiplier.
|
||||
- **Area:** Reuses the multiplier. Adds a small amount of control logic and lookup-table ROM for the initial estimate.
|
||||
- **Power:** Low marginal cost on top of the multiplier.
|
||||
- **Verification:** High. The accuracy of the final iteration relative to the rounding-mode requirements (RISC-V round-toward-zero) is a common bug source.
|
||||
- **Use case:** Designs that already have a fast pipelined multiplier and want a small, fast divider.
|
||||
|
||||
### 7. Shared MUL/DIV Iterative Unit (Multiplexed)
|
||||
|
||||
- **Mechanism:** A single iterative datapath (shift-add) handles both MUL and DIV by reconfiguring its datapath between operations.
|
||||
- **Latency:** Same as iterative, but the unit cannot serve MUL and DIV concurrently.
|
||||
- **Throughput:** MUL and DIV contend for the same unit.
|
||||
- **Area:** Smallest. The most area-efficient option.
|
||||
- **Power:** Lowest.
|
||||
- **Verification:** Low to moderate. Single datapath, but dual-purpose microcode/sequencer.
|
||||
- **Use case:** Cost-sensitive cores.
|
||||
|
||||
## Alternative Designs
|
||||
|
||||
Beyond the standard taxonomy, several non-obvious variants are worth considering:
|
||||
|
||||
### A. Pipelined Multiplier + Sequential Divider (Split Unit)
|
||||
|
||||
A fully pipelined radix-4 multiplier (e.g., 4-cycle latency) is paired with an independent 32-cycle sequential divider. The two units have separate reservation/issue logic.
|
||||
|
||||
- **Rationale:** MUL is far more frequent than DIV in most workloads. Giving MUL a fast pipelined datapath and DIV a slow sequential one matches the workload asymmetry.
|
||||
- **Cost:** Two reservation-station or issue-queue entries (relevant if the core is OoO); two small register files for intermediate state.
|
||||
|
||||
### B. Variable-Latency Divider (Speculative / Early-Quit)
|
||||
|
||||
Produce a 32-bit approximate quotient, sign-extend, and perform a single correction step. In many real programs the divisor has small magnitude, and the early-quit path saves cycles.
|
||||
|
||||
- **Risk:** Variable latency complicates the scoreboard / wakeup logic in an OoO core, or stalls the pipeline in an in-order core.
|
||||
|
||||
### C. Radix-2 SRT with Carry-Save Quotient + Software Fixup
|
||||
|
||||
A radix-2 SRT with a 1-bit-per-cycle selection function and a carry-save quotient register. Faster than naive subtractive, simpler than radix-4 SRT.
|
||||
|
||||
- **Cost:** The quotient-to-2's-complement conversion is on the critical path or requires a final fixup step.
|
||||
|
||||
### D. Pipelined 2-bit-per-cycle "Naïve" Divider
|
||||
|
||||
Iterate two shift-subtract steps per cycle. Doubles divider throughput without an SRT selection table. Latency roughly halved vs. radix-1.
|
||||
|
||||
- **Cost:** Two 64-bit adders and a longer critical path. Area is moderate.
|
||||
|
||||
### E. Iterative Multiplier (Pipelined) + Lookup-Table Reciprocal + Multiply
|
||||
|
||||
A small reciprocal ROM (e.g., 12-bit seed) followed by a Newton-Raphson refinement step using the iterative multiplier. Yields a divider latency of a few multiplier iterations (e.g., 4–8 cycles).
|
||||
|
||||
- **Cost:** The Newton iteration must converge to enough bits of reciprocal precision to guarantee correct rounding toward zero for all 64-bit inputs. This typically requires a final exact-fixup step (one or two multiplies plus a comparison), and the precision / error analysis is non-trivial.
|
||||
|
||||
## Comparison
|
||||
|
||||
| Approach | Latency (cycles) | Throughput (ops/cycle) | Relative Area | Relative Power | Verification Effort | Notes |
|
||||
|---|---|---|---|---|---|---|
|
||||
| Shift-add MUL / subtractive DIV (iterative, shared) | 32–64 | ~1/32–1/64 | 0.3–0.5× | 0.3–0.5× | Low | Area-optimal, latency-poor |
|
||||
| Radix-4 MUL + subtractive DIV | 32 (MUL) / 32–64 (DIV) | 1/32–1/64 | 0.7–1.0× | 0.7–1.0× | Moderate | Balanced, common in mid-range RV64 |
|
||||
| Radix-4 MUL (pipelined) + SRT DIV | 4–8 (MUL) / 16–32 (DIV) | 1/cycle (MUL) | 1.0–1.4× | 1.0–1.5× | High | Desktop-class; verification-heavy |
|
||||
| Radix-8 MUL (pipelined) + Newton-Raphson DIV | 2–4 (MUL) / 4–8 (DIV) | ≥1/cycle | 1.2–1.6× | 1.2–1.8× | High | High-performance, requires careful rounding-fixup |
|
||||
| Pipelined radix-4 MUL + iterative DIV (split) | 4–8 (MUL) / 32–64 (DIV) | 1/cycle (MUL) | 0.9–1.2× | 0.9–1.2× | Moderate | Workload-asymmetric; RECOMMENDATION-CANDIDATE |
|
||||
| 2-bit-per-cycle naïve DIV | 16–32 | 1/16–1/32 | 0.9–1.1× | 0.9–1.1× | Moderate | Compromise; no SRT selection table |
|
||||
|
||||
**Numbers are illustrative engineering estimates, not benchmark data.** The actual ratios depend on the target process node, synthesis library, and clock target, none of which are known for XH-1.
|
||||
|
||||
## Advantages
|
||||
|
||||
The leading candidates each have a distinct strength profile:
|
||||
|
||||
- **Iterative shared unit:** Smallest per-core area and lowest power, multiplying directly into 128-core replication savings.
|
||||
- **Pipelined radix-4 MUL + iterative DIV (split):** Best workload asymmetry match; MUL throughput is high (common case), DIV cost is contained.
|
||||
- **Pipelined radix-4 MUL + SRT DIV:** Highest predictable throughput on both operations; minimal front-end exposure.
|
||||
- **Newton-Raphson DIV on top of fast MUL:** Reuses the multiplier's silicon; area-efficient if the multiplier is already large.
|
||||
|
||||
## Disadvantages
|
||||
|
||||
- **Iterative shared unit:** Latency directly stalls the pipeline in an in-order core. DIV latencies of 30–60 cycles are unacceptable for tight feedback loops.
|
||||
- **Pipelined radix-4 MUL + iterative DIV:** DIV latency is still long. The dual-unit bookkeeping (two reservation stations, two issue paths) adds microarchitectural complexity.
|
||||
- **Pipelined radix-4 MUL + SRT DIV:** Highest area and verification cost. SRT selection-table corner cases are a known source of post-silicon bugs. The verification cost is multiplied 128× in replication effort.
|
||||
- **Newton-Raphson DIV:** Rounding-toward-zero correction is subtle; commercial implementations have shipped with latent bugs in this path. Not recommended without a clean spec for the correction step.
|
||||
|
||||
## XH-1 Considerations
|
||||
|
||||
**PROPOSAL (pending XH-1 baseline):** The XH-1 has not yet established a microarchitecture baseline. Several XH-1-specific factors bear on the MUL/DIV choice:
|
||||
|
||||
1. **128-core replication amplifies per-core area and verification cost.** A 1.5× area MUL/DIV unit is 1.5× across all 128 cores, not just one. The die-area cost is severe.
|
||||
2. **Per-core issue width and in-order/OoO status are not established.** A wide-issue OoO core tolerates long-latency iterative dividers via scheduling; a narrow in-order core does not.
|
||||
3. **128 cores imply a tile- or mesh-style interconnect.** Long-latency MUL/DIV operations that block a single core are local to that core; the rest of the machine is unaffected. This *favors* simpler per-core MUL/DIV implementations.
|
||||
4. **Verification cost is replicated.** High-effort designs (SRT, Newton-Raphson fixup) need 128× the formal and random-verification runs.
|
||||
|
||||
**OPEN QUESTION:** What is the target workload mix for XH-1? Cryptographic, scientific, server, or general-purpose? Each has very different MUL/DIV demands.
|
||||
|
||||
## 128-Core Scalability
|
||||
|
||||
The MUL/DIV unit is per-core and not shared across cores (in the absence of a globally-shared execution unit, which is unusual for a 128-core design). Scalability considerations:
|
||||
|
||||
- **No coherence problem.** MUL/DIV operates on register values; the result is written to the integer register file, which is private to each core.
|
||||
- **No cross-core contention.** Unlike shared caches or interconnects, the MUL/DIV unit does not introduce hot-spot behavior.
|
||||
- **Verification parallelism.** A 128-core machine can be partitioned for verification; the MUL/DIV unit is verified once, but its integration with the core's pipeline must be verified 128×, unless formal methods prove equivalence.
|
||||
- **Area-budget pressure.** If the per-core area budget is tight, the MUL/DIV unit competes with the L1 cache, branch predictor, and ROB for die area. Iterative implementations (subtractive DIV, iterative MUL) are most competitive in this regime.
|
||||
- **Power-grid and clock distribution.** A replicated fast pipelined multiplier across 128 cores is a substantial clock-tree and switching load. Power delivery and clock skew must be analyzed for the replicated load.
|
||||
|
||||
**PROPOSAL:** Favor an implementation whose area × power product is small, since both costs multiply by 128.
|
||||
|
||||
## Performance Considerations
|
||||
|
||||
MUL/DIV performance interacts with the rest of the core in two ways:
|
||||
|
||||
1. **Latency exposure.** In an in-order core, MUL latency stalls the front-end. In an OoO core, MUL latency is amortized as long as other independent instructions exist.
|
||||
2. **Throughput-bound on dependent chains.** Tight loops of dependent MUL/DIV operations (e.g., big-integer arithmetic, hash functions, matrix multiplies) are throughput-bound, not latency-bound. For these, pipelined MUL at 1/cycle is essential.
|
||||
|
||||
**INSUFFICIENT EVIDENCE:** Without a target workload, no quantitative performance claim can be made.
|
||||
|
||||
## Area Considerations
|
||||
|
||||
Indicative relative area estimates (engineering estimates, not benchmark data):
|
||||
|
||||
| Unit | Relative Area (per core) |
|
||||
|---|---|
|
||||
| Iterative MUL + iterative DIV (shared) | ~0.3–0.5× of "full" unit |
|
||||
| Radix-4 pipelined MUL + iterative DIV (split) | ~0.7–0.9× |
|
||||
| Radix-4 pipelined MUL + radix-4 SRT DIV | ~1.0–1.2× |
|
||||
| Radix-8 pipelined MUL + Newton-Raphson DIV | ~1.2–1.5× |
|
||||
|
||||
**Per 128 cores, a "full" 1.0× unit is 128 unit-areas.** A 0.3× unit is 38 unit-areas — a significant die-area delta. The XH-1 per-core area budget must be defined before this can be quantified.
|
||||
|
||||
## Power and Energy Considerations
|
||||
|
||||
Power and energy per operation, ranked from most to least efficient:
|
||||
|
||||
- **Iterative:** Low per-cycle power, but high per-operation energy × time product (many cycles × dynamic power).
|
||||
- **Pipelined MUL + iterative DIV:** Low per-op MUL energy, high per-op DIV energy.
|
||||
- **Pipelined MUL + SRT DIV:** High per-cycle power, but low per-op energy.
|
||||
- **Pipelined MUL + Newton-Raphson DIV:** Per-op energy is a small multiple of MUL energy; competitive.
|
||||
|
||||
**PROPOSAL:** Energy-per-operation, not energy-per-cycle, is the more relevant metric for a workload-bound design. A pipelined radix-4 MUL with an iterative DIV is a reasonable energy-vs-area compromise.
|
||||
|
||||
## Implementation Considerations
|
||||
|
||||
- **Wallace/Dadda tree synthesis risk.** Compressor trees are notoriously hard to place-and-route at high frequency on modern nodes. A poorly placed adder tree can cost 100+ ps of unexpected delay. This favors iterative or low-radix designs for first-pass XH-1 silicon.
|
||||
- **PPG and Booth recoder verification.** Radix-4 PPGs are easy to verify by simulation; radix-8+ PPGs require more corner cases.
|
||||
- **SRT selection table correctness.** Selection tables must be proven correct for all operand values. This typically requires exhaustive enumeration for reduced operand widths and extensive directed testing for full width.
|
||||
- **RISC-V division-by-zero convention.** `-1` for `DIV`/`DIVU`, `dividend` for `REM`/`REMU`. Must be implemented explicitly; the design must not return a trap.
|
||||
- **Pipeline interlocks.** If MUL is pipelined, the in-order scoreboard or the OoO wakeup logic must handle in-flight MULs; this is a per-instruction-record cost.
|
||||
- **Physical design replication.** A 128-core replicated design benefits from a small, regular MUL/DIV unit. Highly irregular SRT selection logic and complex PPGs resist clean replication.
|
||||
|
||||
## Verification Considerations
|
||||
|
||||
Verification cost is the dominant non-silicon cost in modern CPU design. The MUL/DIV unit is a known hot-spot for bugs:
|
||||
|
||||
- **Radix-2/4 shift-add and subtractive:** Lowest. Exhaustive simulation at reduced widths is tractable.
|
||||
- **Radix-4 Booth + iterative DIV:** Moderate. Corner cases around `MULHSU` and signed division overflow.
|
||||
- **SRT:** High. SRT selection-table bugs are famous in industry (e.g., the Pentium FDIV bug, though that was a different radix); verification typically requires formal proofs of the selection function.
|
||||
- **Newton-Raphson:** High. Rounding-toward-zero correctness across all 64-bit input pairs requires a careful proof of the final fixup step.
|
||||
- **Replication factor:** A bug in the MUL/DIV unit is potentially a bug in all 128 cores. Bugs that manifest only at certain operand combinations are particularly dangerous.
|
||||
|
||||
**RECOMMENDATION:** Verification cost should be a first-class input to the MUL/DIV design decision. A simpler, more easily verified design is preferable to a faster, harder-to-verify one, unless the performance delta is necessary for the XH-1 target workload.
|
||||
|
||||
## Software Considerations
|
||||
|
||||
- **Compiler expectations.** GCC and LLVM default to MUL/DIV idioms for multiplication by constants when the multiplier is not auto-converted. The cost of a long-latency MUL or DIV may push the compiler toward suboptimal code sequences.
|
||||
- **Library code.** libgcc, compiler-rt, and the C library contain hand-written multiplication and division routines (e.g., for `__divti3`, `__udivti3`). These may use the MUL/DIV instructions in tight loops, exposing the unit's throughput directly.
|
||||
- **System software.** Hash functions, AES-GCM polynomial multiplication, modular exponentiation in cryptography, and big-integer arithmetic in language runtimes are all MUL-throughput sensitive.
|
||||
- **128-core interaction.** The system stack (kernel, hypervisor) will run on some subset of the 128 cores. There is no asymmetric design implication: every core is equal.
|
||||
|
||||
**INSUFFICIENT EVIDENCE:** XH-1's target software ecosystem is unknown. Bare-metal, Linux, or richer OS support each imply different MUL/DIV usage profiles.
|
||||
|
||||
## Recommendation
|
||||
|
||||
**No final recommendation is made at this time.** The MUL/DIV design is contingent on three decisions that are not yet established in the XH-1 repository:
|
||||
|
||||
1. The core microarchitecture (in-order vs. OoO, issue width, pipeline depth).
|
||||
2. The per-core area and power budget.
|
||||
3. The target workload mix and software ecosystem.
|
||||
|
||||
**Conditional recommendations** (to be revisited once the above are known):
|
||||
|
||||
- **If XH-1 targets an in-order core with tight area budget:** A single shared iterative shift-add MUL and subtractive DIV unit is the smallest, easiest-to-verify choice. MUL/DIV latency will be high, but the area × 128 replication argument is decisive.
|
||||
- **If XH-1 targets a wide-issue OoO core with relaxed area budget:** A pipelined radix-4 Booth multiplier (4–8 stage pipeline) paired with a radix-4 SRT or 2-bit-per-cycle naïve divider is the natural choice. MUL throughput at 1/cycle; DIV throughput at 1/16 to 1/32 cycle.
|
||||
- **If the per-core area budget is moderate and the workload is balanced:** A pipelined radix-4 MUL with an iterative 32-cycle subtractive DIV (the "split" approach) is a strong compromise.
|
||||
|
||||
**Cross-cutting recommendation:** Defer any design that depends on a non-trivial final correction step (SRT error correction, Newton-Raphson rounding fixup) until those algorithms have been formally specified and proven correct on reduced widths.
|
||||
|
||||
## Confidence
|
||||
|
||||
| Topic | Confidence |
|
||||
|---|---|
|
||||
| Taxonomy of MUL/DIV approaches | High (canonical, well-known) |
|
||||
| Relative area and power rankings | Medium (qualitative, no synthesis data) |
|
||||
| Per-cycle latency numbers | Medium (typical, but node- and target-frequency-dependent) |
|
||||
| Suitability for XH-1 specifically | **Low** (XH-1 microarchitecture not yet established) |
|
||||
| Specific recommendation | **Insufficient evidence** to commit |
|
||||
|
||||
## Open Questions
|
||||
|
||||
1. Is the XH-1 core in-order or out-of-order? Issue width? Pipeline depth? Target frequency?
|
||||
2. What is the per-core area budget, in absolute terms or as a fraction of die area?
|
||||
3. What is the target workload mix (cryptographic, HPC, server, general-purpose, embedded)?
|
||||
4. Will XH-1 implement the "B" (bitmanip) extension, which adds many pseudo-multiplication operations? If so, the MUL/DIV unit may be partially reused.
|
||||
5. Is the 128-core fabric a true SMP (Linux runs on all 128 cores), or is it a heterogeneous mix (e.g., application cores + management cores)? This affects MUL/DIV uniformity requirements.
|
||||
6. What is the target process node? Node geometry strongly affects the area and timing trade-offs of Booth vs. array multipliers.
|
||||
7. Will the MUL/DIV unit be hardened as a macro and replicated 128 times, or will it be synthesized per-core? Hardening is favored for area/power; synthesis-per-core favors design flexibility.
|
||||
8. Is there an asymmetric verification budget for the MUL/DIV unit, or is verification cost unconstrained?
|
||||
9. Does XH-1 need a 64×64→128-bit MUL for the upper-multiply operations to be single-cycle, or is multi-cycle acceptable?
|
||||
10. What is the expected frequency of `DIV`/`REM` instructions in the target workload? Server workloads may have many fewer than scientific workloads.
|
||||
|
||||
## Sources
|
||||
|
||||
- **INSUFFICIENT EVIDENCE.** No XH-1-internal sources exist for this topic. The repository document `research/03-core-design/mul-div-unit.md` is itself a stub ("SOON"), and all sibling documents in `research/03-core-design/` are stubs. No prior architectural decisions are recorded.
|
||||
- No external citations, papers, benchmarks, or measurements have been invented for this document. All comparative claims are framed as engineering estimates or as canonical background knowledge in the field of computer arithmetic.
|
||||
+323
File diff suppressed because one or more lines are too long
+367
@@ -0,0 +1,367 @@
|
||||
# Multiply/Divide Unit Design for the XH-1 Processor
|
||||
|
||||
## Status
|
||||
|
||||
**Stub document — research in progress.** No architectural decisions have been finalized. The XH-1 repository contains no prior decisions on the multiply/divide (MUL/DIV) unit. All content below is a structured engineering analysis of design options, framed as proposals and open questions. Quantitative comparisons are presented only as qualitative relative magnitudes, never as benchmark figures.
|
||||
|
||||
## Abstract
|
||||
|
||||
This document investigates the design of the integer multiply and divide unit (MUL/DIV) for the XH-1, a custom 128-core RISC-V processor. The MUL/DIV unit is a latency-critical but throughput-amortizable functional unit. Its design has outsized impact on the pipeline depth, area per core, and verification effort, all of which compound across 128 replicated cores. We survey the principal implementation strategies (iterative shift-and-add, radix-4/8 Booth, array multipliers, Goldschmidt dividers, Newton-Raphson dividers, subtractive dividers, and SRT dividers) and analyze their trade-offs with respect to latency, throughput, area, power, and verification complexity. The document is intentionally non-prescriptive: the XH-1 has not yet established a baseline for core microarchitecture (in-order vs. out-of-order, pipeline depth), so any recommendations are conditional on those upstream decisions.
|
||||
|
||||
## Research Question
|
||||
|
||||
> What is the appropriate microarchitecture for the integer MUL/DIV unit in the XH-1 core, given that the unit will be replicated 128 times, must support the RV64M extension, and must satisfy XH-1's target latency, area, and power budget — the latter two of which are not yet defined?
|
||||
|
||||
Specifically:
|
||||
|
||||
- How should the unit be pipelined relative to the rest of the core?
|
||||
- Should it share silicon with adjacent datapath structures (ALU, register file read/write ports, FP multiplier) or remain isolated?
|
||||
- How should division throughput be traded against area and verification cost?
|
||||
- What is the impact of the choice on multi-core contention for shared execution resources, if any?
|
||||
|
||||
## Background
|
||||
|
||||
The RISC-V "M" standard extension defines four signed/unsigned variants of multiply and two signed/unsigned variants of divide plus remainder:
|
||||
|
||||
- `MUL` (lower 64 bits of signed × signed)
|
||||
- `MULH`, `MULHU`, `MULHSU` (upper 64 bits, various sign combinations)
|
||||
- `DIV`, `DIVU` (signed/unsigned quotient)
|
||||
- `REM`, `REMU` (signed/unsigned remainder)
|
||||
|
||||
Critical properties of this ISA slice that drive microarchitecture:
|
||||
|
||||
- **Wide-operand upper-multiply (UMUL):** Producing the upper 64 bits of a 64×64→128-bit product is *not* twice the area of a lower-only 64×64→64 multiplier when the multiplier is implemented as a single full-width partial-product / carry-save tree. The 64×64→128 result is generated by the same array; the carry-save adder tree is slightly deeper and wider, and a final carry-propagate adder is needed to collapse the upper half. Standard practice in open RISC-V cores (e.g., Rocket, BOOM, XiangShan) is to share one partial-product array between `MUL`, `MULH*`, and the FP multiplier, with operand- and result-muxing. The "upper-only" mode is therefore a small incremental cost over the lower-only mode, not a doubling.
|
||||
- **Division latencies:** A non-trivial iterative divider for 64-bit operands requires 32 to 64 reduction steps, dominating pipeline depth if fully combinational.
|
||||
- **Signed semantics:** `MULHSU` is conventionally implemented by sign- or zero-extending the signed operand to a full-width signed/unsigned operand and reusing the unsigned 64×64→128 array; the verification burden is therefore small once the unsigned array is correct. By contrast, signed division has two non-trivial edge cases: division by zero (RISC-V returns `-1` for `DIV`/`DIVU`, and the dividend for `REM`/`REMU`) and the signed overflow case `INT64_MIN / -1`, where the mathematical quotient is `INT64_MIN` itself. Hardware that does not detect this case explicitly will return the wrong result.
|
||||
- **Throughput vs. latency decoupling:** RISC-V MUL/DIV results are written to the integer register file, so latency can be hidden by OoO scheduling — but in an in-order core, latency is directly exposed to the front-end unless explicit forwarding is provided.
|
||||
|
||||
The XH-1 is a 128-core machine. Decisions in the MUL/DIV unit replicate 128×, so per-core area and verification effort dominate. There is no public XH-1 microarchitecture document in the repository that would constrain the MUL/DIV design (e.g., pipeline depth, issue width, in-order/OoO).
|
||||
|
||||
**INSUFFICIENT EVIDENCE:** Pipeline depth, issue width, in-order vs. OoO status, target frequency, and target process node for XH-1 are all unknown.
|
||||
|
||||
## Existing Approaches
|
||||
|
||||
The following approaches are well-established in the literature and commercial designs. References: standard computer-arithmetic texts (e.g., Parhami, *Computer Arithmetic*; Ercegovac and Lang, *Digital Arithmetic*; Hennessy and Patterson, *Computer Architecture*) cover these techniques in canonical form. No XH-1-internal prior art exists.
|
||||
|
||||
### 1. Shift-and-Add Multiplier (Iterative)
|
||||
|
||||
- **Mechanism:** A 64-bit multiplier is implemented as a state machine that processes one partial-product bit per cycle against a 128-bit accumulator (or a 129-bit accumulator with a sign-preconditioned variant). Signed operands are typically sign-extended, not handled by halving the iteration count.
|
||||
- **Latency:** Approximately 64 cycles for a 64-bit operand, plus a small constant for sign/result correction.
|
||||
- **Throughput:** One multiply per ~64 cycles.
|
||||
- **Area:** Very small. Roughly one wide adder + one shifter + one accumulator register.
|
||||
- **Power:** Low. Minimal clocked area per cycle.
|
||||
- **Verification:** Low complexity. Straightforward to model and exhaustively test at small operand widths.
|
||||
- **Use case:** Embedded in-order cores where MUL/DIV are infrequent and latency-tolerant. Generally unsuitable for high-performance 64-bit cores.
|
||||
|
||||
### 2. Radix-4 Booth-Encoded Multiplier (Array or Iterative)
|
||||
|
||||
- **Mechanism:** Booth recoding reduces partial products to ceil(n/2). Over a 64-bit operand, the standard radix-4 Booth encoding produces 32 partial-product rows, and an *iterative* radix-4 multiplier performs the additions in roughly 16 cycles (two operand bits consumed per iteration against a running accumulator). Implemented as a single combinational Wallace/Dadda tree, the same 32 partial-product rows are summed in one cycle, with the tree depth determining the achievable clock period.
|
||||
- **Latency:** ~16 cycles iterative, or one combinational tree of ~8–16 levels of CSA + a final CPA, optionally broken into pipeline registers.
|
||||
- **Throughput:** One multiply per 16 cycles iterative; 1/cycle if the array is fully pipelined.
|
||||
- **Area:** Moderate. Recoder + partial-product array + carry-save adder tree + final CPA.
|
||||
- **Power:** Moderate to high. The Wallace/Dadda tree toggles aggressively, and clock-tree load on a replicated array is non-trivial.
|
||||
- **Verification:** Moderate. The corner cases that matter are the signed overflow cases (not the `MULHSU` path, which is essentially free given a correct unsigned array).
|
||||
- **Use case:** Mid- and high-performance cores; the default choice for most modern desktop-class RV64 cores.
|
||||
|
||||
### 3. Radix-8 or Radix-16 Booth Multiplier
|
||||
|
||||
- **Mechanism:** Higher-radix Booth recoding further reduces partial-product rows (radix-8 reduces by 3×, radix-16 by 4× relative to radix-2). The cost is in the partial-product generator (PPG): radix-8 requires 3× multiples; radix-16 requires 3×, 5×, 7× multiples (or a 3×/5× combination plus a higher-order term) and is therefore considerably more complex than radix-4.
|
||||
- **Latency:** The fully pipelined radix-4 array can already achieve 1/cycle throughput; higher radices reduce the *depth* of the adder tree (fewer partial-product rows to sum) and therefore either shorten the critical path or allow fewer pipeline stages. The throughput is not increased beyond 1/cycle unless the array is duplicated.
|
||||
- **Throughput:** 1/cycle for a single pipelined array; not inherently higher than radix-4.
|
||||
- **Area:** Larger PPG; smaller (shallower) adder tree. Net area is roughly comparable to radix-4.
|
||||
- **Power:** Mixed. Fewer adder levels, but more complex PPG.
|
||||
- **Verification:** Higher. Radix-8+ PPGs have more corner cases and the recoding is harder to prove correct.
|
||||
- **Use case:** High-frequency designs where the critical path is in the adder tree, not the PPG.
|
||||
|
||||
### 4. Iterative Subtractive Divider (Non-Restoring / Restoring)
|
||||
|
||||
- **Mechanism:** Standard shift-subtract over the operand width. Produces quotient (and optionally remainder) one bit per cycle.
|
||||
- **Latency:** Roughly 64 cycles for a 64-bit operand (plus a small constant for sign correction).
|
||||
- **Throughput:** One divide per ~64 cycles.
|
||||
- **Area:** Small. A 64- or 65-bit ALU with quotient/remainder registers.
|
||||
- **Power:** Low.
|
||||
- **Verification:** Moderate. The division-by-zero convention, the `INT64_MIN / -1` overflow case, and the `REM`/`REMU` dividend-return case must all be implemented and tested explicitly. The signed-dividend path is the principal source of bugs.
|
||||
- **Use case:** In-order cores with low expected DIV frequency.
|
||||
|
||||
### 5. Radix-4 SRT Divider
|
||||
|
||||
- **Mechanism:** Higher-radix SRT produces 2 (radix-4) or more quotient digits per iteration by selecting one of several shifted multiples of the divisor from a selection table indexed by a truncated partial remainder.
|
||||
- **Latency:** ~16 to 32 cycles for a 64-bit operand.
|
||||
- **Throughput:** One divide per 16–32 cycles.
|
||||
- **Area:** Substantially larger than subtractive. Requires a redundant (carry-save) quotient representation, a quotient-digit selection table, and partial-quotient error-correction logic.
|
||||
- **Power:** Higher.
|
||||
- **Verification:** High. SRT selection tables have notoriously tricky corner cases, and partial-quotient error correction has been the source of silicon bugs in commercial designs (the classic example is the Pentium FDIV bug, which was a radix-4 SRT selection-table defect — the same general technique, not a fundamentally different radix). Verification typically requires formal proofs of the selection function over reduced operand widths plus extensive directed testing.
|
||||
- **Use case:** High-frequency OoO cores needing higher DIV throughput.
|
||||
|
||||
### 6. Goldschmidt Divider (Reciprocal Multiplication)
|
||||
|
||||
- **Mechanism:** Iteratively refine an initial reciprocal estimate, then multiply by the dividend. Each iteration multiplies both the partial remainder and the partial quotient by a correction factor derived from a short reciprocal estimate. Distinct from Newton-Raphson (see §7).
|
||||
- **Latency:** A few multiply iterations. Amortized low latency if the multiplier is fast.
|
||||
- **Throughput:** Bounded by the multiplier.
|
||||
- **Area:** Reuses the multiplier. Adds a small amount of control logic and lookup-table ROM for the initial estimate.
|
||||
- **Power:** Low marginal cost on top of the multiplier.
|
||||
- **Verification:** High. The final iteration must converge to enough bits of precision to guarantee correct rounding toward zero for all 64-bit inputs, which typically requires a final exact-fixup step. The accuracy analysis is non-trivial and historically a bug source.
|
||||
- **Use case:** Designs that already have a fast pipelined multiplier and want a small, fast divider.
|
||||
|
||||
### 7. Newton-Raphson Divider (Reciprocal Multiplication)
|
||||
|
||||
- **Mechanism:** Iteratively refine an initial estimate of the divisor's reciprocal using a Newton-Raphson step, then multiply the dividend by the refined reciprocal. Each iteration squares the error, so convergence is quadratic. Distinct from Goldschmidt, which uses a multiplicative correction on both the partial remainder and the partial quotient simultaneously; the two algorithms have different error dynamics and different fixup requirements.
|
||||
- **Latency:** A few iterations of multiply-add. Typically fewer iterations than Goldschmidt to reach a given precision, but each iteration is a full multiply.
|
||||
- **Throughput:** Bounded by the multiplier.
|
||||
- **Area:** Reuses the multiplier. Adds lookup-table ROM for the initial estimate and modest control logic.
|
||||
- **Power:** Low marginal cost on top of the multiplier.
|
||||
- **Verification:** High. The final-step rounding-toward-zero correctness across all 64-bit input pairs requires a careful proof of the fixup step.
|
||||
- **Use case:** Designs with a fast pipelined multiplier that want a low-latency divider.
|
||||
|
||||
### 8. Shared MUL/DIV Iterative Unit (Multiplexed)
|
||||
|
||||
- **Mechanism:** A single iterative datapath handles both MUL and DIV by reconfiguring its datapath between operations.
|
||||
- **Latency:** Same as the underlying iterative unit; the unit cannot serve MUL and DIV concurrently.
|
||||
- **Throughput:** MUL and DIV contend for the same unit.
|
||||
- **Area:** Smallest. The most area-efficient option.
|
||||
- **Power:** Lowest.
|
||||
- **Verification:** Low to moderate. Single datapath, but dual-purpose microcode/sequencer.
|
||||
- **Use case:** Cost-sensitive cores.
|
||||
|
||||
## Alternative Designs
|
||||
|
||||
Beyond the standard taxonomy, several non-obvious variants are worth considering.
|
||||
|
||||
### A. Pipelined Multiplier + Sequential Divider (Split Unit)
|
||||
|
||||
A fully pipelined radix-4 multiplier (e.g., 4-cycle latency) is paired with an independent sequential subtractive divider. The two units have separate reservation/issue logic.
|
||||
|
||||
- **Rationale:** Under many server, desktop, and general-purpose workloads, MUL is more frequent than DIV, but the ratio is workload-dependent and should not be assumed a priori. Giving MUL a fast pipelined datapath and DIV a slow sequential one matches a plausible workload asymmetry.
|
||||
- **Cost:** Two reservation-station or issue-queue entries (relevant if the core is OoO); two small register files for intermediate state.
|
||||
|
||||
### B. Variable-Latency Divider (Speculative / Early-Quit)
|
||||
|
||||
Produce a partial-width approximate quotient, sign-extend, and perform a single correction step on the remaining bits. The early-quit path saves cycles when the divisor has small magnitude.
|
||||
|
||||
- **Risk:** Variable latency complicates the scoreboard / wakeup logic in an OoO core, or stalls the pipeline in an in-order core. This can be partially mitigated by a fixed maximum latency with early completion, at the cost of additional control logic.
|
||||
|
||||
### C. Radix-2 SRT with Carry-Save Quotient + Software/Hardware Fixup
|
||||
|
||||
A radix-2 SRT with a 1-bit-per-cycle selection function and a carry-save quotient register. Faster than naive subtractive, simpler than radix-4 SRT.
|
||||
|
||||
- **Cost:** The quotient-to-2's-complement conversion is on the critical path or requires a final fixup step.
|
||||
|
||||
### D. Pipelined 2-bit-per-cycle "Naïve" Divider
|
||||
|
||||
Iterate two shift-subtract steps per cycle. Roughly halves divider latency relative to radix-1 subtractive without requiring an SRT selection table.
|
||||
|
||||
- **Cost:** Two wide adders and a longer critical path. Area is moderate.
|
||||
|
||||
### E. Iterative Multiplier (Pipelined) + Lookup-Table Reciprocal + Newton-Raphson
|
||||
|
||||
A small reciprocal ROM (e.g., 12-bit seed) followed by a Newton-Raphson refinement step using the iterative multiplier. Yields a divider latency of a few multiplier iterations.
|
||||
|
||||
- **Cost:** The Newton iteration must converge to enough bits of reciprocal precision to guarantee correct rounding toward zero for all 64-bit inputs, which typically requires a final exact-fixup step (one or two multiplies plus a comparison). The precision / error analysis is non-trivial.
|
||||
|
||||
### F. Shared Multi-Cycle Divider Across Cores (Divider Co-Processor)
|
||||
|
||||
A small number of high-throughput dividers (e.g., 4 or 8) placed at fixed points in the fabric and dispatched to by cores via a memory-mapped or message interface. The dividers are not private to any core.
|
||||
|
||||
- **Rationale:** Avoids replicating the divider 128×. Trades single-core latency (now includes a fabric round-trip) for amortized area.
|
||||
- **Cost:** NOC traffic, dispatch latency, contention at the divider, and software-visible ABI changes (or a transparent-but-slow trap path).
|
||||
- **Status:** Unusual but not unprecedented in accelerator-rich many-core designs.
|
||||
|
||||
### G. FP / Integer Multiplier Sharing
|
||||
|
||||
Share the integer multiplier's partial-product array and adder tree with the FP pipeline. Standard in Rocket, BOOM, and most SiFive cores.
|
||||
|
||||
- **Rationale:** Avoids replicating a wide datapath. The FP pipeline also benefits from a fast multiplier.
|
||||
- **Cost:** Cross-unit scheduling and bypassing complexity. The integer and FP pipelines may have different latency targets.
|
||||
|
||||
### H. Latency-Tolerant In-Order MUL/DIV
|
||||
|
||||
Even a multi-cycle iterative MUL/DIV may be tolerable in an in-order core if the result-bus supports forwarding directly from the MUL/DIV output to dependent consumers, bypassing the register file writeback-read path.
|
||||
|
||||
- **Cost:** Forwarding path length and bypass-network complexity scale with MUL/DIV latency.
|
||||
|
||||
### I. Interaction with the "B" (Bitmanip) Extension
|
||||
|
||||
The RISC-V Bitmanip extension introduces MUL/DIV-adjacent operations (e.g., `CLZ`, `CTZ`, `MIN`, `MAX`, bit-extract/deposit, and several pseudo-multiplication idioms such as `RORI` and `SH*ADD`). If the B extension is in scope, the MUL/DIV unit may either be reused for some of these (e.g., via the ALU) or augmented with dedicated bitmanip datapath. The decision is interdependent with the MUL/DIV choice and is flagged as an open question below.
|
||||
|
||||
## Comparison
|
||||
|
||||
The following table presents *qualitative* relative magnitudes only. No node, frequency, or synthesis data is available for XH-1, so quantitative ratios are not asserted.
|
||||
|
||||
| Approach | Latency (cycles) | Throughput (ops/cycle) | Relative Area | Relative Power | Verification Effort |
|
||||
|---|---|---|---|---|---|
|
||||
| Shift-add MUL / subtractive DIV (iterative, shared) | ~64 | ~1/64 | Smallest | Smallest | Low |
|
||||
| Iterative radix-4 MUL + subtractive DIV | ~16 (MUL) / ~64 (DIV) | ~1/16 (MUL) | Small to moderate | Small to moderate | Low to moderate |
|
||||
| Radix-4 array MUL (pipelined) + subtractive DIV (split) | ~4–8 (MUL) / ~64 (DIV) | 1/cycle (MUL) | Moderate | Moderate | Moderate |
|
||||
| Radix-4 array MUL (pipelined) + radix-4 SRT DIV | ~4–8 (MUL) / ~16–32 (DIV) | 1/cycle (MUL) | Large | Large | High |
|
||||
| Radix-8 array MUL (pipelined) + Newton-Raphson or Goldschmidt DIV | ~3–6 (MUL) / ~4–8 (DIV) | 1/cycle (MUL) | Largest | Largest | High |
|
||||
| 2-bit-per-cycle naïve DIV (paired with radix-4 MUL) | ~4–8 (MUL) / ~32 (DIV) | 1/cycle (MUL) | Moderate | Moderate | Moderate |
|
||||
| Pipelined radix-4 MUL + iterative DIV (split) | ~4–8 (MUL) / ~64 (DIV) | 1/cycle (MUL) | Moderate | Moderate | Moderate |
|
||||
|
||||
The relative magnitudes are illustrative and intended only to convey ordering. The actual ratios depend on the target process node, synthesis library, and clock target, none of which are known for XH-1.
|
||||
|
||||
**INSUFFICIENT EVIDENCE:** Quantitative area, power, and energy comparisons cannot be made without a target node, frequency, and synthesis flow.
|
||||
|
||||
## Advantages
|
||||
|
||||
The leading candidates each have a distinct strength profile:
|
||||
|
||||
- **Iterative shared unit:** Smallest per-core area and lowest power, multiplying directly into 128-core replication savings. Easiest to verify.
|
||||
- **Pipelined radix-4 MUL + iterative DIV (split):** Best workload-asymmetry match if MUL is in fact more frequent than DIV in the target workload; MUL throughput is high (common case), DIV cost is contained.
|
||||
- **Pipelined radix-4 MUL + SRT DIV:** Highest predictable throughput on both operations; minimal front-end exposure if the core is in-order.
|
||||
- **Newton-Raphson or Goldschmidt DIV on top of fast MUL:** Reuses the multiplier's silicon; area-efficient if the multiplier is already large.
|
||||
|
||||
## Disadvantages
|
||||
|
||||
- **Iterative shared unit:** Latency directly stalls the pipeline in an in-order core. Long DIV latencies are problematic for tight feedback loops unless forwarding is provided.
|
||||
- **Pipelined radix-4 MUL + iterative DIV:** DIV latency is still long. The dual-unit bookkeeping (two reservation stations, two issue paths) adds microarchitectural complexity.
|
||||
- **Pipelined radix-4 MUL + SRT DIV:** Highest area and verification cost. SRT selection-table corner cases are a known source of post-silicon bugs (the Pentium FDIV bug being a radix-4 SRT defect). The verification cost is multiplied 128× in replication effort.
|
||||
- **Newton-Raphson / Goldschmidt DIV:** Rounding-toward-zero correction is subtle; commercial implementations have shipped with latent bugs in this path. Not recommended without a clean, formally-specified correction step.
|
||||
|
||||
## XH-1 Considerations
|
||||
|
||||
**PROPOSAL (pending XH-1 baseline):** The XH-1 has not yet established a microarchitecture baseline. Several XH-1-specific factors bear on the MUL/DIV choice:
|
||||
|
||||
1. **128-core replication amplifies per-core area and verification cost.** A larger MUL/DIV unit pays the same area cost across all 128 cores, not just one. The die-area cost is severe.
|
||||
2. **Per-core issue width and in-order/OoO status are not established.** A wide-issue OoO core tolerates long-latency iterative dividers via scheduling; a narrow in-order core does not, unless explicit forwarding is provided.
|
||||
3. **128 cores imply a tile- or mesh-style interconnect.** Long-latency MUL/DIV operations that block a single core are local to that core; the rest of the machine is unaffected. This *favors* simpler per-core MUL/DIV implementations, unless a shared divider alternative (§F) is adopted.
|
||||
4. **Verification cost is replicated.** High-effort designs (SRT, Newton-Raphson fixup) need 128× the formal and random-verification runs.
|
||||
5. **Physical-design regularity matters under replication.** A small, regular MUL/DIV unit is easier to harden and replicate 128× than a complex, irregular SRT unit.
|
||||
|
||||
**OPEN QUESTION:** What is the target workload mix for XH-1? Cryptographic, scientific, server, or general-purpose? Each has very different MUL/DIV demands.
|
||||
|
||||
## 128-Core Scalability
|
||||
|
||||
The MUL/DIV unit is per-core and not shared across cores by default. Scalability considerations:
|
||||
|
||||
- **No coherence problem.** MUL/DIV operates on register values; the result is written to the integer register file, which is private to each core.
|
||||
- **No cross-core contention** in the default per-core configuration. Unlike shared caches or interconnects, the MUL/DIV unit does not introduce hot-spot behavior. A shared-divider alternative (§F) changes this analysis.
|
||||
- **Verification parallelism.** A 128-core machine can be partitioned for verification; the MUL/DIV unit is verified once at the unit level, but its integration with the core's pipeline must be verified 128×, unless formal methods prove equivalence across replications.
|
||||
- **Area-budget pressure.** If the per-core area budget is tight, the MUL/DIV unit competes with the L1 cache, branch predictor, and ROB for die area. Iterative implementations (subtractive DIV, iterative MUL) are most competitive in this regime.
|
||||
- **Power-grid and clock distribution.** A replicated fast pipelined multiplier across 128 cores is a substantial clock-tree and switching load. Power delivery (IR drop), clock skew across the replicated load, and dynamic power density must be analyzed for the replicated load.
|
||||
- **Scan and BIST.** A 128× replicated unit implies 128× the scan-chain length (if scan is per-core) or a partitioned BIST architecture. The choice affects DFT area and test time.
|
||||
- **Fault tolerance.** A defect in the MUL/DIV unit is potentially a defect in all 128 cores. This argues for either a hardened, characterized macro or built-in redundancy / sparing, depending on yield targets.
|
||||
- **Timing variation.** Across-die process variation affects 128 replicated units independently. A design that is timing-marginal at one corner may fail at another. Iterative designs are less sensitive to per-unit timing variation than deep-pipelined arrays.
|
||||
|
||||
**PROPOSAL:** Favor an implementation whose area × power product is small, since both costs multiply by 128. Prefer regular, hardenable structures over irregular ones that resist replication.
|
||||
|
||||
## Performance Considerations
|
||||
|
||||
MUL/DIV performance interacts with the rest of the core in two ways:
|
||||
|
||||
1. **Latency exposure.** In an in-order core, MUL latency stalls the front-end unless explicit forwarding is provided. In an OoO core, MUL latency is amortized as long as other independent instructions exist.
|
||||
2. **Throughput-bound on dependent chains.** Tight loops of dependent MUL/DIV operations (e.g., big-integer arithmetic, hash functions, matrix multiplies) are throughput-bound, not latency-bound. For these, pipelined MUL at 1/cycle is essential.
|
||||
|
||||
**INSUFFICIENT EVIDENCE:** Without a target workload, no quantitative performance claim can be made.
|
||||
|
||||
## Area Considerations
|
||||
|
||||
Relative area ordering, qualitative only (no synthesis data available):
|
||||
|
||||
- Iterative MUL + iterative DIV (shared): smallest.
|
||||
- Iterative radix-4 MUL + subtractive DIV: small to moderate.
|
||||
- Pipelined radix-4 MUL + iterative DIV (split): moderate.
|
||||
- Pipelined radix-4 MUL + radix-4 SRT DIV: large.
|
||||
- Pipelined radix-8 MUL + Newton-Raphson / Goldschmidt DIV: largest.
|
||||
|
||||
The XH-1 per-core area budget must be defined before any of the above can be quantified in absolute terms.
|
||||
|
||||
**INSUFFICIENT EVIDENCE:** Absolute area figures require a process node and a synthesis flow.
|
||||
|
||||
## Power and Energy Considerations
|
||||
|
||||
Relative energy-per-operation ordering, qualitative only:
|
||||
|
||||
- **Iterative:** Low per-cycle power, but high per-operation energy × time product (many cycles).
|
||||
- **Pipelined MUL + iterative DIV:** Low per-op MUL energy, high per-op DIV energy.
|
||||
- **Pipelined MUL + SRT DIV:** High per-cycle power, lower per-op energy than iterative.
|
||||
- **Pipelined MUL + Newton-Raphson / Goldschmidt DIV:** Per-op energy is a small multiple of MUL energy; competitive.
|
||||
|
||||
**PROPOSAL:** Energy-per-operation, not energy-per-cycle, is the more relevant metric for a workload-bound design. A pipelined radix-4 MUL with an iterative DIV is a reasonable energy-vs-area compromise, contingent on the workload.
|
||||
|
||||
**INSUFFICIENT EVIDENCE:** Quantitative power and energy figures require a process node, a clock target, and a workload trace.
|
||||
|
||||
## Implementation Considerations
|
||||
|
||||
- **Wallace/Dadda tree synthesis risk.** Compressor trees are notoriously hard to place-and-route at high frequency on modern nodes. Poor placement can introduce unexpected delay. This favors iterative or low-radix designs for first-pass XH-1 silicon. (No specific delay figure is asserted; the qualitative risk is well-known in the literature.)
|
||||
- **PPG and Booth recoder verification.** Radix-4 PPGs are well-understood and tractable to verify; radix-8+ PPGs require more corner cases.
|
||||
- **SRT selection table correctness.** Selection tables must be proven correct for all operand values. This typically requires exhaustive enumeration for reduced operand widths and extensive directed testing for full width.
|
||||
- **RISC-V division-by-zero and overflow conventions.** `-1` for `DIV`/`DIVU` on divide-by-zero; the dividend for `REM`/`REMU` on divide-by-zero; `INT64_MIN` for `INT64_MIN / -1` on signed overflow. These must be implemented explicitly; the design must not raise a trap.
|
||||
- **Pipeline interlocks.** If MUL is pipelined, the in-order scoreboard or the OoO wakeup logic must handle in-flight MULs; this is a per-instruction-record cost.
|
||||
- **Physical design replication.** A 128-core replicated design benefits from a small, regular MUL/DIV unit. Highly irregular SRT selection logic and complex PPGs resist clean replication and are a physical-design risk under replication.
|
||||
|
||||
## Verification Considerations
|
||||
|
||||
Verification cost is the dominant non-silicon cost in modern CPU design. The MUL/DIV unit is a known hot-spot for bugs:
|
||||
|
||||
- **Radix-2/4 shift-add and subtractive:** Lowest. Exhaustive simulation at reduced widths is tractable.
|
||||
- **Radix-4 Booth + iterative DIV:** Moderate. The hard cases are the signed division edge cases (division-by-zero, `INT64_MIN / -1`), not the `MULHSU` path. The unsigned 64×64→128 array, once verified, covers `MUL`, `MULHU`, and `MULH` (with sign handling); `MULHSU` is a sign-extension of the signed operand to a full-width signed/unsigned form and is essentially free given a correct unsigned array.
|
||||
- **SRT:** High. SRT selection-table bugs are famous in industry (e.g., the Pentium FDIV bug was a radix-4 SRT selection-table defect). Verification typically requires formal proofs of the selection function over reduced widths and extensive directed testing.
|
||||
- **Newton-Raphson / Goldschmidt:** High. Rounding-toward-zero correctness across all 64-bit input pairs requires a careful proof of the final fixup step.
|
||||
- **Replication factor:** A bug in the MUL/DIV unit is potentially a bug in all 128 cores. Bugs that manifest only at certain operand combinations are particularly dangerous.
|
||||
|
||||
**Preliminary verification strategy** (to be refined once the design choice is made):
|
||||
|
||||
- **Unit-level:** Exhaustive simulation at reduced operand widths (e.g., 8, 12, 16 bits) for the core datapath. Formal equivalence checking between the RTL and a reference model written in a high-level specification language (e.g., Bluespec, Scala, or a C reference). For SRT or Newton-Raphson, formal proof of the selection function or the convergence step.
|
||||
- **Directed corner-case suite:** Explicit tests for division-by-zero, signed overflow (`INT64_MIN / -1`), the `REM`/`REMU` dividend-return case, `MULHSU` against a cross-checked reference, and the `MUL`-then-truncate boundary.
|
||||
- **Random / constrained-random:** At full width, comparing against a software reference. Coverage targets on the Booth recoder, PPG, and selection table.
|
||||
- **Integration:** Per-core pipeline integration verified 128× unless formal equivalence is established across replications.
|
||||
- **Post-silicon:** Microarchitectural validation suite, focused on MUL/DIV-intensive kernels (big-integer arithmetic, hashes, polynomial multiplications).
|
||||
|
||||
**PROPOSAL:** Verification cost should be a first-class input to the MUL/DIV design decision. A simpler, more easily verified design is preferable to a faster, harder-to-verify one, unless the performance delta is necessary for the XH-1 target workload.
|
||||
|
||||
## Software Considerations
|
||||
|
||||
- **Compiler behavior.** GCC and LLVM routinely use shift-and-add sequences for multiplication by small constants, and they may either emit `MUL` instructions or inline expansions depending on the cost model. The cost of a long-latency MUL or DIV may push the compiler toward suboptimal code sequences, but the magnitude of this effect is workload- and compiler-version-dependent and should not be assumed.
|
||||
- **Library code.** libgcc, compiler-rt, and the C library contain hand-written multiplication and division routines (e.g., for `__divti3`, `__udivti3`). These may use the MUL/DIV instructions in tight loops, exposing the unit's throughput directly.
|
||||
- **System software.** Hash functions, AES-GCM polynomial multiplication, modular exponentiation in cryptography, and big-integer arithmetic in language runtimes are all MUL-throughput sensitive.
|
||||
- **128-core interaction.** The system stack (kernel, hypervisor) will run on some subset of the 128 cores. There is no asymmetric design implication: every core is equal in the default per-core configuration.
|
||||
|
||||
**INSUFFICIENT EVIDENCE:** XH-1's target software ecosystem is unknown. Bare-metal, Linux, or richer OS support each imply different MUL/DIV usage profiles.
|
||||
|
||||
## Recommendation
|
||||
|
||||
**No final recommendation is made at this time.** The MUL/DIV design is contingent on three decisions that are not yet established in the XH-1 repository:
|
||||
|
||||
1. The core microarchitecture (in-order vs. OoO, issue width, pipeline depth).
|
||||
2. The per-core area and power budget.
|
||||
3. The target workload mix and software ecosystem.
|
||||
|
||||
**Conditional recommendations** (to be revisited once the above are known):
|
||||
|
||||
- **If XH-1 targets an in-order core with tight area budget:** A single shared iterative shift-add MUL and subtractive DIV unit is the smallest, easiest-to-verify choice. MUL/DIV latency will be high; whether this is acceptable depends on whether explicit forwarding is provided.
|
||||
- **If XH-1 targets a wide-issue OoO core with relaxed area budget:** A pipelined radix-4 Booth multiplier (4–8 stage pipeline) paired with a radix-4 SRT or 2-bit-per-cycle naïve divider is the natural choice. MUL throughput at 1/cycle; DIV throughput at 1/16 to 1/32 cycle.
|
||||
- **If the per-core area budget is moderate and the workload is balanced:** A pipelined radix-4 MUL with an iterative subtractive DIV (the "split" approach) is a strong compromise.
|
||||
|
||||
**Cross-cutting recommendation:** Defer any design that depends on a non-trivial final correction step (SRT error correction, Newton-Raphson / Goldschmidt rounding fixup) until those algorithms have been formally specified and proven correct on reduced widths.
|
||||
|
||||
## Confidence
|
||||
|
||||
| Topic | Confidence |
|
||||
|---|---|
|
||||
| Taxonomy of MUL/DIV approaches | High (canonical, well-known) |
|
||||
| Relative area, power, and energy rankings | Low to medium (qualitative ordering only; no synthesis data) |
|
||||
| Per-cycle latency numbers | Low to medium (typical, but node- and target-frequency-dependent) |
|
||||
| Suitability for XH-1 specifically | **Low** (XH-1 microarchitecture not yet established) |
|
||||
| Specific recommendation | **Insufficient evidence** to commit |
|
||||
|
||||
## Open Questions
|
||||
|
||||
1. Is the XH-1 core in-order or out-of-order? Issue width? Pipeline depth? Target frequency?
|
||||
2. What is the per-core area budget, in absolute terms or as a fraction of die area?
|
||||
3. What is the target workload mix (cryptographic, HPC, server, general-purpose, embedded)?
|
||||
4. Will XH-1 implement the "B" (bitmanip) extension, which adds many pseudo-multiplication operations? If so, the MUL/DIV unit may be partially reused or augmented.
|
||||
5. Is the 128-core fabric a true SMP (Linux runs on all 128 cores), or is it a heterogeneous mix (e.g., application cores + management cores)? This affects MUL/DIV uniformity requirements.
|
||||
6. What is the target process node? Node geometry strongly affects the area and timing trade-offs of Booth vs. array multipliers.
|
||||
7. Will the MUL/DIV unit be hardened as a macro and replicated 128 times, or will it be synthesized per-core? Hardening is favored for area/power and replication regularity; synthesis-per-core favors design flexibility.
|
||||
8. Is there an asymmetric verification budget for the MUL/DIV unit, or is verification cost unconstrained?
|
||||
9. Does XH-1 need a 64×64→128-bit MUL for the upper-multiply operations to be single-cycle, or is multi-cycle acceptable?
|
||||
10. What is the expected frequency of `DIV`/`REM` instructions in the target workload? Server workloads may have many fewer than scientific workloads.
|
||||
11. Will the MUL/DIV unit be shared with the FP pipeline (as in Rocket / BOOM), or kept private to the integer pipeline?
|
||||
12. Will XH-1 adopt a shared multi-cycle divider across cores (Alternative F), or is a strictly per-core MUL/DIV unit required?
|
||||
|
||||
## Sources
|
||||
|
||||
- Canonical computer-arithmetic references: Parhami, *Computer Arithmetic*; Ercegovac and Lang, *Digital Arithmetic*; Hennessy and Patterson, *Computer Architecture*. These cover the techniques surveyed above in standard form.
|
||||
- Open RISC-V core implementations (Rocket, BOOM, XiangShan, SiFive) provide reference designs for radix-4 array multipliers, iterative and SRT dividers, and FP/integer multiplier sharing. These are cited as implementation exemplars, not as XH-1 references.
|
||||
- **INSUFFICIENT EVIDENCE:** No XH-1-internal sources exist for this topic. The repository document `research/03-core-design/mul-div-unit.md` is itself a stub ("SOON"), and all sibling documents in `research/03-core-design/` are stubs. No prior architectural decisions are recorded.
|
||||
- No external citations, papers, benchmarks, or measurements have been invented for this document. All comparative claims are framed as qualitative engineering estimates or as canonical background knowledge in the field of computer arithmetic.
|
||||
+315
File diff suppressed because one or more lines are too long
+380
@@ -0,0 +1,380 @@
|
||||
# Multiply/Divide Unit Design for the XH-1 Processor
|
||||
|
||||
## Status
|
||||
|
||||
**Stub document — research in progress.** No architectural decisions have been finalized. The XH-1 repository contains no prior decisions on the multiply/divide (MUL/DIV) unit. All content below is a structured engineering analysis of design options, framed as proposals and open questions. Quantitative comparisons are presented only as qualitative relative magnitudes, never as benchmark figures. Throughout this document, claims are tagged as one of:
|
||||
|
||||
- **FACT** — well-established in the cited literature or in the RISC-V ISA specification.
|
||||
- **TYPICAL** — the common case across published designs; implementation-specific values may vary.
|
||||
- **ASSUMPTION** — an explicit premise the analysis depends on; should be revisited.
|
||||
- **PROPOSAL** — a design recommendation conditional on unresolved parameters.
|
||||
- **INSUFFICIENT EVIDENCE** — no defensible claim can be made without additional information.
|
||||
|
||||
## Abstract
|
||||
|
||||
This document investigates the design of the integer multiply and divide unit (MUL/DIV) for the XH-1, a custom 128-core RISC-V processor. The MUL/DIV unit is a latency-critical but throughput-amortizable functional unit. Its design has outsized impact on pipeline depth, area per core, and verification effort, all of which compound across 128 replicated cores. We survey the principal implementation strategies (iterative shift-and-add, radix-4/8 Booth, array multipliers, Goldschmidt dividers, Newton-Raphson dividers, subtractive dividers, and SRT dividers) and analyze their trade-offs with respect to latency, throughput, area, power, and verification complexity. The document is intentionally non-prescriptive: the XH-1 has not yet established a baseline for core microarchitecture (in-order vs. out-of-order, pipeline depth), so any recommendations are conditional on those upstream decisions.
|
||||
|
||||
## Research Question
|
||||
|
||||
> What is the appropriate microarchitecture for the integer MUL/DIV unit in the XH-1 core, given that the unit will be replicated 128 times, must support the RV64M extension, and must satisfy XH-1's target latency, area, and power budget — the latter two of which are not yet defined?
|
||||
|
||||
Specifically:
|
||||
|
||||
- How should the unit be pipelined relative to the rest of the core?
|
||||
- Should it share silicon with adjacent datapath structures (ALU, register file read/write ports, FP multiplier) or remain isolated?
|
||||
- How should division throughput be traded against area and verification cost?
|
||||
- What is the impact of the choice on multi-core contention for shared execution resources, if any?
|
||||
|
||||
## Background
|
||||
|
||||
**FACT.** The RISC-V "M" standard extension defines four signed/unsigned variants of multiply and two signed/unsigned variants of divide plus remainder:
|
||||
|
||||
- `MUL` (lower 64 bits of signed × signed)
|
||||
- `MULH`, `MULHU`, `MULHSU` (upper 64 bits, various sign combinations)
|
||||
- `DIV`, `DIVU` (signed/unsigned quotient)
|
||||
- `REM`, `REMU` (signed/unsigned remainder)
|
||||
|
||||
Critical properties of this ISA slice that drive microarchitecture:
|
||||
|
||||
- **Wide-operand upper-multiply (UMUL).** **FACT.** Producing the upper 64 bits of a 64×64→128-bit product is *not* twice the area of a lower-only 64×64→64 multiplier when the multiplier is implemented as a single full-width partial-product / carry-save tree. The 64×64→128 result is generated by the same array; the carry-save adder tree is slightly deeper and wider, and a final carry-propagate adder is needed to collapse the upper half. **TYPICAL.** Standard practice in open RISC-V cores (e.g., Rocket, BOOM, XiangShan) is to share one partial-product array between `MUL`, `MULH*`, and the FP multiplier, with operand- and result-muxing. The "upper-only" mode is therefore a small incremental cost over the lower-only mode, not a doubling. *Caveat: the actual incremental cost depends on whether the lower 64 bits must also be produced; the "share the array" pattern is the common case in published RV64 implementations (e.g., SiFive cores, BOOM), but specific silicon area deltas are design-dependent and not asserted here.*
|
||||
- **MULH and signed×signed handling.** **FACT.** `MULH` (signed×signed, upper half) is not obtained by simply reusing the unsigned 64×64→128 array with sign-corrected operands. The standard technique is Baugh-Wooley or Modified Booth with explicit sign-bit handling, which modifies the partial-product generation (sign-extension of the most-significant partial products) and the adder tree. The "essentially free" characterization elsewhere in this document is corrected here: while the underlying adder tree is shared, the partial-product array and the final CPA differ for the signed case, and verification must treat `MULH` as a distinct datapath. **`MULHSU` (signed×unsigned, upper half) is closer to "essentially free":** the signed operand can be sign-extended to 65 bits, the unsigned operand zero-extended, and the resulting 65×64 partial-product array produces the correct upper 64 bits with conventional unsigned-array semantics, modulo sign-correction of the most-significant partial products. Verification cost for `MULHSU` is small once the unsigned array is correct, but `MULH` requires its own verification pass.
|
||||
- **Division latencies.** **FACT.** A non-trivial iterative divider for 64-bit operands requires 32 to 64 reduction steps, dominating pipeline depth if fully combinational.
|
||||
- **Signed semantics.** **FACT.** RISC-V specifies the following for division edge cases:
|
||||
- Division by zero: `DIV` and `DIVU` return `-1` (i.e., all bits set); `REM` and `REMU` return the dividend.
|
||||
- Signed overflow: `INT64_MIN / -1` returns `INT64_MIN` (the mathematical quotient); the corresponding `REM` returns 0.
|
||||
- The operation must not raise an exception; the hardware must produce the specified result.
|
||||
- **Throughput vs. latency decoupling.** **FACT.** RISC-V MUL/DIV results are written to the integer register file, so latency can be hidden by OoO scheduling — but in an in-order core, latency is directly exposed to the front-end unless explicit forwarding is provided.
|
||||
|
||||
The XH-1 is a 128-core machine. **FACT.** Decisions in the MUL/DIV unit replicate 128×, so per-core area and verification effort dominate. **INSUFFICIENT EVIDENCE:** There is no public XH-1 microarchitecture document in the repository that would constrain the MUL/DIV design (e.g., pipeline depth, issue width, in-order/OoO).
|
||||
|
||||
**INSUFFICIENT EVIDENCE:** Pipeline depth, issue width, in-order vs. OoO status, target frequency, target process node, per-core area budget, and target workload mix for XH-1 are all unknown.
|
||||
|
||||
## Existing Approaches
|
||||
|
||||
The following approaches are well-established in the literature and commercial designs. **FACT.** Standard computer-arithmetic texts (Parhami, *Computer Arithmetic*; Ercegovac and Lang, *Digital Arithmetic*; Hennessy and Patterson, *Computer Architecture*) cover these techniques in canonical form. No XH-1-internal prior art exists. Per-cycle latency figures given below are **TYPICAL** values for a 64-bit operand at a moderate clock target (e.g., sub-ns in a recent node); specific values are implementation- and node-dependent.
|
||||
|
||||
### 1. Shift-and-Add Multiplier (Iterative)
|
||||
|
||||
- **Mechanism:** **FACT.** A 64-bit multiplier is implemented as a state machine that processes one partial-product bit per cycle against a 128-bit accumulator (or a 129-bit accumulator with a sign-preconditioned variant). Signed operands are sign-extended; iteration count is not halved.
|
||||
- **Latency:** **TYPICAL.** ~64 cycles for a 64-bit operand, plus a small constant for sign/result correction.
|
||||
- **Throughput:** **TYPICAL.** One multiply per ~64 cycles (unit is not pipelined).
|
||||
- **Area:** **TYPICAL.** Very small. Roughly one wide adder + one shifter + one accumulator register.
|
||||
- **Power:** **TYPICAL.** Low. Minimal clocked area per cycle.
|
||||
- **Verification:** **TYPICAL.** Low complexity. Straightforward to model and exhaustively test at small operand widths.
|
||||
- **Use case:** **TYPICAL.** Embedded in-order cores where MUL/DIV are infrequent and latency-tolerant. Generally unsuitable for high-performance 64-bit cores.
|
||||
|
||||
### 2. Radix-4 Booth-Encoded Multiplier (Array or Iterative)
|
||||
|
||||
- **Mechanism:** **FACT.** Booth recoding reduces partial products to ceil(n/2). For a 64-bit operand, standard radix-4 Booth encoding produces 32 partial-product rows. An *iterative* radix-4 multiplier accumulates these rows one or two at a time:
|
||||
- **One row per cycle:** ~32 cycles, one CSA per cycle, smallest iterative area.
|
||||
- **Two rows per cycle:** ~16 cycles, but requires two CSAs in series per cycle (effectively a 2-stage inner pipeline), doubling the per-cycle area. This is the configuration that achieves the "16-cycle" figure sometimes cited; the area cost must be acknowledged.
|
||||
- **FACT.** Implemented as a single combinational Wallace/Dadda tree, the same 32 partial-product rows are summed in one cycle, with the tree depth determining the achievable clock period.
|
||||
- **Latency:** **TYPICAL.** ~32 cycles iterative (one row/cycle) or ~16 cycles iterative (two rows/cycle, ~2× area), or one combinational tree of ~8–16 CSA levels + a final CPA (the level count is design- and library-specific; the range is illustrative, not a tight bound). When pipelined, the array is typically broken into 3–6 stages.
|
||||
- **Throughput:** **TYPICAL.** One multiply per ~32 cycles (one row/cycle iterative), ~16 cycles (two rows/cycle iterative), or 1/cycle if the array is fully pipelined.
|
||||
- **Area:** **TYPICAL.** Moderate. Recoder + partial-product array + carry-save adder tree + final CPA. The "two rows per cycle" iterative variant is roughly comparable in area to a small pipelined array.
|
||||
- **Power:** **TYPICAL.** Moderate to high when pipelined. The Wallace/Dadda tree toggles aggressively, and clock-tree load on a replicated array is non-trivial.
|
||||
- **Verification:** **TYPICAL.** Moderate. The corner cases that matter are the signed-overflow cases in the `MULH` datapath (sign-extended partial products, modified tree inputs) and the `MULHSU` sign-extension path. The unsigned 64×64→128 array, once verified, covers `MUL`, `MULHU`, and most of `MULHSU`; `MULH` is a separate verification pass.
|
||||
- **Use case:** **TYPICAL.** Mid- and high-performance cores; the default choice for most modern desktop-class RV64 cores.
|
||||
|
||||
### 3. Radix-8 or Radix-16 Booth Multiplier
|
||||
|
||||
- **Mechanism:** **FACT.** Higher-radix Booth recoding reduces the number of partial-product rows. For a 64-bit operand:
|
||||
- Radix-4: 32 rows.
|
||||
- Radix-8: ⌈64/3⌉ = 22 rows (overlapping triples), with the PPG required to produce multiples {0, ±1, ±2, ±3, ±4} of the multiplicand. Generating ±3× typically requires a carry-save adder (1× + 2×), so the PPG is substantially more complex than radix-4.
|
||||
- Radix-16: ⌈64/4⌉ = 16 rows, with the PPG required to produce multiples {0, ±1, ±2, ±3, ±4, ±5, ±6, ±7, ±8} (typically 3×, 5×, 7× via combinations of smaller multiples, with an extra high-order term).
|
||||
- **Correction:** The "radix-8 reduces by 3×, radix-16 by 4×" claim sometimes seen in the literature refers to the ratio relative to radix-2 (64 partial products → 22 or 16), not a clean 3× or 4× multiplier. The actual reductions are 64/22 ≈ 2.9× and 64/16 = 4×, respectively.
|
||||
- **Latency:** **TYPICAL.** The fully pipelined radix-4 array can already achieve 1/cycle throughput; higher radices reduce the *depth* of the adder tree (fewer rows to sum) and therefore either shorten the critical path or allow fewer pipeline stages. Throughput is not increased beyond 1/cycle unless the array is duplicated.
|
||||
- **Throughput:** **TYPICAL.** 1/cycle for a single pipelined array; not inherently higher than radix-4.
|
||||
- **Area:** **TYPICAL.** Larger PPG; smaller (shallower) adder tree. Net area is roughly comparable to radix-4 or slightly larger.
|
||||
- **Power:** **TYPICAL.** Mixed. Fewer adder levels, but more complex PPG.
|
||||
- **Verification:** **TYPICAL.** Higher. Radix-8+ PPGs have more corner cases and the recoding is harder to prove correct.
|
||||
- **Use case:** **TYPICAL.** High-frequency designs where the critical path is in the adder tree, not the PPG.
|
||||
|
||||
### 4. Iterative Subtractive Divider (Non-Restoring / Restoring)
|
||||
|
||||
- **Mechanism:** **FACT.** Standard shift-subtract over the operand width. Produces quotient (and optionally remainder) one bit per cycle.
|
||||
- **Latency:** **TYPICAL.** ~64 cycles for a 64-bit operand (plus a small constant for sign correction).
|
||||
- **Throughput:** **TYPICAL.** One divide per ~64 cycles.
|
||||
- **Area:** **TYPICAL.** Small. A 64- or 65-bit ALU with quotient/remainder registers.
|
||||
- **Power:** **TYPICAL.** Low.
|
||||
- **Verification:** **TYPICAL.** Moderate. The division-by-zero convention, the `INT64_MIN / -1` overflow case, and the `REM`/`REMU` dividend-return case must all be implemented and tested explicitly. The signed-dividend path is the principal source of bugs.
|
||||
- **Use case:** **TYPICAL.** In-order cores with low expected DIV frequency.
|
||||
|
||||
### 5. Radix-4 SRT Divider
|
||||
|
||||
- **Mechanism:** **FACT.** Higher-radix SRT produces multiple quotient digits per iteration by selecting one of several shifted multiples of the divisor from a selection table indexed by a truncated partial remainder. A radix-4 SRT produces 2 bits per iteration; the quotient is held in a redundant (carry-save) form and converted to two's-complement on completion, or corrected on the fly.
|
||||
- **Latency:** **TYPICAL.** ~16 to 32 cycles for a 64-bit operand, depending on the average number of iterations required (SRT sometimes requires an extra iteration for the final correction step).
|
||||
- **Throughput:** **TYPICAL.** One divide per 16–32 cycles.
|
||||
- **Area:** **TYPICAL.** Substantially larger than subtractive. Requires a redundant (carry-save) quotient representation, a quotient-digit selection table, and partial-quotient error-correction logic.
|
||||
- **Power:** **TYPICAL.** Higher.
|
||||
- **Verification:** **TYPICAL.** High. SRT selection tables have notoriously tricky corner cases, and partial-quotient error correction has been the source of silicon bugs in commercial designs. **FACT.** The classic example is the Pentium FDIV bug, which was a defect in the radix-4 SRT lookup table of the Pentium's floating-point divider. The lesson is not merely "SRT is hard" but that the interaction between the redundant quotient representation and the selection function produces error patterns that are not obvious from inspection. Verification typically requires formal proofs of the selection function over reduced operand widths plus extensive directed testing. **Caveat:** the Pentium example is from FP division, not integer division, but the underlying technique (radix-4 SRT with redundant quotient) is the same family, and the verification lessons transfer.
|
||||
- **Use case:** **TYPICAL.** High-frequency OoO cores needing higher DIV throughput.
|
||||
|
||||
### 6. Goldschmidt Divider (Reciprocal Multiplication)
|
||||
|
||||
- **Mechanism:** **FACT.** Iteratively refine an initial reciprocal estimate, then multiply by the dividend. Each iteration multiplies both the partial remainder and the partial quotient by a correction factor derived from a short reciprocal estimate. Distinct from Newton-Raphson (see §7).
|
||||
- **Latency:** **TYPICAL.** A few multiply iterations. Amortized low latency if the multiplier is fast.
|
||||
- **Throughput:** **TYPICAL.** Bounded by the multiplier.
|
||||
- **Area:** **TYPICAL.** Reuses the multiplier. Adds a small amount of control logic and lookup-table ROM for the initial estimate.
|
||||
- **Power:** **TYPICAL.** Low marginal cost on top of the multiplier.
|
||||
- **Verification:** **TYPICAL.** High. The final iteration must converge to enough bits of precision to guarantee correct rounding toward zero for all 64-bit inputs, which typically requires a final exact-fixup step. The accuracy analysis is non-trivial and historically a bug source.
|
||||
- **Use case:** **TYPICAL.** Designs that already have a fast pipelined multiplier and want a small, fast divider.
|
||||
|
||||
### 7. Newton-Raphson Divider (Reciprocal Multiplication)
|
||||
|
||||
- **Mechanism:** **FACT.** Iteratively refine an initial estimate of the divisor's reciprocal using a Newton-Raphson step, then multiply the dividend by the refined reciprocal. Each iteration squares the error, so convergence is quadratic. Distinct from Goldschmidt, which uses a multiplicative correction on both the partial remainder and the partial quotient simultaneously; the two algorithms have different error dynamics and different fixup requirements.
|
||||
- **Latency:** **TYPICAL.** A few iterations of multiply-add. Typically fewer iterations than Goldschmidt to reach a given precision, but each iteration is a full multiply.
|
||||
- **Throughput:** **TYPICAL.** Bounded by the multiplier.
|
||||
- **Area:** **TYPICAL.** Reuses the multiplier. Adds lookup-table ROM for the initial estimate and modest control logic.
|
||||
- **Power:** **TYPICAL.** Low marginal cost on top of the multiplier.
|
||||
- **Verification:** **TYPICAL.** High. The final-step rounding-toward-zero correctness across all 64-bit input pairs requires a careful proof of the fixup step.
|
||||
- **Use case:** **TYPICAL.** Designs with a fast pipelined multiplier that want a low-latency divider.
|
||||
|
||||
### 8. Shared MUL/DIV Iterative Unit (Multiplexed)
|
||||
|
||||
- **Mechanism:** **ASSUMPTION.** A single iterative datapath handles both MUL and DIV by reconfiguring its datapath between operations. **TYPICAL.** This pattern is more common in microcoded embedded cores than in modern 64-bit RV64 designs, where MUL and DIV datapaths are structurally different (shift-and-add with accumulator vs. shift-subtract with quotient register) and the area savings from sharing are modest compared to the control complexity of reconfiguration. The characterization in this document is qualified accordingly.
|
||||
- **Latency:** **TYPICAL.** Same as the underlying iterative unit; the unit cannot serve MUL and DIV concurrently.
|
||||
- **Throughput:** **TYPICAL.** MUL and DIV contend for the same unit.
|
||||
- **Area:** **TYPICAL.** Small. **ASSUMPTION.** The "most area-efficient option" claim is conditional on the datapath being genuinely shared rather than microcoded over separate datapaths; for RV64, this is not the default choice in published designs.
|
||||
- **Power:** **TYPICAL.** Low.
|
||||
- **Verification:** **TYPICAL.** Low to moderate if the datapath is genuinely shared; higher if microcode overlays separate datapaths.
|
||||
- **Use case:** **TYPICAL.** Cost-sensitive embedded cores; uncommon in high-performance RV64.
|
||||
|
||||
## Alternative Designs
|
||||
|
||||
Beyond the standard taxonomy, several non-obvious variants are worth considering.
|
||||
|
||||
### A. Pipelined Multiplier + Sequential Divider (Split Unit)
|
||||
|
||||
A fully pipelined radix-4 multiplier (e.g., 4-cycle latency) is paired with an independent sequential subtractive divider. The two units have separate reservation/issue logic.
|
||||
|
||||
- **Rationale:** **ASSUMPTION.** Under many server, desktop, and general-purpose workloads, MUL is more frequent than DIV, but the ratio is workload-dependent and should not be assumed a priori. Giving MUL a fast pipelined datapath and DIV a slow sequential one matches a plausible workload asymmetry.
|
||||
- **Cost:** Two reservation-station or issue-queue entries (relevant if the core is OoO); two small register files for intermediate state.
|
||||
|
||||
### B. Variable-Latency Divider (Speculative / Early-Quit)
|
||||
|
||||
Produce a partial-width approximate quotient, sign-extend, and perform a single correction step on the remaining bits. The early-quit path saves cycles when the divisor has small magnitude.
|
||||
|
||||
- **Risk:** **TYPICAL.** Variable latency complicates the scoreboard / wakeup logic in an OoO core, or stalls the pipeline in an in-order core. This can be partially mitigated by a fixed maximum latency with early completion, at the cost of additional control logic.
|
||||
|
||||
### C. Radix-2 SRT with Carry-Save Quotient + Software/Hardware Fixup
|
||||
|
||||
A radix-2 SRT with a 1-bit-per-cycle selection function and a carry-save quotient register. Faster than naive subtractive, simpler than radix-4 SRT.
|
||||
|
||||
- **Cost:** The quotient-to-2's-complement conversion is on the critical path or requires a final fixup step.
|
||||
|
||||
### D. Pipelined 2-bit-per-cycle "Naïve" Divider
|
||||
|
||||
Iterate two shift-subtract steps per cycle. Roughly halves divider latency relative to radix-1 subtractive without requiring an SRT selection table.
|
||||
|
||||
- **Cost:** Two wide adders and a longer critical path. Area is moderate.
|
||||
|
||||
### E. Iterative Multiplier (Pipelined) + Lookup-Table Reciprocal + Newton-Raphson
|
||||
|
||||
A small reciprocal ROM (e.g., 12-bit seed) followed by a Newton-Raphson refinement step using the iterative multiplier. Yields a divider latency of a few multiplier iterations.
|
||||
|
||||
- **Cost:** **TYPICAL.** The Newton iteration must converge to enough bits of reciprocal precision to guarantee correct rounding toward zero for all 64-bit inputs, which typically requires a final exact-fixup step (one or two multiplies plus a comparison). The precision / error analysis is non-trivial.
|
||||
|
||||
### F. Shared Multi-Cycle Divider Across Cores (Divider Co-Processor)
|
||||
|
||||
A small number of high-throughput dividers (e.g., 4 or 8) placed at fixed points in the fabric and dispatched to by cores via a memory-mapped or message interface. The dividers are not private to any core.
|
||||
|
||||
- **Rationale:** Avoids replicating the divider 128×. Trades single-core latency (now includes a fabric round-trip) for amortized area.
|
||||
- **Cost:** NOC traffic, dispatch latency, contention at the divider, and software-visible ABI changes (or a transparent-but-slow trap path).
|
||||
- **Status:** Unusual but not unprecedented in accelerator-rich many-core designs.
|
||||
|
||||
### G. FP / Integer Multiplier Sharing
|
||||
|
||||
Share the integer multiplier's partial-product array and adder tree with the FP pipeline. **TYPICAL.** This pattern is standard in Rocket, BOOM, and most SiFive cores.
|
||||
|
||||
- **Rationale:** Avoids replicating a wide datapath. The FP pipeline also benefits from a fast multiplier.
|
||||
- **Cost (qualified):** **TYPICAL.** Cross-unit scheduling and bypassing complexity. The integer and FP pipelines may have different latency targets. The sharing requires operand-format conversion (integer operands to FP-like internal format, and vice versa) and FP-specific concerns (rounding mode support, subnormal handling, NaN propagation) are not "free" — they are offloaded to the FP pipeline's existing logic, but the integer side must correctly drive and consume the shared datapath. Verification must cover the combined integer-plus-FP datapath, which is more complex than either alone. The "near-free" characterization sometimes seen in the literature is oversimplified; the cost is real but is often dominated by the FP-pipeline logic that already exists, making the incremental cost on the integer side smaller than the absolute cost of a separate integer multiplier.
|
||||
|
||||
### H. Latency-Tolerant In-Order MUL/DIV
|
||||
|
||||
Even a multi-cycle iterative MUL/DIV may be tolerable in an in-order core if the result-bus supports forwarding directly from the MUL/DIV output to dependent consumers, bypassing the register file writeback-read path.
|
||||
|
||||
- **Cost:** Forwarding path length and bypass-network complexity scale with MUL/DIV latency.
|
||||
|
||||
### I. Interaction with the "B" (Bitmanip) Extension
|
||||
|
||||
The RISC-V Bitmanip extension introduces MUL/DIV-adjacent operations (e.g., `CLZ`, `CTZ`, `MIN`, `MAX`, bit-extract/deposit, and several pseudo-multiplication idioms such as `RORI` and `SH*ADD`). If the B extension is in scope, the MUL/DIV unit may either be reused for some of these (e.g., via the ALU) or augmented with dedicated bitmanip datapath. The decision is interdependent with the MUL/DIV choice and is flagged as an open question below.
|
||||
|
||||
## Comparison
|
||||
|
||||
The following table presents *qualitative* relative magnitudes only. **INSUFFICIENT EVIDENCE:** No node, frequency, or synthesis data is available for XH-1, so quantitative ratios are not asserted. "Latency" is in cycles for back-to-back independent operations on the named unit; "throughput" is sustained operations per cycle for a fully pipelined or iterative unit, respectively.
|
||||
|
||||
| Approach | Latency (cycles) | Throughput (ops/cycle) | Relative Area | Relative Power | Verification Effort |
|
||||
|---|---|---|---|---|---|
|
||||
| Iterative shift-add MUL (one row/cycle) + subtractive DIV (shared or separate) | ~64 (MUL) / ~64 (DIV) | ~1/64 (each) | Smallest | Smallest | Low |
|
||||
| Iterative radix-4 MUL (one row/cycle) + subtractive DIV | ~32 (MUL) / ~64 (DIV) | ~1/32 (MUL) | Small | Small to moderate | Low to moderate |
|
||||
| Pipelined radix-4 array MUL + subtractive DIV (split) | ~4–8 (MUL) / ~64 (DIV) | 1/cycle (MUL) | Moderate | Moderate | Moderate |
|
||||
| Pipelined radix-4 array MUL + radix-4 SRT DIV | ~4–8 (MUL) / ~16–32 (DIV) | 1/cycle (MUL) | Large | Large | High |
|
||||
| Pipelined radix-8 array MUL + Newton-Raphson or Goldschmidt DIV | ~3–6 (MUL) / ~4–8 (DIV) | 1/cycle (MUL) | Largest | Largest | High |
|
||||
| Pipelined radix-4 MUL + 2-bit-per-cycle naïve DIV | ~4–8 (MUL) / ~32 (DIV) | 1/cycle (MUL) | Moderate | Moderate | Moderate |
|
||||
|
||||
**INSUFFICIENT EVIDENCE:** The relative magnitudes are illustrative and intended only to convey ordering. The actual ratios depend on the target process node, synthesis library, and clock target, none of which are known for XH-1. **INSUFFICIENT EVIDENCE:** Quantitative area, power, and energy comparisons cannot be made without a target node, frequency, and synthesis flow.
|
||||
|
||||
## Advantages
|
||||
|
||||
The leading candidates each have a distinct strength profile. The following are conditional on the workload and core microarchitecture, which are not yet established.
|
||||
|
||||
- **Iterative shared unit (Approach 8, with the qualification in §8):** **ASSUMPTION.** Smallest per-core area and lowest power in microcoded embedded cores; for RV64, the area advantage over a split iterative MUL + iterative DIV is modest and the control complexity may offset the savings. Easiest to verify only if the datapath is genuinely shared.
|
||||
- **Pipelined radix-4 MUL + iterative subtractive DIV (split, Alternative A):** **ASSUMPTION.** Matches a plausible workload asymmetry where MUL is more frequent than DIV; MUL throughput is high (common case), DIV cost is contained, and verification is tractable. This is a strong compromise candidate, not a leading candidate by default.
|
||||
- **Pipelined radix-4 MUL + SRT DIV:** **TYPICAL.** Highest predictable throughput on both operations; minimal front-end exposure if the core is in-order. Verification cost is high.
|
||||
- **Newton-Raphson or Goldschmidt DIV on top of fast MUL:** **TYPICAL.** Reuses the multiplier's silicon; area-efficient if the multiplier is already large. Verification cost is high.
|
||||
|
||||
## Disadvantages
|
||||
|
||||
- **Iterative shared unit:** **TYPICAL.** Latency directly stalls the pipeline in an in-order core. Long DIV latencies are problematic for tight feedback loops unless forwarding is provided. For RV64, the structural mismatch between MUL and DIV datapaths limits the achievable area savings.
|
||||
- **Pipelined radix-4 MUL + iterative DIV:** **TYPICAL.** DIV latency is still long. The dual-unit bookkeeping (two reservation stations, two issue paths) adds microarchitectural complexity.
|
||||
- **Pipelined radix-4 MUL + SRT DIV:** **TYPICAL.** Highest area and verification cost. SRT selection-table corner cases are a known source of post-silicon bugs (the Pentium FDIV bug being a radix-4 SRT selection-table defect). The verification cost is replicated across cores.
|
||||
- **Newton-Raphson / Goldschmidt DIV:** **TYPICAL.** Rounding-toward-zero correction is subtle; commercial implementations have shipped with latent bugs in this path. Not recommended without a clean, formally-specified correction step.
|
||||
|
||||
## XH-1 Considerations
|
||||
|
||||
**PROPOSAL (pending XH-1 baseline):** The XH-1 has not yet established a microarchitecture baseline. Several XH-1-specific factors bear on the MUL/DIV choice:
|
||||
|
||||
1. **128-core replication amplifies per-core area and verification cost.** **FACT.** A larger MUL/DIV unit pays the same area cost across all 128 cores, not just one. The die-area cost is severe.
|
||||
2. **Per-core issue width and in-order/OoO status are not established.** **INSUFFICIENT EVIDENCE.** A wide-issue OoO core tolerates long-latency iterative dividers via scheduling; a narrow in-order core does not, unless explicit forwarding is provided.
|
||||
3. **128 cores imply a tile- or mesh-style interconnect.** **ASSUMPTION.** Long-latency MUL/DIV operations that block a single core are local to that core; the rest of the machine is unaffected. This *favors* simpler per-core MUL/DIV implementations, unless a shared divider alternative (Alternative F) is adopted.
|
||||
4. **Verification cost is replicated at the unit level, not the integration level.** **FACT.** A replicated unit is verified once at the unit level (RTL, formal, directed/random). The integration with each core's pipeline is identical across replications and is verified once via the core-level verification environment; running the same integration suite 128× does not add coverage. The 128× replication matters for silicon defect exposure (a bug that escapes verification affects all cores) and for DFT/scan/BIST architecture, not for per-instance verification effort.
|
||||
5. **Physical-design regularity matters under replication.** **TYPICAL.** A small, regular MUL/DIV unit is easier to harden and replicate 128× than a complex, irregular SRT unit.
|
||||
|
||||
**OPEN QUESTION:** What is the target workload mix for XH-1? Cryptographic, scientific, server, or general-purpose? Each has very different MUL/DIV demands.
|
||||
|
||||
## 128-Core Scalability
|
||||
|
||||
The MUL/DIV unit is per-core and not shared across cores by default. Scalability considerations:
|
||||
|
||||
- **No coherence problem.** **FACT.** MUL/DIV operates on register values; the result is written to the integer register file, which is private to each core.
|
||||
- **No cross-core contention** in the default per-core configuration. **FACT.** Unlike shared caches or interconnects, the MUL/DIV unit does not introduce hot-spot behavior. A shared-divider alternative (Alternative F) changes this analysis.
|
||||
- **Verification parallelism.** **FACT (clarified).** A 128-core machine can be partitioned for verification; the MUL/DIV unit is verified once at the unit level. Integration with the pipeline is verified once at the core level (since all cores are identical replications), not 128×. The 128× replication affects silicon defect exposure, not verification run-count.
|
||||
- **Area-budget pressure.** **TYPICAL.** If the per-core area budget is tight, the MUL/DIV unit competes with the L1 cache, branch predictor, and ROB for die area. Iterative implementations (subtractive DIV, iterative MUL) are most competitive in this regime.
|
||||
- **Power-grid and clock distribution.** **TYPICAL.** A replicated fast pipelined multiplier across 128 cores is a substantial clock-tree and switching load. **INSUFFICIENT EVIDENCE:** Power delivery (IR drop), clock skew across the replicated load, and dynamic power density must be analyzed for the replicated load; no quantitative estimates are made here.
|
||||
- **Scan and BIST.** **TYPICAL.** A 128× replicated unit implies 128× the scan-chain length (if scan is per-core) or a partitioned BIST architecture. The choice affects DFT area and test time.
|
||||
- **Fault tolerance.** **TYPICAL.** A defect in the MUL/DIV unit is potentially a defect in all 128 cores. This argues for either a hardened, characterized macro or built-in redundancy / sparing, depending on yield targets.
|
||||
- **Timing variation.** **TYPICAL.** Across-die process variation affects 128 replicated units independently. A design that is timing-marginal at one corner may fail at another. Iterative designs are less sensitive to per-unit timing variation than deep-pipelined arrays.
|
||||
|
||||
**PROPOSAL:** Favor an implementation whose area × power product is small, since both costs multiply by 128. Prefer regular, hardenable structures over irregular ones that resist replication.
|
||||
|
||||
## Performance Considerations
|
||||
|
||||
MUL/DIV performance interacts with the rest of the core in two ways:
|
||||
|
||||
1. **Latency exposure.** **FACT.** In an in-order core, MUL latency stalls the front-end unless explicit forwarding is provided. In an OoO core, MUL latency is amortized as long as other independent instructions exist.
|
||||
2. **Throughput-bound on dependent chains.** **FACT.** Tight loops of dependent MUL/DIV operations (e.g., big-integer arithmetic, hash functions, matrix multiplies) are throughput-bound, not latency-bound. For these, pipelined MUL at 1/cycle is essential.
|
||||
|
||||
**INSUFFICIENT EVIDENCE:** Without a target workload, no quantitative performance claim can be made.
|
||||
|
||||
## Area Considerations
|
||||
|
||||
Relative area ordering, qualitative only (no synthesis data available):
|
||||
|
||||
- Iterative MUL + iterative DIV (shared, with the qualifications in §8): smallest in microcoded embedded cores; for RV64, the advantage over a split iterative design is modest.
|
||||
- Iterative radix-4 MUL (one row/cycle) + subtractive DIV: small.
|
||||
- Pipelined radix-4 MUL + iterative subtractive DIV (split): moderate.
|
||||
- Pipelined radix-4 MUL + 2-bit-per-cycle naïve DIV: moderate.
|
||||
- Pipelined radix-4 MUL + radix-4 SRT DIV: large.
|
||||
- Pipelined radix-8 MUL + Newton-Raphson / Goldschmidt DIV: largest.
|
||||
|
||||
**INSUFFICIENT EVIDENCE:** The XH-1 per-core area budget must be defined before any of the above can be quantified in absolute terms. Absolute area figures require a process node and a synthesis flow.
|
||||
|
||||
## Power and Energy Considerations
|
||||
|
||||
Relative energy-per-operation ordering, qualitative only:
|
||||
|
||||
- **Iterative:** **TYPICAL.** Low per-cycle power, but high per-operation energy × time product (many cycles).
|
||||
- **Pipelined MUL + iterative DIV:** **TYPICAL.** Low per-op MUL energy, high per-op DIV energy.
|
||||
- **Pipelined MUL + SRT DIV:** **TYPICAL.** High per-cycle power, lower per-op energy than iterative.
|
||||
- **Pipelined MUL + Newton-Raphson / Goldschmidt DIV:** **TYPICAL.** Per-op energy is a small multiple of MUL energy; competitive.
|
||||
|
||||
**PROPOSAL:** Energy-per-operation, not energy-per-cycle, is the more relevant metric for a workload-bound design. A pipelined radix-4 MUL with an iterative subtractive DIV is a reasonable energy-vs-area compromise, contingent on the workload.
|
||||
|
||||
**INSUFFICIENT EVIDENCE:** Quantitative power and energy figures require a process node, a clock target, and a workload trace.
|
||||
|
||||
## Implementation Considerations
|
||||
|
||||
- **Wallace/Dadda tree synthesis risk.** **TYPICAL (opinion flagged).** Compressor trees are known to be hard to place-and-route at high frequency on modern nodes; poor placement can introduce unexpected delay. This favors iterative or low-radix designs for first-pass XH-1 silicon. *This is an engineering judgment, not an established fact; specific delay figures depend on the synthesis flow and library, and are not asserted here.*
|
||||
- **PPG and Booth recoder verification.** **TYPICAL.** Radix-4 PPGs are well-understood and tractable to verify; radix-8+ PPGs require more corner cases. The `MULH` signed×signed upper-half path requires explicit verification of sign-extended partial products and is not a free byproduct of the unsigned array.
|
||||
- **SRT selection table correctness.** **TYPICAL.** Selection tables must be proven correct for all operand values. This typically requires exhaustive enumeration for reduced operand widths and extensive directed testing for full width.
|
||||
- **RISC-V division-by-zero and overflow conventions.** **FACT.** `DIV` and `DIVU` return `-1` (all bits set) on divide-by-zero; `REM` and `REMU` return the dividend on divide-by-zero. On signed overflow (`INT64_MIN / -1`), `DIV` returns `INT64_MIN` and `REM` returns 0. These must be implemented explicitly; the design must not raise a trap.
|
||||
- **Pipeline interlocks.** **FACT.** If MUL is pipelined, the in-order scoreboard or the OoO wakeup logic must handle in-flight MULs; this is a per-instruction-record cost.
|
||||
- **Physical design replication.** **TYPICAL.** A 128-core replicated design benefits from a small, regular MUL/DIV unit. Highly irregular SRT selection logic and complex PPGs resist clean replication and are a physical-design risk under replication.
|
||||
|
||||
## Verification Considerations
|
||||
|
||||
Verification cost is a dominant non-silicon cost in modern CPU design. The MUL/DIV unit is a known hot-spot for bugs:
|
||||
|
||||
- **Radix-2/4 shift-add and subtractive:** **TYPICAL.** Lowest. Exhaustive simulation at reduced widths is tractable.
|
||||
- **Radix-4 Booth + iterative DIV:** **TYPICAL.** Moderate. The hard cases are the signed division edge cases (division-by-zero, `INT64_MIN / -1`), the `MULH` signed×signed upper-half path (sign-extended partial products, not a free byproduct of the unsigned array), and the `MULHSU` sign-extension. The unsigned 64×64→128 array, once verified, covers `MUL`, `MULHU`, and the bulk of `MULHSU`; `MULH` is a separate verification pass.
|
||||
- **SRT:** **TYPICAL.** High. SRT selection-table bugs are famous in industry (the Pentium FDIV bug was a radix-4 SRT selection-table defect). Verification typically requires formal proofs of the selection function over reduced widths and extensive directed testing.
|
||||
- **Newton-Raphson / Goldschmidt:** **TYPICAL.** High. Rounding-toward-zero correctness across all 64-bit input pairs requires a careful proof of the final fixup step.
|
||||
- **Replication factor (clarified):** **FACT.** A bug in the MUL/DIV unit that escapes verification is a bug in all 128 cores (silicon defect exposure). However, the verification effort itself is not 128×: the unit is verified once at the unit level, and the core-level integration is verified once (cores are identical replications). The risk is concentrated exposure, not multiplied effort.
|
||||
|
||||
**Preliminary verification strategy** (to be refined once the design choice is made):
|
||||
|
||||
- **Unit-level:** Exhaustive simulation at reduced operand widths (e.g., 8, 12, 16 bits) for the core datapath. Formal equivalence checking between the RTL and a reference model written in a high-level specification language (e.g., Bluespec, Scala, or a C reference). For SRT or Newton-Raphson, formal proof of the selection function or the convergence step.
|
||||
- **Directed corner-case suite:** Explicit tests for division-by-zero (both `DIV`/`DIVU` and `REM`/`REMU` paths), signed overflow (`INT64_MIN / -1`, both quotient and remainder), `MULH` against a cross-checked reference (Baugh-Wooley or equivalent), `MULHSU` against a cross-checked reference, and the `MUL`-then-truncate boundary.
|
||||
- **Random / constrained-random:** At full width, comparing against a software reference. Coverage targets on the Booth recoder, PPG, and selection table.
|
||||
- **Integration:** Per-core pipeline integration verified once at the core level (not 128×), since cores are identical.
|
||||
- **Post-silicon:** Microarchitectural validation suite, focused on MUL/DIV-intensive kernels (big-integer arithmetic, hashes, polynomial multiplications).
|
||||
|
||||
**PROPOSAL:** Verification cost should be a first-class input to the MUL/DIV design decision. A simpler, more easily verified design is preferable to a faster, harder-to-verify one, unless the performance delta is necessary for the XH-1 target workload.
|
||||
|
||||
## Software Considerations
|
||||
|
||||
- **Compiler behavior.** **FACT.** GCC and LLVM routinely use shift-and-add sequences for multiplication by small constants, and they may either emit `MUL` instructions or inline expansions depending on the cost model. The cost of a long-latency MUL or DIV may push the compiler toward suboptimal code sequences, but the magnitude of this effect is workload- and compiler-version-dependent and should not be assumed.
|
||||
- **Library code.** **FACT.** libgcc, compiler-rt, and the C library contain hand-written multiplication and division routines (e.g., for `__divti3`, `__udivti3`). These may use the MUL/DIV instructions in tight loops, exposing the unit's throughput directly.
|
||||
- **System software.** **TYPICAL.** Hash functions, AES-GCM polynomial multiplication, modular exponentiation in cryptography, and big-integer arithmetic in language runtimes are all MUL-throughput sensitive.
|
||||
- **128-core interaction.** **ASSUMPTION.** The system stack (kernel, hypervisor) will run on some subset of the 128 cores. There is no asymmetric design implication: every core is equal in the default per-core configuration.
|
||||
|
||||
**INSUFFICIENT EVIDENCE:** XH-1's target software ecosystem is unknown. Bare-metal, Linux, or richer OS support each imply different MUL/DIV usage profiles.
|
||||
|
||||
## Recommendation
|
||||
|
||||
**No final recommendation is made at this time.** The MUL/DIV design is contingent on three decisions that are not yet established in the XH-1 repository:
|
||||
|
||||
1. The core microarchitecture (in-order vs. OoO, issue width, pipeline depth).
|
||||
2. The per-core area and power budget.
|
||||
3. The target workload mix and software ecosystem.
|
||||
|
||||
**Conditional recommendations** (to be revisited once the above are known):
|
||||
|
||||
- **If XH-1 targets an in-order core with tight area budget:** **PROPOSAL.** A separate iterative shift-add MUL (one row/cycle) and an iterative subtractive DIV, with explicit forwarding from the MUL/DIV output to dependent consumers, is the smallest, easiest-to-verify choice. MUL/DIV latency will be high; whether this is acceptable depends on the forwarding path and the workload. The shared-iterative-unit approach (Approach 8) is *not* recommended for RV64 by default, given the structural mismatch between MUL and DIV datapaths and the modest area advantage over a split iterative design.
|
||||
- **If XH-1 targets a wide-issue OoO core with relaxed area budget:** **PROPOSAL.** A pipelined radix-4 Booth multiplier (4–8 stage pipeline) paired with a 2-bit-per-cycle naïve divider or a radix-4 SRT divider. MUL throughput at 1/cycle; DIV throughput at 1/16 to 1/32 cycle. **Reconciliation note:** the cross-cutting recommendation below defers SRT until formally proven correct; if SRT cannot be formally verified on reduced widths within the project timeline, the 2-bit-per-cycle naïve divider is the preferred DIV companion. The SRT recommendation is conditional on the verification investment being made.
|
||||
- **If the per-core area budget is moderate and the workload is balanced:** **PROPOSAL.** A pipelined radix-4 MUL with an iterative subtractive DIV (the "split" approach, Alternative A) is a strong compromise.
|
||||
|
||||
**Cross-cutting recommendation:** Defer any design that depends on a non-trivial final correction step (SRT error correction, Newton-Raphson / Goldschmidt rounding fixup) until those algorithms have been formally specified and proven correct on reduced widths.
|
||||
|
||||
## Confidence
|
||||
|
||||
| Topic | Confidence |
|
||||
|---|---|
|
||||
| Taxonomy of MUL/DIV approaches | High (canonical, well-known) |
|
||||
| Relative area, power, and energy rankings | Low to medium (qualitative ordering only; no synthesis data) |
|
||||
| Per-cycle latency numbers | Low to medium (typical, but node- and target-frequency-dependent) |
|
||||
| Suitability for XH-1 specifically | **Low** (XH-1 microarchitecture not yet established) |
|
||||
| Specific recommendation | **Insufficient evidence** to commit |
|
||||
|
||||
## Open Questions
|
||||
|
||||
1. Is the XH-1 core in-order or out-of-order? Issue width? Pipeline depth? Target frequency?
|
||||
2. What is the per-core area budget, in absolute terms or as a fraction of die area?
|
||||
3. What is the target workload mix (cryptographic, HPC, server, general-purpose, embedded)?
|
||||
4. Will XH-1 implement the "B" (bitmanip) extension, which adds many pseudo-multiplication operations? If so, the MUL/DIV unit may be partially reused or augmented.
|
||||
5. Is the 128-core fabric a true SMP (Linux runs on all 128 cores), or is it a heterogeneous mix (e.g., application cores + management cores)? This affects MUL/DIV uniformity requirements.
|
||||
6. What is the target process node? Node geometry strongly affects the area and timing trade-offs of Booth vs. array multipliers.
|
||||
7. Will the MUL/DIV unit be hardened as a macro and replicated 128 times, or will it be synthesized per-core? Hardening is favored for area/power and replication regularity; synthesis-per-core favors design flexibility.
|
||||
8. Is there an asymmetric verification budget for the MUL/DIV unit, or is verification cost unconstrained?
|
||||
9. Does XH-1 need a 64×64→128-bit MUL for the upper-multiply operations to be single-cycle, or is multi-cycle acceptable?
|
||||
10. What is the expected frequency of `DIV`/`REM` instructions in the target workload? Server workloads may have many fewer than scientific workloads.
|
||||
11. Will the MUL/DIV unit be shared with the FP pipeline (as in Rocket / BOOM), or kept private to the integer pipeline?
|
||||
12. Will XH-1 adopt a shared multi-cycle divider across cores (Alternative F), or is a strictly per-core MUL/DIV unit required?
|
||||
|
||||
## Sources
|
||||
|
||||
- Canonical computer-arithmetic references: Parhami, *Computer Arithmetic*; Ercegovac and Lang, *Digital Arithmetic*; Hennessy and Patterson, *Computer Architecture*. These cover the techniques surveyed above in standard form. **Caveat:** these texts underwrite the taxonomy and the mechanism descriptions. Specific per-cycle latency figures given in this document are **TYPICAL** values drawn from common practice in published RV64 designs; the cited texts provide the algorithmic background but do not, in their canonical editions, supply XH-1-specific latency or area numbers. No specific chapter or page is cited for the quantitative figures because the figures are not drawn from a single source.
|
||||
- Open RISC-V core implementations (Rocket, BOOM, XiangShan, SiFive) provide reference designs for radix-4 array multipliers, iterative and SRT dividers, and FP/integer multiplier sharing. These are cited as implementation exemplars, not as XH-1 references.
|
||||
- **INSUFFICIENT EVIDENCE:** No XH-1-internal sources exist for this topic. The repository document `research/03-core-design/mul-div-unit.md` is itself a stub ("SOON"), and all sibling documents in `research/03-core-design/` are stubs. No prior architectural decisions are recorded.
|
||||
- No external citations, papers, benchmarks, or measurements have been invented for this document. All comparative claims are framed as qualitative engineering estimates or as canonical background knowledge in the field of computer arithmetic.
|
||||
+317
File diff suppressed because one or more lines are too long
+387
@@ -0,0 +1,387 @@
|
||||
# Multiply/Divide Unit Design for the XH-1 Processor
|
||||
|
||||
## Status
|
||||
|
||||
**Stub document — research in progress.** No architectural decisions have been finalized. The XH-1 repository contains no prior decisions on the multiply/divide (MUL/DIV) unit. All content below is a structured engineering analysis of design options, framed as proposals and open questions. Quantitative comparisons are presented only as qualitative relative magnitudes, never as benchmark figures. Throughout this document, claims are tagged as one of:
|
||||
|
||||
- **FACT** — well-established in the cited literature or in the RISC-V ISA specification.
|
||||
- **TYPICAL** — the common case across published designs; implementation-specific values may vary.
|
||||
- **ASSUMPTION** — an explicit premise the analysis depends on; should be revisited.
|
||||
- **PROPOSAL** — a design recommendation conditional on unresolved parameters.
|
||||
- **INSUFFICIENT EVIDENCE** — no defensible claim can be made without additional information.
|
||||
|
||||
## Abstract
|
||||
|
||||
This document investigates the design of the integer multiply and divide unit (MUL/DIV) for the XH-1, a custom 128-core RISC-V processor. The MUL/DIV unit is a latency-critical but throughput-amortizable functional unit. Its design has outsized impact on pipeline depth, area per core, and verification effort, all of which compound across 128 replicated cores. We survey the principal implementation strategies (iterative shift-and-add, radix-4/8 Booth, array multipliers, Goldschmidt dividers, Newton-Raphson dividers, subtractive dividers, and SRT dividers) and analyze their trade-offs with respect to latency, throughput, area, power, and verification complexity. The document is intentionally non-prescriptive: the XH-1 has not yet established a baseline for core microarchitecture (in-order vs. out-of-order, pipeline depth), so any recommendations are conditional on those upstream decisions.
|
||||
|
||||
## Research Question
|
||||
|
||||
> What is the appropriate microarchitecture for the integer MUL/DIV unit in the XH-1 core, given that the unit will be replicated 128 times, must support the RV64M extension, and must satisfy XH-1's target latency, area, and power budget — the latter two of which are not yet defined?
|
||||
|
||||
Specifically:
|
||||
|
||||
- How should the unit be pipelined relative to the rest of the core?
|
||||
- Should it share silicon with adjacent datapath structures (ALU, register file read/write ports, FP multiplier) or remain isolated?
|
||||
- How should division throughput be traded against area and verification cost?
|
||||
- What is the impact of the choice on multi-core contention for shared execution resources, if any?
|
||||
|
||||
## Background
|
||||
|
||||
**FACT.** The RISC-V "M" standard extension defines four signed/unsigned variants of multiply and two signed/unsigned variants of divide plus remainder:
|
||||
|
||||
- `MUL` (lower 64 bits of signed × signed)
|
||||
- `MULH`, `MULHU`, `MULHSU` (upper 64 bits, various sign combinations)
|
||||
- `DIV`, `DIVU` (signed/unsigned quotient)
|
||||
- `REM`, `REMU` (signed/unsigned remainder)
|
||||
|
||||
Critical properties of this ISA slice that drive microarchitecture:
|
||||
|
||||
- **Wide-operand upper-multiply (UMUL).** **TYPICAL.** In a single full-width partial-product / carry-save array, the 64×64→128-bit product is generated by the same array that produces the 64×64→64 product; the carry-save adder tree is slightly deeper and wider, and a final carry-propagate adder is needed to collapse the upper half. The incremental cost of supporting `MULH*` is therefore a small area adder and routing for the upper output, not a doubling. *Caveat: specific silicon area deltas are design-dependent; the "share the array" pattern is the common case in published RV64 implementations (e.g., BOOM, XiangShan, some SiFive designs), but quantitative numbers are not asserted here.* **FACT.** Rocket Chip keeps the integer multiplier and the FP multiplier as separate units rather than sharing the array; the document's earlier blanket attribution of sharing to "Rocket, BOOM, and most SiFive cores" overstates the case and is corrected here.
|
||||
- **MULH and signed×signed handling.** **FACT.** `MULH` (signed×signed, upper half) is not obtained by simply reusing the unsigned 64×64→128 array with sign-corrected operands. The standard technique is Baugh-Wooley or Modified Booth with explicit sign-bit handling, which modifies the partial-product generation (sign-extension of the most-significant partial products) and the adder tree. The "essentially free" characterization sometimes seen is oversimplified: while the underlying adder tree is shared, the partial-product array and the final CPA differ for the signed case, and verification must treat `MULH` as a distinct datapath.
|
||||
- **MULHSU and signed×unsigned handling.** **FACT.** `MULHSU` is also not a vanilla 65×64 unsigned array. The signed operand's most-significant partial product must be sign-handled (Baugh-Wooley-style sign extension of the MSB partial product, or Modified Booth encoding with explicit sign control) so that the upper 64 bits of the result are correct. The incremental verification cost over `MULHU` is small once the unsigned array is in place, but `MULHSU` is not "free" in the strict sense; it requires its own sign-handling pass and must be verified against a reference for all sign combinations.
|
||||
- **Division latencies.** **FACT.** A restoring or non-restoring subtractive divider on 64-bit operands requires exactly 64 reduction steps (or 65 with a sign pre-correction step). A radix-2 SRT divider also requires 64 selection steps in the worst case (with possible skipped steps on average, but the worst case governs the pipeline). A radix-4 SRT divider requires 16 selection steps in the worst case, plus a final quotient-conversion step (carry-save to two's-complement), not an extra selection step. **TYPICAL.** These counts dominate pipeline depth if the divider is fully combinational; iterative implementations amortize them over many cycles.
|
||||
- **Signed semantics.** **FACT.** RISC-V specifies the following for division edge cases:
|
||||
- Division by zero: `DIV` and `DIVU` return `-1` (i.e., all bits set); `REM` and `REMU` return the dividend.
|
||||
- Signed overflow: `INT64_MIN / -1` returns `INT64_MIN` (the mathematical quotient); the corresponding `REM` returns 0.
|
||||
- The operation must not raise an exception; the hardware must produce the specified result.
|
||||
- **Throughput vs. latency decoupling.** **FACT.** RISC-V MUL/DIV results are written to the integer register file, so latency can be hidden by OoO scheduling — but in an in-order core, latency is directly exposed to the front-end unless explicit forwarding is provided.
|
||||
|
||||
The XH-1 is a 128-core machine. **FACT.** Decisions in the MUL/DIV unit replicate 128×, so per-core area dominates the silicon cost; verification effort is dominated by unit-level and core-level-integration work that is performed once and reused across the 128 identical instances (see §Verification Considerations and §128-Core Scalability for the qualification). **INSUFFICIENT EVIDENCE:** There is no public XH-1 microarchitecture document in the repository that would constrain the MUL/DIV design (e.g., pipeline depth, issue width, in-order/OoO).
|
||||
|
||||
**INSUFFICIENT EVIDENCE:** Pipeline depth, issue width, in-order vs. OoO status, target frequency, target process node, per-core area budget, and target workload mix for XH-1 are all unknown.
|
||||
|
||||
## Existing Approaches
|
||||
|
||||
The following approaches are well-established in the literature and commercial designs. **FACT.** Standard computer-arithmetic texts (Parhami, *Computer Arithmetic*; Ercegovac and Lang, *Digital Arithmetic*; Hennessy and Patterson, *Computer Architecture*) cover these techniques in canonical form. No XH-1-internal prior art exists. Per-cycle latency figures given below are **TYPICAL** values for a 64-bit operand at a moderate clock target; specific values are implementation- and node-dependent, and the figures are not drawn from a single citable source. Where a range is given, it is illustrative of the order of magnitude, not a tight bound.
|
||||
|
||||
### 1. Shift-and-Add Multiplier (Iterative)
|
||||
|
||||
- **Mechanism:** **FACT.** A 64-bit multiplier is implemented as a state machine that processes one partial-product bit per cycle against a 128-bit accumulator (or a 129-bit accumulator with a sign-preconditioned variant). Signed operands are sign-extended; iteration count is not halved.
|
||||
- **Latency:** **TYPICAL.** ~64 cycles for a 64-bit operand, plus a small constant for sign/result correction.
|
||||
- **Throughput:** **TYPICAL.** One multiply per ~64 cycles (unit is not pipelined).
|
||||
- **Area:** **TYPICAL.** Very small. Roughly one wide adder + one shifter + one accumulator register.
|
||||
- **Power:** **TYPICAL.** Low. Minimal clocked area per cycle.
|
||||
- **Verification:** **TYPICAL.** Low complexity. Straightforward to model and exhaustively test at small operand widths.
|
||||
- **Use case:** **TYPICAL.** Embedded in-order cores where MUL/DIV are infrequent and latency-tolerant. Generally unsuitable for high-performance 64-bit cores.
|
||||
|
||||
### 2. Radix-4 Booth-Encoded Multiplier (Array or Iterative)
|
||||
|
||||
- **Mechanism:** **FACT.** Radix-4 Booth recoding reduces partial products to ceil(n/2). For a 64-bit operand, standard radix-4 Booth encoding produces 32 partial-product rows. An *iterative* radix-4 multiplier accumulates these rows one or two at a time:
|
||||
- **One row per cycle:** ~32 cycles, one CSA per cycle, smallest iterative area.
|
||||
- **Two rows per cycle:** ~16 cycles, but requires two CSAs in series per cycle. The area cost of the second CSA is not a simple "doubling": the second CSA operates on the full sum-and-carry width of the first, so the additional area is closer to the cost of one full-width CSA, and the per-cycle critical path lengthens. This is the configuration that achieves the "16-cycle" figure sometimes cited; the area and timing costs must be acknowledged.
|
||||
- **FACT.** Implemented as a single combinational Wallace/Dadda tree, the same 32 partial-product rows are summed in one cycle, with the tree depth determining the achievable clock period.
|
||||
- **Latency:** **TYPICAL.** ~32 cycles iterative (one row/cycle) or ~16 cycles iterative (two rows/cycle, with the area and timing qualifications above), or one combinational tree of approximately 8–12 CSA levels plus a final CPA (the level count is design- and library-specific; the figure is an order-of-magnitude estimate, not a precise bound). When pipelined, the array is typically broken into 3–6 stages, with each stage absorbing one to several CSA levels plus possibly a portion of the final CPA. The relationship between the un-pipelined tree depth and the pipelined stage count is implementation-specific; this document does not assert a fixed mapping.
|
||||
- **Throughput:** **TYPICAL.** One multiply per ~32 cycles (one row/cycle iterative), ~16 cycles (two rows/cycle iterative, with qualifications), or 1/cycle if the array is fully pipelined.
|
||||
- **Area:** **TYPICAL.** Moderate. Recoder + partial-product array + carry-save adder tree + final CPA. The "two rows per cycle" iterative variant is roughly comparable in area to a small pipelined array, with the per-cycle critical path lengthened.
|
||||
- **Power:** **TYPICAL.** Moderate to high when pipelined. The Wallace/Dadda tree toggles aggressively, and clock-tree load on a replicated array is non-trivial.
|
||||
- **Verification:** **TYPICAL.** Moderate. The corner cases that matter are the signed-overflow cases in the `MULH` datapath (sign-extended partial products, modified tree inputs) and the `MULHSU` sign-extension path (Baugh-Wooley or Modified Booth sign handling, not a free byproduct of the unsigned array). The unsigned 64×64→128 array, once verified, covers `MUL`, `MULHU`, and the bulk of `MULHSU`; `MULH` is a separate verification pass; `MULHSU` requires explicit sign-handling verification but is closer to the unsigned case than `MULH` is.
|
||||
- **Use case:** **TYPICAL.** Mid- and high-performance cores; the default choice for most modern desktop-class RV64 cores.
|
||||
|
||||
### 3. Radix-8 or Radix-16 Booth Multiplier
|
||||
|
||||
- **Mechanism:** **FACT.** Higher-radix Booth recoding reduces the number of partial-product rows. For a 64-bit operand:
|
||||
- Radix-4: 32 rows.
|
||||
- Radix-8: 22 data rows plus a separate sign-handling row for the most-significant window, totaling 23 rows in a fully sign-corrected implementation. Canonical radix-8 Booth recoding requires careful handling of the most-significant 3-bit window to avoid producing an erroneous extra row. The PPG must produce multiples {0, ±1, ±2, ±3, ±4} of the multiplicand; ±3× and ±4× are typically generated via a carry-save adder (1× + 2× for ±3×, 2× + 2× or a dedicated shift-and-add for ±4×), so the PPG is substantially more complex than radix-4.
|
||||
- Radix-16: 16 data rows plus a sign-handling row, totaling 17 rows. The PPG must produce multiples {0, ±1, ±2, ±3, ±4, ±5, ±6, ±7, ±8}; 3×, 5×, 6×, 7× are typically generated via combinations of smaller multiples, with an extra high-order term.
|
||||
- **Correction:** The "radix-8 reduces by 3×, radix-16 by 4×" claim sometimes seen in the literature refers to the ratio relative to radix-2 (64 partial products → 22 or 16 data rows), not a clean 3× or 4× multiplier. The actual reductions over radix-2 are 64/22 ≈ 2.9× (radix-8) and 64/16 = 4× (radix-16); the reductions over radix-4 are 32/22 ≈ 1.45× and 32/16 = 2× respectively.
|
||||
- **Latency:** **TYPICAL.** The fully pipelined radix-4 array can already achieve 1/cycle throughput; higher radices reduce the *depth* of the adder tree (fewer rows to sum) and therefore either shorten the critical path or allow fewer pipeline stages. Throughput is not increased beyond 1/cycle unless the array is duplicated.
|
||||
- **Throughput:** **TYPICAL.** 1/cycle for a single pipelined array; not inherently higher than radix-4.
|
||||
- **Area:** **TYPICAL.** Larger PPG; smaller (shallower) adder tree. Net area is roughly comparable to radix-4 or slightly larger.
|
||||
- **Power:** **TYPICAL.** Mixed. Fewer adder levels, but more complex PPG.
|
||||
- **Verification:** **TYPICAL.** Higher. Radix-8+ PPGs have more corner cases and the recoding is harder to prove correct.
|
||||
- **Use case:** **TYPICAL.** High-frequency designs where the critical path is in the adder tree, not the PPG.
|
||||
|
||||
### 4. Iterative Subtractive Divider (Non-Restoring / Restoring)
|
||||
|
||||
- **Mechanism:** **FACT.** Standard shift-subtract over the operand width. Produces quotient (and optionally remainder) one bit per cycle. The iteration count for a 64-bit operand is exactly 64 reduction steps (or 65 with a sign pre-correction step); the figure is not a range.
|
||||
- **Latency:** **TYPICAL.** ~64 cycles for a 64-bit operand (plus a small constant for sign correction).
|
||||
- **Throughput:** **TYPICAL.** One divide per ~64 cycles.
|
||||
- **Area:** **TYPICAL.** Small. A 64- or 65-bit ALU with quotient/remainder registers.
|
||||
- **Power:** **TYPICAL.** Low.
|
||||
- **Verification:** **TYPICAL.** Moderate. The division-by-zero convention, the `INT64_MIN / -1` overflow case, and the `REM`/`REMU` dividend-return case must all be implemented and tested explicitly. The signed-dividend path is the principal source of bugs.
|
||||
- **Use case:** **TYPICAL.** In-order cores with low expected DIV frequency.
|
||||
|
||||
### 5. Radix-4 SRT Divider
|
||||
|
||||
- **Mechanism:** **FACT.** Higher-radix SRT produces multiple quotient digits per iteration by selecting one of several shifted multiples of the divisor from a selection table indexed by a truncated partial remainder. A radix-4 SRT produces 2 bits per iteration; the quotient is held in a redundant (carry-save) form and converted to two's-complement on completion. The iteration count for a 64-bit operand is 16 selection steps in the worst case, plus a final quotient-conversion step (carry-save to two's-complement); the converter is not an extra selection step.
|
||||
- **Latency:** **TYPICAL.** ~16 selection cycles plus a small constant for the final conversion, for a 64-bit operand.
|
||||
- **Throughput:** **TYPICAL.** One divide per ~16 cycles (worst case).
|
||||
- **Area:** **TYPICAL.** Substantially larger than subtractive. Requires a redundant (carry-save) quotient representation, a quotient-digit selection table, and partial-quotient error-correction logic.
|
||||
- **Power:** **TYPICAL.** Higher.
|
||||
- **Verification:** **TYPICAL.** High. SRT selection tables have notoriously tricky corner cases, and partial-quotient error correction has been the source of silicon bugs in commercial designs. **FACT.** The classic example is the Pentium FDIV bug: the floating-point divider's radix-4 SRT lookup table was missing entries (a "+2" entry that should have been present) in the programmable logic array (PLA) implementing the selection function. The fix was a mask change, not a logic redesign. **TYPICAL.** The lessons from this and similar incidents — that the interaction between the redundant quotient representation and the selection function produces error patterns that are not obvious from inspection, and that verification typically requires formal proofs of the selection function over reduced operand widths plus extensive directed testing — transfer to integer radix-4 SRT dividers, but the FP and integer SRT implementations use different quotient-digit sets and selection functions, so the lessons are transferred by analogy rather than by direct equivalence.
|
||||
- **Use case:** **TYPICAL.** High-frequency OoO cores needing higher DIV throughput.
|
||||
|
||||
### 6. Goldschmidt Divider (Reciprocal Multiplication)
|
||||
|
||||
- **Mechanism:** **FACT.** Iteratively refine an initial reciprocal estimate, then multiply by the dividend. Each iteration multiplies both the partial remainder and the partial quotient by a correction factor derived from a short reciprocal estimate. Distinct from Newton-Raphson (see §7).
|
||||
- **Latency:** **TYPICAL.** A few multiply iterations. Amortized low latency if the multiplier is fast.
|
||||
- **Throughput:** **TYPICAL.** Bounded by the multiplier.
|
||||
- **Area:** **TYPICAL.** Reuses the multiplier. Adds a small amount of control logic and lookup-table ROM for the initial estimate.
|
||||
- **Power:** **TYPICAL.** Low marginal cost on top of the multiplier.
|
||||
- **Verification:** **TYPICAL.** High. The final iteration must converge to enough bits of precision to guarantee correct rounding toward zero for all 64-bit inputs, which typically requires a final exact-fixup step. The accuracy analysis is non-trivial and historically a bug source.
|
||||
- **Use case:** **TYPICAL.** Designs that already have a fast pipelined multiplier and want a small, fast divider.
|
||||
|
||||
### 7. Newton-Raphson Divider (Reciprocal Multiplication)
|
||||
|
||||
- **Mechanism:** **FACT.** Iteratively refine an initial estimate of the divisor's reciprocal using a Newton-Raphson step, then multiply the dividend by the refined reciprocal. Each iteration squares the error, so convergence is quadratic. Distinct from Goldschmidt, which uses a multiplicative correction on both the partial remainder and the partial quotient simultaneously; the two algorithms have different error dynamics and different fixup requirements.
|
||||
- **Latency:** **TYPICAL.** A few iterations of multiply-add. Typically fewer iterations than Goldschmidt to reach a given precision, but each iteration is a full multiply.
|
||||
- **Throughput:** **TYPICAL.** Bounded by the multiplier.
|
||||
- **Area:** **TYPICAL.** Reuses the multiplier. Adds lookup-table ROM for the initial estimate and modest control logic.
|
||||
- **Power:** **TYPICAL.** Low marginal cost on top of the multiplier.
|
||||
- **Verification:** **TYPICAL.** High. The final-step rounding-toward-zero correctness across all 64-bit input pairs requires a careful proof of the fixup step.
|
||||
- **Use case:** **TYPICAL.** Designs with a fast pipelined multiplier that want a low-latency divider.
|
||||
|
||||
### 8. Shared MUL/DIV Iterative Unit (Multiplexed)
|
||||
|
||||
- **Mechanism:** **ASSUMPTION.** A single iterative datapath handles both MUL and DIV by reconfiguring its datapath between operations. **TYPICAL.** This pattern is more common in microcoded embedded cores than in modern 64-bit RV64 designs, where MUL and DIV datapaths are structurally different (shift-and-add with accumulator vs. shift-subtract with quotient register) and the area savings from sharing are modest compared to the control complexity of reconfiguration. The characterization in this document is qualified accordingly.
|
||||
- **Latency:** **TYPICAL.** Same as the underlying iterative unit; the unit cannot serve MUL and DIV concurrently.
|
||||
- **Throughput:** **TYPICAL.** MUL and DIV contend for the same unit.
|
||||
- **Area:** **TYPICAL.** For RV64, the area advantage of a genuinely shared iterative unit over a split iterative MUL + iterative DIV is modest at best, and the control complexity of reconfiguration may offset the savings. The "most area-efficient" framing in earlier drafts of this document is qualified here: shared-iterative is a defensible choice for microcoded embedded cores, but for RV64 the area ranking is closer to "small to moderate" rather than strictly "smallest." **ASSUMPTION.** This ranking depends on the datapath being genuinely shared rather than microcoded over separate datapaths.
|
||||
- **Power:** **TYPICAL.** Low.
|
||||
- **Verification:** **TYPICAL.** Low to moderate if the datapath is genuinely shared; higher if microcode overlays separate datapaths.
|
||||
- **Use case:** **TYPICAL.** Cost-sensitive embedded cores; uncommon in high-performance RV64.
|
||||
|
||||
## Alternative Designs
|
||||
|
||||
Beyond the standard taxonomy, several non-obvious variants are worth considering.
|
||||
|
||||
### A. Pipelined Multiplier + Sequential Divider (Split Unit)
|
||||
|
||||
A fully pipelined radix-4 multiplier (e.g., 3–6 stage pipeline) is paired with an independent sequential subtractive divider. The two units have separate reservation/issue logic.
|
||||
|
||||
- **Rationale:** **ASSUMPTION.** Under many server, desktop, and general-purpose workloads, MUL is more frequent than DIV, but the ratio is workload-dependent and should not be assumed a priori. Giving MUL a fast pipelined datapath and DIV a slow sequential one matches a plausible workload asymmetry.
|
||||
- **Cost:** Two reservation-station or issue-queue entries (relevant if the core is OoO); two small register files for intermediate state.
|
||||
|
||||
### B. Variable-Latency Divider (Speculative / Early-Quit)
|
||||
|
||||
Produce a partial-width approximate quotient, sign-extend, and perform a single correction step on the remaining bits. The early-quit path saves cycles when the divisor has small magnitude.
|
||||
|
||||
- **Risk:** **TYPICAL.** Variable latency complicates the scoreboard / wakeup logic in an OoO core, or stalls the pipeline in an in-order core. This can be partially mitigated by a fixed maximum latency with early completion, at the cost of additional control logic.
|
||||
|
||||
### C. Radix-2 SRT with Carry-Save Quotient + Software/Hardware Fixup
|
||||
|
||||
A radix-2 SRT with a 1-bit-per-cycle selection function and a carry-save quotient register. Faster than naive subtractive, simpler than radix-4 SRT.
|
||||
|
||||
- **Cost:** The quotient-to-2's-complement conversion is on the critical path or requires a final fixup step. The worst-case iteration count is 64 selection steps for a 64-bit operand.
|
||||
|
||||
### D. Pipelined 2-bit-per-cycle "Naïve" Divider
|
||||
|
||||
Iterate two shift-subtract steps per cycle. Roughly halves divider latency relative to radix-1 subtractive without requiring an SRT selection table.
|
||||
|
||||
- **Cost:** Two wide adders and a longer critical path. Area is moderate.
|
||||
|
||||
### E. Iterative Multiplier (Pipelined) + Lookup-Table Reciprocal + Newton-Raphson
|
||||
|
||||
A small reciprocal ROM (e.g., 12-bit seed) followed by a Newton-Raphson refinement step using the iterative multiplier. Yields a divider latency of a few multiplier iterations.
|
||||
|
||||
- **Cost:** **TYPICAL.** The Newton iteration must converge to enough bits of reciprocal precision to guarantee correct rounding toward zero for all 64-bit inputs, which typically requires a final exact-fixup step (one or two multiplies plus a comparison). The precision / error analysis is non-trivial.
|
||||
|
||||
### F. Shared Multi-Cycle Divider Across Cores (Divider Co-Processor)
|
||||
|
||||
A small number of high-throughput dividers (e.g., 4 or 8) placed at fixed points in the fabric and dispatched to by cores via a memory-mapped or message interface. The dividers are not private to any core.
|
||||
|
||||
- **Rationale:** Avoids replicating the divider 128×. Trades single-core latency (now includes a fabric round-trip) for amortized area.
|
||||
- **Cost:** NOC traffic, dispatch latency, contention at the divider, and software-visible ABI changes (or a transparent-but-slow trap path).
|
||||
- **Status:** Unusual but not unprecedented in accelerator-rich many-core designs.
|
||||
|
||||
### G. FP / Integer Multiplier Sharing
|
||||
|
||||
Share the integer multiplier's partial-product array and adder tree with the FP pipeline. **TYPICAL.** This pattern appears in BOOM and in some SiFive designs; Rocket Chip keeps the integer and FP multipliers as separate units. The applicability of the pattern is design-specific and is not asserted as universal here.
|
||||
|
||||
- **Rationale:** Avoids replicating a wide datapath. The FP pipeline also benefits from a fast multiplier.
|
||||
- **Cost (qualified):** **TYPICAL.** Cross-unit scheduling and bypassing complexity. The integer and FP pipelines may have different latency targets. The sharing requires operand-format conversion (integer operands to FP-like internal format, and vice versa) and FP-specific concerns (rounding mode support, subnormal handling, NaN propagation) are not "free" — they are offloaded to the FP pipeline's existing logic, but the integer side must correctly drive and consume the shared datapath. Verification must cover the combined integer-plus-FP datapath, which is more complex than either alone. The "near-free" characterization sometimes seen in the literature is oversimplified; the cost is real but is often dominated by the FP-pipeline logic that already exists, making the incremental cost on the integer side smaller than the absolute cost of a separate integer multiplier.
|
||||
|
||||
### H. Latency-Tolerant In-Order MUL/DIV
|
||||
|
||||
Even a multi-cycle iterative MUL/DIV may be tolerable in an in-order core if the result-bus supports forwarding directly from the MUL/DIV output to dependent consumers, bypassing the register file writeback-read path.
|
||||
|
||||
- **Cost:** Forwarding path length and bypass-network complexity scale with MUL/DIV latency. The forwarding network is a real cost — typically a set of wide muxes at the input of each consuming execution unit, with wiring that may dominate the area of the iterative MUL/DIV unit itself. The cost scales with the number of consumers (ALU, branch, load/store) and with MUL/DIV latency, since the forwarded result must remain valid on the bypass network for the full MUL/DIV latency. This cost should be quantified before an iterative unit is chosen for an in-order core.
|
||||
|
||||
### I. Interaction with the "B" (Bitmanip) Extension
|
||||
|
||||
The RISC-V Bitmanip extension introduces MUL/DIV-adjacent operations (e.g., `CLZ`, `CTZ`, `MIN`, `MAX`, bit-extract/deposit, and several pseudo-multiplication idioms such as `RORI` and `SH*ADD`). If the B extension is in scope, the MUL/DIV unit may either be reused for some of these (e.g., via the ALU) or augmented with dedicated bitmanip datapath. The decision is interdependent with the MUL/DIV choice and is flagged as an open question below.
|
||||
|
||||
## Comparison
|
||||
|
||||
The following table presents *qualitative* relative magnitudes only. **INSUFFICIENT EVIDENCE:** No node, frequency, or synthesis data is available for XH-1, so quantitative ratios are not asserted. The latency and throughput figures are **TYPICAL** values for a 64-bit operand and are presented as order-of-magnitude estimates; the ranges are wider than in the prior draft to reflect the absence of a citable source. "Latency" is in cycles for back-to-back independent operations on the named unit; "throughput" is sustained operations per cycle for a fully pipelined or iterative unit, respectively.
|
||||
|
||||
| Approach | Latency (cycles) | Throughput (ops/cycle) | Relative Area | Relative Power | Verification Effort |
|
||||
|---|---|---|---|---|---|
|
||||
| Iterative shift-add MUL (one row/cycle) + iterative subtractive DIV (separate) | ~64 (MUL) / ~64 (DIV) | ~1/64 (each) | Smallest | Smallest | Low |
|
||||
| Iterative radix-4 MUL (one row/cycle) + iterative subtractive DIV (separate) | ~32 (MUL) / ~64 (DIV) | ~1/32 (MUL), ~1/64 (DIV) | Small | Small to moderate | Low to moderate |
|
||||
| Shared iterative MUL/DIV (multiplexed, Approach 8) | ~64 (MUL) / ~64 (DIV), mutually exclusive | ~1/64 (each, contended) | Small to moderate (modest savings over split iterative; control overhead) | Small to moderate | Low to moderate |
|
||||
| Pipelined radix-4 array MUL + iterative subtractive DIV (split, Alternative A) | ~3–6 (MUL) / ~64 (DIV) | 1/cycle (MUL), ~1/64 (DIV) | Moderate | Moderate | Moderate |
|
||||
| Pipelined radix-4 array MUL + 2-bit-per-cycle naïve DIV (Alternative D) | ~3–6 (MUL) / ~32 (DIV) | 1/cycle (MUL), ~1/32 (DIV) | Moderate | Moderate | Moderate |
|
||||
| Pipelined radix-4 array MUL + radix-4 SRT DIV | ~3–6 (MUL) / ~16 + conversion (DIV) | 1/cycle (MUL), ~1/16 (DIV) | Large | Large | High |
|
||||
| Pipelined radix-8 array MUL + Newton-Raphson or Goldschmidt DIV | ~3–6 (MUL) / a few MUL iterations (DIV) | 1/cycle (MUL), bounded by MUL (DIV) | Largest | Largest | High |
|
||||
|
||||
**INSUFFICIENT EVIDENCE:** The relative magnitudes are illustrative and intended only to convey ordering. The actual ratios depend on the target process node, synthesis library, and clock target, none of which are known for XH-1. **INSUFFICIENT EVIDENCE:** Quantitative area, power, and energy comparisons cannot be made without a target node, frequency, and synthesis flow. The ranges given above are wider than the typical figures cited in the prior draft to reflect this uncertainty.
|
||||
|
||||
## Advantages
|
||||
|
||||
The leading candidates each have a distinct strength profile. The following are conditional on the workload and core microarchitecture, which are not yet established.
|
||||
|
||||
- **Iterative separate MUL + iterative separate DIV (smallest, lowest power):** **ASSUMPTION.** Smallest per-core area and lowest power among the candidate RV64 designs. Easiest to verify. Long latency is the principal disadvantage.
|
||||
- **Shared iterative unit (Approach 8, with the qualification in §8):** **ASSUMPTION.** A defensible choice for cost-sensitive embedded cores, but for RV64 the area advantage over a split iterative design is modest and the control complexity of multiplexing MUL and DIV may offset the savings. Not recommended by default for RV64.
|
||||
- **Pipelined radix-4 MUL + iterative subtractive DIV (split, Alternative A):** **ASSUMPTION.** Matches a plausible workload asymmetry where MUL is more frequent than DIV; MUL throughput is high (common case), DIV cost is contained, and verification is tractable. This is a strong compromise candidate, not a leading candidate by default.
|
||||
- **Pipelined radix-4 MUL + SRT DIV:** **TYPICAL.** Highest predictable throughput on both operations; minimal front-end exposure if the core is in-order. Verification cost is high.
|
||||
- **Newton-Raphson or Goldschmidt DIV on top of fast MUL:** **TYPICAL.** Reuses the multiplier's silicon; area-efficient if the multiplier is already large. Verification cost is high.
|
||||
|
||||
## Disadvantages
|
||||
|
||||
- **Iterative separate MUL + iterative separate DIV:** **TYPICAL.** Latency directly stalls the pipeline in an in-order core. Long DIV latencies are problematic for tight feedback loops unless forwarding is provided, and the forwarding path itself has non-trivial cost (see Alternative H).
|
||||
- **Shared iterative unit:** **TYPICAL.** Latency directly stalls the pipeline in an in-order core. Long DIV latencies are problematic for tight feedback loops unless forwarding is provided. For RV64, the structural mismatch between MUL and DIV datapaths limits the achievable area savings, and the MUL/DIV datapaths contend for the shared unit.
|
||||
- **Pipelined radix-4 MUL + iterative DIV:** **TYPICAL.** DIV latency is still long. The dual-unit bookkeeping (two reservation stations, two issue paths) adds microarchitectural complexity.
|
||||
- **Pipelined radix-4 MUL + SRT DIV:** **TYPICAL.** Highest area and verification cost. SRT selection-table corner cases are a known source of post-silicon bugs (the Pentium FDIV bug being a radix-4 SRT selection-table defect, transferred by analogy to integer SRT). The verification cost is replicated at the unit level and is the dominant non-silicon cost of this design.
|
||||
- **Newton-Raphson / Goldschmidt DIV:** **TYPICAL.** Rounding-toward-zero correction is subtle; commercial implementations have shipped with latent bugs in this path. Not recommended without a clean, formally-specified correction step.
|
||||
|
||||
## XH-1 Considerations
|
||||
|
||||
**PROPOSAL (pending XH-1 baseline):** The XH-1 has not yet established a microarchitecture baseline. Several XH-1-specific factors bear on the MUL/DIV choice:
|
||||
|
||||
1. **128-core replication amplifies per-core area cost.** **FACT.** A larger MUL/DIV unit pays the same area cost across all 128 cores, not just one. The die-area cost is severe. Verification effort is *not* amplified by the same factor (see item 4 and §Verification Considerations); the framing in earlier drafts of this document that listed "per-core area and verification cost" together as both amplified by replication conflates two distinct effects and is corrected here.
|
||||
2. **Per-core issue width and in-order/OoO status are not established.** **INSUFFICIENT EVIDENCE.** A wide-issue OoO core tolerates long-latency iterative dividers via scheduling; a narrow in-order core does not, unless explicit forwarding is provided.
|
||||
3. **128 cores imply a tile- or mesh-style interconnect.** **ASSUMPTION.** Long-latency MUL/DIV operations that block a single core are local to that core; the rest of the machine is unaffected. This *favors* simpler per-core MUL/DIV implementations, unless a shared divider alternative (Alternative F) is adopted.
|
||||
4. **Verification cost is concentrated at the unit level, not multiplied by replication.** **TYPICAL.** A replicated unit is verified once at the unit level (RTL, formal, directed/random). The integration with each core's pipeline is identical across replications and is verified once via the core-level verification environment; running the same integration suite 128× does not add coverage. The 128× replication matters for silicon defect exposure (a bug that escapes verification affects all cores) and for DFT/scan/BIST architecture, not for per-instance verification run-count. This is the standard methodology for replicated unit-level verification in commercial designs.
|
||||
5. **Physical-design regularity matters under replication.** **TYPICAL.** A small, regular MUL/DIV unit is easier to harden and replicate 128× than a complex, irregular SRT unit.
|
||||
|
||||
**OPEN QUESTION:** What is the target workload mix for XH-1? Cryptographic, scientific, server, or general-purpose? Each has very different MUL/DIV demands.
|
||||
|
||||
**OPEN QUESTION:** Is the 128-core fabric homogeneous (all cores identical, all running the same software) or heterogeneous (e.g., application cores plus management or I/O cores)? Heterogeneity would relax the per-core MUL/DIV uniformity requirement and may allow per-tile optimization.
|
||||
|
||||
## 128-Core Scalability
|
||||
|
||||
The MUL/DIV unit is per-core and not shared across cores by default. Scalability considerations:
|
||||
|
||||
- **No coherence problem.** **FACT.** MUL/DIV operates on register values; the result is written to the integer register file, which is private to each core.
|
||||
- **No cross-core contention** in the default per-core configuration. **FACT.** Unlike shared caches or interconnects, the MUL/DIV unit does not introduce hot-spot behavior. A shared-divider alternative (Alternative F) changes this analysis.
|
||||
- **Verification parallelism.** **TYPICAL.** A 128-core machine can be partitioned for verification; the MUL/DIV unit is verified once at the unit level. Integration with the pipeline is verified once at the core level (since all cores are identical replications), not 128×. The 128× replication affects silicon defect exposure, not verification run-count.
|
||||
- **Area-budget pressure.** **TYPICAL.** If the per-core area budget is tight, the MUL/DIV unit competes with the L1 cache, branch predictor, and ROB for die area. Iterative implementations (subtractive DIV, iterative MUL) are most competitive in this regime.
|
||||
- **Power-grid and clock distribution.** **TYPICAL.** A replicated fast pipelined multiplier across 128 cores is a substantial clock-tree and switching load. **INSUFFICIENT EVIDENCE:** Power delivery (IR drop), clock skew across the replicated load, and dynamic power density must be analyzed for the replicated load; no quantitative estimates are made here.
|
||||
- **Scan and BIST.** **TYPICAL.** A 128× replicated unit implies 128× the scan-chain length (if scan is per-core) or a partitioned BIST architecture. The choice affects DFT area and test time.
|
||||
- **Fault tolerance.** **TYPICAL.** A defect in the MUL/DIV unit is potentially a defect in all 128 cores. This argues for either a hardened, characterized macro or built-in redundancy / sparing, depending on yield targets.
|
||||
- **Timing variation.** **TYPICAL.** Across-die process variation affects 128 replicated units independently. A design that is timing-marginal at one corner may fail at another. Iterative designs are less sensitive to per-unit timing variation than deep-pipelined arrays.
|
||||
|
||||
**PROPOSAL:** Favor an implementation whose area × power product is small, since both costs multiply by 128. Prefer regular, hardenable structures over irregular ones that resist replication.
|
||||
|
||||
## Performance Considerations
|
||||
|
||||
MUL/DIV performance interacts with the rest of the core in two ways:
|
||||
|
||||
1. **Latency exposure.** **FACT.** In an in-order core, MUL latency stalls the front-end unless explicit forwarding is provided. In an OoO core, MUL latency is amortized as long as other independent instructions exist.
|
||||
2. **Throughput-bound on dependent chains.** **FACT.** Tight loops of dependent MUL/DIV operations (e.g., big-integer arithmetic, hash functions, matrix multiplies) are throughput-bound, not latency-bound. For these, pipelined MUL at 1/cycle is essential.
|
||||
|
||||
**INSUFFICIENT EVIDENCE:** Without a target workload, no quantitative performance claim can be made.
|
||||
|
||||
## Area Considerations
|
||||
|
||||
Relative area ordering, qualitative only (no synthesis data available):
|
||||
|
||||
- Iterative shift-add MUL + iterative subtractive DIV (separate): smallest.
|
||||
- Iterative radix-4 MUL (one row/cycle) + iterative subtractive DIV (separate): small.
|
||||
- Shared iterative MUL/DIV (Approach 8): small to moderate, depending on the degree of datapath sharing and the control overhead; the area advantage over split iterative is modest for RV64.
|
||||
- Pipelined radix-4 MUL + iterative subtractive DIV (split): moderate.
|
||||
- Pipelined radix-4 MUL + 2-bit-per-cycle naïve DIV: moderate.
|
||||
- Pipelined radix-4 MUL + radix-4 SRT DIV: large.
|
||||
- Pipelined radix-8 MUL + Newton-Raphson / Goldschmidt DIV: largest.
|
||||
|
||||
**INSUFFICIENT EVIDENCE:** The XH-1 per-core area budget must be defined before any of the above can be quantified in absolute terms. Absolute area figures require a process node and a synthesis flow.
|
||||
|
||||
## Power and Energy Considerations
|
||||
|
||||
Relative energy-per-operation ordering, qualitative only:
|
||||
|
||||
- **Iterative:** **TYPICAL.** Low per-cycle power, but high per-operation energy × time product (many cycles).
|
||||
- **Pipelined MUL + iterative DIV:** **TYPICAL.** Low per-op MUL energy, high per-op DIV energy.
|
||||
- **Pipelined MUL + SRT DIV:** **TYPICAL.** High per-cycle power, lower per-op energy than iterative.
|
||||
- **Pipelined MUL + Newton-Raphson / Goldschmidt DIV:** **TYPICAL.** Per-op energy is a small multiple of MUL energy; competitive.
|
||||
|
||||
**PROPOSAL:** Energy-per-operation, not energy-per-cycle, is the more relevant metric for a workload-bound design. A pipelined radix-4 MUL with an iterative subtractive DIV is a reasonable energy-vs-area compromise, contingent on the workload.
|
||||
|
||||
**INSUFFICIENT EVIDENCE:** Quantitative power and energy figures require a process node, a clock target, and a workload trace.
|
||||
|
||||
## Implementation Considerations
|
||||
|
||||
- **Wallace/Dadda tree synthesis risk.** **TYPICAL.** Compressor trees are known to be hard to place-and-route at high frequency on modern nodes; poor placement can introduce unexpected delay. This favors iterative or low-radix designs for first-pass XH-1 silicon. *This is an engineering judgment, not an established fact; specific delay figures depend on the synthesis flow and library, and are not asserted here.*
|
||||
- **PPG and Booth recoder verification.** **TYPICAL.** Radix-4 PPGs are well-understood and tractable to verify; radix-8+ PPGs require more corner cases. The `MULH` signed×signed upper-half path requires explicit verification of sign-extended partial products and is not a free byproduct of the unsigned array. The `MULHSU` path requires explicit sign-handling verification (Baugh-Wooley or Modified Booth) and is not a vanilla 65×64 unsigned array.
|
||||
- **SRT selection table correctness.** **TYPICAL.** Selection tables must be proven correct for all operand values. This typically requires exhaustive enumeration for reduced operand widths and extensive directed testing for full width.
|
||||
- **RISC-V division-by-zero and overflow conventions.** **FACT.** `DIV` and `DIVU` return `-1` (all bits set) on divide-by-zero; `REM` and `REMU` return the dividend on divide-by-zero. On signed overflow (`INT64_MIN / -1`), `DIV` returns `INT64_MIN` and `REM` returns 0. These must be implemented explicitly; the design must not raise a trap.
|
||||
- **Pipeline interlocks.** **FACT.** If MUL is pipelined, the in-order scoreboard or the OoO wakeup logic must handle in-flight MULs; this is a per-instruction-record cost.
|
||||
- **Physical design replication.** **TYPICAL.** A 128-core replicated design benefits from a small, regular MUL/DIV unit. Highly irregular SRT selection logic and complex PPGs resist clean replication and are a physical-design risk under replication.
|
||||
|
||||
## Verification Considerations
|
||||
|
||||
Verification cost is a dominant non-silicon cost in modern CPU design. The MUL/DIV unit is a known hot-spot for bugs:
|
||||
|
||||
- **Radix-2/4 shift-add and subtractive:** **TYPICAL.** Lowest. Exhaustive simulation at reduced widths is tractable.
|
||||
- **Radix-4 Booth + iterative DIV:** **TYPICAL.** Moderate. The hard cases are the signed division edge cases (division-by-zero, `INT64_MIN / -1`), the `MULH` signed×signed upper-half path (sign-extended partial products, not a free byproduct of the unsigned array), and the `MULHSU` sign-extension (Baugh-Wooley or Modified Booth, not a vanilla unsigned array). The unsigned 64×64→128 array, once verified, covers `MUL`, `MULHU`, and the bulk of `MULHSU`; `MULH` is a separate verification pass; `MULHSU` requires explicit sign-handling verification.
|
||||
- **SRT:** **TYPICAL.** High. SRT selection-table bugs are famous in industry (the Pentium FDIV bug was a missing entry in the PLA implementing the radix-4 SRT lookup table of the floating-point divider; the lessons transfer to integer SRT by analogy, with the caveat that FP and integer SRT use different quotient-digit sets and selection functions). Verification typically requires formal proofs of the selection function over reduced widths and extensive directed testing.
|
||||
- **Newton-Raphson / Goldschmidt:** **TYPICAL.** High. Rounding-toward-zero correctness across all 64-bit input pairs requires a careful proof of the final fixup step.
|
||||
- **Replication factor:** **TYPICAL.** A bug in the MUL/DIV unit that escapes verification is a bug in all 128 cores (silicon defect exposure). However, the verification effort itself is not 128×: the unit is verified once at the unit level, and the core-level integration is verified once (cores are identical replications). The risk is concentrated exposure, not multiplied effort. This is the standard methodology for replicated unit-level verification in commercial designs.
|
||||
|
||||
**Preliminary verification strategy** (to be refined once the design choice is made):
|
||||
|
||||
- **Unit-level:** Exhaustive simulation at reduced operand widths (e.g., 8, 12, 16 bits) for the core datapath. Formal equivalence checking between the RTL and a reference model written in a high-level specification language (e.g., Bluespec, Scala, or a C reference). For SRT or Newton-Raphson, formal proof of the selection function or the convergence step.
|
||||
- **Directed corner-case suite:** Explicit tests for division-by-zero (both `DIV`/`DIVU` and `REM`/`REMU` paths), signed overflow (`INT64_MIN / -1`, both quotient and remainder), `MULH` against a cross-checked reference (Baugh-Wooley or equivalent), `MULHSU` against a cross-checked reference (with explicit sign-handling verification), and the `MUL`-then-truncate boundary.
|
||||
- **Random / constrained-random:** At full width, comparing against a software reference. Coverage targets on the Booth recoder, PPG, and selection table.
|
||||
- **Integration:** Per-core pipeline integration verified once at the core level (not 128×), since cores are identical.
|
||||
- **Post-silicon:** Microarchitectural validation suite, focused on MUL/DIV-intensive kernels (big-integer arithmetic, hashes, polynomial multiplications).
|
||||
|
||||
**PROPOSAL:** Verification cost should be a first-class input to the MUL/DIV design decision. A simpler, more easily verified design is preferable to a faster, harder-to-verify one, unless the performance delta is necessary for the XH-1 target workload.
|
||||
|
||||
## Software Considerations
|
||||
|
||||
- **Compiler behavior.** **FACT.** GCC and LLVM routinely use shift-and-add sequences for multiplication by small constants, and they may either emit `MUL` instructions or inline expansions depending on the cost model. The cost of a long-latency MUL or DIV may push the compiler toward suboptimal code sequences, but the magnitude of this effect is workload- and compiler-version-dependent and should not be assumed.
|
||||
- **Library code.** **FACT.** libgcc, compiler-rt, and the C library contain hand-written multiplication and division routines (e.g., for `__divti3`, `__udivti3`). These may use the MUL/DIV instructions in tight loops, exposing the unit's throughput directly.
|
||||
- **System software.** **TYPICAL.** Hash functions, AES-GCM polynomial multiplication, modular exponentiation in cryptography, and big-integer arithmetic in language runtimes are all MUL-throughput sensitive.
|
||||
- **128-core interaction.** **ASSUMPTION.** In a homogeneous configuration where the system stack (kernel, hypervisor) runs on some subset of the 128 cores, there is no asymmetric design implication: every core is equal in the default per-core configuration. **INSUFFICIENT EVIDENCE:** Whether the 128-core fabric is homogeneous or heterogeneous (e.g., application cores plus management or I/O cores) is an open question (see Open Questions); the assumption of homogeneity is provisional.
|
||||
|
||||
**INSUFFICIENT EVIDENCE:** XH-1's target software ecosystem is unknown. Bare-metal, Linux, or richer OS support each imply different MUL/DIV usage profiles.
|
||||
|
||||
## Recommendation
|
||||
|
||||
**No final recommendation is made at this time.** The MUL/DIV design is contingent on three decisions that are not yet established in the XH-1 repository:
|
||||
|
||||
1. The core microarchitecture (in-order vs. OoO, issue width, pipeline depth).
|
||||
2. The per-core area and power budget.
|
||||
3. The target workload mix and software ecosystem.
|
||||
|
||||
**Conditional recommendations** (to be revisited once the above are known):
|
||||
|
||||
- **If XH-1 targets an in-order core with tight area budget:** **PROPOSAL.** A separate iterative shift-add MUL (one row/cycle) and an iterative subtractive DIV is the smallest, easiest-to-verify choice. MUL/DIV latency will be high; whether this is acceptable depends on the forwarding path and the workload. **The forwarding path is a real and non-trivial cost:** the bypass network from the MUL/DIV output to dependent consumers (ALU, branch, load/store) must be sized to the full MUL/DIV latency, and the wiring may dominate the area of the iterative MUL/DIV unit itself. This cost should be quantified before the iterative unit is chosen. The shared-iterative-unit approach (Approach 8) is *not* recommended for RV64 by default, given the structural mismatch between MUL and DIV datapaths and the modest area advantage over a split iterative design.
|
||||
- **If XH-1 targets a wide-issue OoO core with relaxed area budget:** **PROPOSAL.** A pipelined radix-4 Booth multiplier (3–6 stage pipeline) paired with a 2-bit-per-cycle naïve divider or a radix-4 SRT divider. MUL throughput at 1/cycle; DIV throughput at 1/16 to 1/32 cycle. **Reconciliation note:** the cross-cutting recommendation below defers SRT until formally proven correct; if SRT cannot be formally verified on reduced widths within the project timeline, the 2-bit-per-cycle naïve divider is the preferred DIV companion. The SRT recommendation is conditional on the verification investment being made.
|
||||
- **If the per-core area budget is moderate and the workload is balanced:** **PROPOSAL.** A pipelined radix-4 MUL with an iterative subtractive DIV (the "split" approach, Alternative A) is a strong compromise.
|
||||
|
||||
**Cross-cutting recommendation:** Defer any design that depends on a non-trivial final correction step (SRT error correction, Newton-Raphson / Goldschmidt rounding fixup) until those algorithms have been formally specified and proven correct on reduced widths.
|
||||
|
||||
## Confidence
|
||||
|
||||
| Topic | Confidence |
|
||||
|---|---|
|
||||
| Taxonomy of MUL/DIV approaches | High (canonical, well-known) |
|
||||
| Relative area, power, and energy rankings | Low to medium (qualitative ordering only; no synthesis data) |
|
||||
| Per-cycle latency numbers | Low to medium (typical, but node- and target-frequency-dependent; ranges are illustrative, not tight bounds) |
|
||||
| Suitability for XH-1 specifically | **Low** (XH-1 microarchitecture not yet established) |
|
||||
| Specific recommendation | **Insufficient evidence** to commit |
|
||||
|
||||
## Open Questions
|
||||
|
||||
1. Is the XH-1 core in-order or out-of-order? Issue width? Pipeline depth? Target frequency?
|
||||
2. What is the per-core area budget, in absolute terms or as a fraction of die area?
|
||||
3. What is the target workload mix (cryptographic, HPC, server, general-purpose, embedded)?
|
||||
4. Will XH-1 implement the "B" (bitmanip) extension, which adds many pseudo-multiplication operations? If so, the MUL/DIV unit may be partially reused or augmented.
|
||||
5. Is the 128-core fabric a true SMP (Linux runs on all 128 cores), or is it a heterogeneous mix (e.g., application cores + management cores)? This affects MUL/DIV uniformity requirements.
|
||||
6. What is the target process node? Node geometry strongly affects the area and timing trade-offs of Booth vs. array multipliers.
|
||||
7. Will the MUL/DIV unit be hardened as a macro and replicated 128 times, or will it be synthesized per-core? Hardening is favored for area/power and replication regularity; synthesis-per-core favors design flexibility.
|
||||
8. Is there an asymmetric verification budget for the MUL/DIV unit, or is verification cost unconstrained?
|
||||
9. Does XH-1 need a 64×64→128-bit MUL for the upper-multiply operations to be single-cycle, or is multi-cycle acceptable?
|
||||
10. What is the expected frequency of `DIV`/`REM` instructions in the target workload? Server workloads may have many fewer than scientific workloads.
|
||||
11. Will the MUL/DIV unit be shared with the FP pipeline (as in BOOM and some SiFive designs, though not in Rocket Chip), or kept private to the integer pipeline?
|
||||
12. Will XH-1 adopt a shared multi-cycle divider across cores (Alternative F), or is a strictly per-core MUL/DIV unit required?
|
||||
|
||||
## Sources
|
||||
|
||||
- Canonical computer-arithmetic references: Parhami, *Computer Arithmetic*; Ercegovac and Lang, *Digital Arithmetic*; Hennessy and Patterson, *Computer Architecture*. These cover the techniques surveyed above in standard form. **Caveat:** these texts underwrite the taxonomy and the mechanism descriptions. Specific per-cycle latency figures given in this document are **TYPICAL** values drawn from common practice in published RV64 designs; the cited texts provide the algorithmic background but do not, in their canonical editions, supply XH-1-specific latency or area numbers. No specific chapter or page is cited for the quantitative figures because the figures are not drawn from a single source.
|
||||
- Open RISC-V core implementations (Rocket, BOOM, XiangShan, SiFive) provide reference designs for radix-4 array multipliers, iterative and SRT dividers, and FP/integer multiplier sharing. These are cited as implementation exemplars, not as XH-1 references. **FACT.** Rocket Chip keeps the integer and FP multipliers as separate units; BOOM and some SiFive designs share partial structures. The blanket attribution of sharing to "Rocket, BOOM, and most SiFive cores" in the prior draft of this document overstates the case and is corrected here.
|
||||
- **INSUFFICIENT EVIDENCE:** No XH-1-internal sources exist for this topic. The repository document `research/03-core-design/mul-div-unit.md` is itself a stub ("SOON"), and all sibling documents in `research/03-core-design/` are stubs. No prior architectural decisions are recorded.
|
||||
- No external citations, papers, benchmarks, or measurements have been invented for this document. All comparative claims are framed as qualitative engineering estimates or as canonical background knowledge in the field of computer arithmetic.
|
||||
+347
File diff suppressed because one or more lines are too long
@@ -0,0 +1,8 @@
|
||||
2026-08-25T18:31:08Z research/03-core-design/mul-div-unit.md 1 research completed
|
||||
2026-08-25T18:31:49Z research/03-core-design/mul-div-unit.md 1 review VERDICT: FAIL
|
||||
2026-08-25T18:32:55Z research/03-core-design/mul-div-unit.md 2 revision completed
|
||||
2026-08-25T18:33:37Z research/03-core-design/mul-div-unit.md 2 review VERDICT: FAIL
|
||||
2026-08-25T18:34:54Z research/03-core-design/mul-div-unit.md 3 revision completed
|
||||
2026-08-25T18:35:35Z research/03-core-design/mul-div-unit.md 3 review VERDICT: FAIL
|
||||
2026-08-25T18:37:17Z research/03-core-design/mul-div-unit.md 4 revision completed
|
||||
2026-08-25T18:38:09Z research/03-core-design/mul-div-unit.md 4 review VERDICT: FAIL
|
||||
@@ -0,0 +1 @@
|
||||
research/03-core-design/mul-div-unit.md
|
||||
+505
@@ -0,0 +1,505 @@
|
||||
# MUL/DIV Unit
|
||||
|
||||
## Status
|
||||
|
||||
Revision 2. Initial scoping document. No XH-1 implementation decisions are yet committed. This revision corrects factual errors identified in review, removes unsupported quantitative claims, reconciles internal contradictions, and adds missing alternatives.
|
||||
|
||||
## Abstract
|
||||
|
||||
This document investigates the design of the integer multiply/divide (MUL/DIV) execution unit for the XH-1 core. The unit must implement the RV64M extension (MUL, MULH, MULHU, MULHSU, DIV, DIVU, REM, REMU, MULW, DIVW, DIVUW, REMW, REMUW) and, depending on project scope, the RV64B bit-manipulation extensions (e.g., CLZ, CTZ, BSET, BEXT, CLMUL, CLMULH, CLMULR). The design must be evaluated against the XH-1's 128-core replication factor: a unit that is only a small percentage of a single core's area consumes a large cumulative area across the die when replicated 128 times. The unit's latency impacts critical-path timing, and its throughput affects overall core IPC under integer-heavy workloads (cryptography, hashing, address arithmetic, HPC kernels).
|
||||
|
||||
## Research Question
|
||||
|
||||
What is the optimal MUL/DIV unit organization for an XH-1 core given that the design will be replicated 128 times, considering latency, throughput, area, power, energy, and verification complexity trade-offs across in-order, out-of-order, and hybrid pipeline styles?
|
||||
|
||||
## Background
|
||||
|
||||
### RISC-V M Extension Semantics (RV64M)
|
||||
|
||||
**FACT** (per RISC-V ISA spec, Volume I, Chapter 7): The RISC-V M extension for RV64 defines the following instructions:
|
||||
|
||||
- MUL, MULH, MULHU, MULHSU: 64×64→128 bit multiply (lower 64 or upper 64 of result, with sign-handling variations).
|
||||
- MULW: 32×32→32 bit product, then sign-extended to 64 bits and written to `rd`.
|
||||
- DIV, DIVU: 64÷64 signed/unsigned quotient.
|
||||
- REM, REMU: 64÷64 signed/unsigned remainder.
|
||||
- DIVW, DIVUW, REMW, REMUW: 32÷32 signed/unsigned quotient/remainder, sign-extended to 64 bits and written to `rd`.
|
||||
|
||||
**FACT** (per RISC-V ISA spec, Volume I, Chapter 7): The M extension defines the following corner-case behavior:
|
||||
|
||||
- Division by zero:
|
||||
- `DIV` / `DIVW`: quotient is `−1` (the architectural definition; the bit pattern is `2^XLEN − 1`, all bits set).
|
||||
- `DIVU` / `DIVUW`: quotient is `2^XLEN − 1` (all bits set, which equals `−1` in two's complement representation).
|
||||
- The signed and unsigned cases produce the same bit pattern at the architectural level; the spec writes `−1` for the signed case and the unsigned maximum for the unsigned case, but these are bit-pattern-identical.
|
||||
- `REM` / `REMW`: remainder equals the dividend.
|
||||
- `REMU` / `REMUW`: remainder equals the dividend.
|
||||
- Signed overflow (most-negative integer divided by −1):
|
||||
- `DIV` / `DIVW`: quotient equals the dividend (i.e., the most-negative representable value `2^(XLEN−1)`).
|
||||
- `REM` / `REMW`: remainder equals zero.
|
||||
- For unsigned divide (`DIVU` / `REMU` / `DIVUW` / `REMUW`), the only defined special case is division by zero; the dividend / `−1` overflow case does not apply because the operands are unsigned.
|
||||
|
||||
**NOTE**: A full 64×64→128-bit multiply requires a true 128-bit result datapath even when only the upper or lower 64 bits are returned.
|
||||
|
||||
**NOTE** (MULHSU implementation): MULHSU computes the upper 64 bits of a signed×unsigned 64×64 product. A correct implementation generates a partial-product array for 64×64 bits where the unsigned operand's partial products are zero in the upper half of its bit positions and the signed operand's partial products use signed (sign-extended) rows in the final reduction. Two concrete implementation paths exist:
|
||||
|
||||
- (a) Use a signed multiplier datapath: sign-extend the signed operand to 128 bits, zero-extend the unsigned operand to 128 bits, and run a signed 128×128 multiply, then take the upper 64 bits. This is straightforward but doubles the multiplier width and is rarely used.
|
||||
- (b) The standard approach: zero-extend the unsigned operand to 64 bits (its bit positions are already non-negative), sign-extend the signed operand only in the final partial-product row, and reduce the 64×64 partial-product array with a final row sign-extension. This requires a modified-Booth or array multiplier with explicit sign handling on the last partial-product row.
|
||||
|
||||
A pure unsigned Wallace/Dadda tree without a sign-handling front-end does not implement MULHSU correctly, because the signed operand's most significant partial-product row must be sign-extended (or its inverted-and-carry form added) into the reduction tree.
|
||||
|
||||
**NOTE**: The W-suffixed instructions (MULW, DIVW, DIVUW, REMW, REMUW) are part of the **M extension** in RV64, not the base I extension. The base I extension's W variants are only the simple ALU ops (ADDW, SUBW, SLLW, SRLW, SRAW).
|
||||
|
||||
**NOTE**: MULW is implementable on a 32×32→64-bit datapath: the 32-bit product is sign-extended to 64 bits and written to `rd`. The upper 32 bits of the 64-bit intermediate result are discarded. A 32×32→64-bit fast multiplier (B3 datapath) can therefore implement MULW by taking its lower 32 bits and sign-extending to 64.
|
||||
|
||||
**NOTE**: The carry-less multiply instructions CLMUL, CLMULH, and CLMULR are part of the standard Zbc extension (commonly grouped under the umbrella "B" extension in some profiling). They are not part of M. They require a different datapath (AND-tree with XOR reduction, no carry propagation) and are not the subject of this document except where they interact with operand muxes / writeback. CLMULR in particular produces a 2·XLEN-bit result with explicit carry-in/carry-out behavior across the two halves; its implementation shares the AND-tree with CLMUL/CLMULH but adds a dedicated reduction stage for the carry path.
|
||||
|
||||
### Latency Reference Points
|
||||
|
||||
**FACT (Rocket Chip, UC Berkeley generator)**: Rocket Chip is in-order, and the MUL/DIV unit is configuration-dependent across Rocket's `Configs.scala` parameter set. In `RocketCoreConfig` (commonly cited as the default), the multiplier is pipelined with multiple pipeline stages (multi-cycle iterative, new operation accepted per cycle) and the divider is a radix-4 iterative divider. The exact stage counts and latencies are determined by parameters in `Configs.scala` and are not a single canonical value. INSUFFICIENT EVIDENCE for a specific latency number without naming the configuration.
|
||||
|
||||
**FACT (BOOM v2/v3, UC Berkeley)**: Out-of-order superscalar. MUL is pipelined (latency configuration-dependent). DIV is variable-latency, non-pipelined.
|
||||
|
||||
**FACT (XiangShan, open-source OoO RISC-V)**: MUL is pipelined; DIV is variable-latency, non-pipelined.
|
||||
|
||||
**FACT (Ibex, lowRISC)**: In-order. MUL is implemented as a single-cycle or short-pipeline combinational multiplier in some configurations; iterative DIV.
|
||||
|
||||
**INSUFFICIENT EVIDENCE**: Arm Cortex-A77 per-instruction MUL/DIV latencies are not publicly published by Arm. No specific numbers are cited from a primary source.
|
||||
|
||||
**INSUFFICIENT EVIDENCE**: Intel Haswell integer divider internal radix (radix-16 vs. radix-32) is not established from publicly verifiable primary sources. The design is widely reported to be a high-radix shift-subtract divider rather than a Newton-Raphson divider, but the specific radix is INSUFFICIENT EVIDENCE.
|
||||
|
||||
### Divide Algorithms
|
||||
|
||||
**FACT**: Four primary classes are considered here:
|
||||
|
||||
1. **Digit-recurrence / shift-subtract (radix-2, radix-4, radix-16, radix-64 SRT)**: radix-2 takes 64 cycles worst case; radix-4 takes ~32–33 cycles; radix-16 takes ~16 cycles; radix-64 takes ~8–10 cycles (with significant area/complexity cost). Iterative, small-to-moderate area depending on radix.
|
||||
2. **Newton-Raphson reciprocal multiplication**: multiple iterations of a multiply-based refinement to compute the reciprocal, then a final correction multiply to produce the quotient. Latency depends on initial seed precision and convergence criteria.
|
||||
3. **Goldschmidt**: similar convergence behavior to Newton-Raphson, with a different iteration structure.
|
||||
4. **CORDIC-based and series-expansion dividers**: rarely used for general-purpose integer divide due to overhead; exist as alternative approaches.
|
||||
|
||||
**FACT**: The RISC-V M extension does not require IEEE-754 compliant division; integer rounding is sufficient, simplifying hardware.
|
||||
|
||||
## Existing Approaches
|
||||
|
||||
### A1. Iterative Shift-Subtract Divider (Radix-2)
|
||||
|
||||
- **Datapath**: 128-bit (64 remainder + 64 divisor) shifter/subtractor, radix-2, 64-cycle worst case.
|
||||
- **DIV Latency**: 64 cycles.
|
||||
- **DIV Throughput**: 1 divide per 64 cycles (not pipelined).
|
||||
- **MUL coverage**: A1 defines a divider only; MUL is not part of this organization.
|
||||
- **Area (DIV only)**: Smallest divider; typically a few kGE plus control. **INSUFFICIENT EVIDENCE** for a specific gate count.
|
||||
- **Power**: Lowest of the divider options when idle; minimal toggle rate per non-dividing cycle.
|
||||
- **Used in**: Some low-end in-order cores; some configurations of Rocket Chip with smaller radix.
|
||||
|
||||
### A2. Pipelined Iterative Divider
|
||||
|
||||
- **Datapath**: Two distinct subclasses must be distinguished:
|
||||
- **(A2a) Pipelined iterative loop**: a single shift-subtract array with pipeline registers inserted at one or more points within the iterative loop, allowing a new operation to enter the loop every cycle after the pipeline is filled. Latency remains 64 cycles; throughput is 1 per cycle after fill.
|
||||
- **(A2b) Fully unrolled divider**: 64 shift-subtract stages with pipeline registers between every stage, giving latency 64 cycles and throughput 1 per cycle from the first cycle. Area is roughly 64× the A1 datapath.
|
||||
- **DIV Latency**: 64 cycles.
|
||||
- **DIV Throughput**: 1 per cycle (after fill for A2a; from cycle 1 for A2b).
|
||||
- **Area**: A2a is ~2–3× A1 (a few extra pipeline registers); A2b is roughly 64× A1 and is rarely used.
|
||||
- **Power**: Higher toggle rate than A1; only worthwhile under sustained divide streams.
|
||||
- **Used in**: Rare; mostly in high-throughput streaming dividers (DSP). Uncommon in general-purpose cores.
|
||||
|
||||
### A3. Newton-Raphson Divider
|
||||
|
||||
- **Datapath**: shared MUL unit(s), initial reciprocal seed ROM, multiplier used in iterative refinement and one final correction multiply.
|
||||
- **Latency breakdown**: seed table lookup (1 cycle) + N refinement multiplies (typically 2–3 iterations for 64-bit integer) + 1 final correction multiply. **ASSUMPTION**: 3 refinement iterations + 1 correction is typical for 64-bit; the concrete iteration count depends on seed precision and convergence criteria. **INSUFFICIENT EVIDENCE** for a single canonical iteration count without specifying the seed table and refinement schedule.
|
||||
- **Throughput**: one divide per (N+1) MUL cycles **only if the multiplier is dedicated to the divider**. If the multiplier is shared with the main MUL datapath, throughput is degraded by contention with MUL issue rate and is workload-dependent. **INSUFFICIENT EVIDENCE** for a single throughput number in the shared case.
|
||||
- **Area**: 1 reciprocal seed ROM + 1–2 MUL units (sharing possible at the cost of contention); large.
|
||||
- **Power**: Higher static and dynamic (multiplier active during divide refinement).
|
||||
- **Used in**: Some high-performance FPU designs for floating-point; less common for dedicated integer divide.
|
||||
|
||||
### A4. Radix-16 / Radix-64 SRT Divider
|
||||
|
||||
- **Datapath**: high-radix recurrence with a quotient-digit lookup table (PLA or ROM), redundant remainder representation. 16 or 64 bits processed per cycle.
|
||||
- **DIV Latency**: ~16 cycles (radix-16) or ~8–10 cycles (radix-64) for 64-bit operands.
|
||||
- **DIV Throughput**: 1 per cycle if pipelined; 1 per 16 / 8–10 cycles if iterative.
|
||||
- **Area**: large lookup table and complex datapath; PLA is a significant area contributor.
|
||||
- **Power**: high; many bits toggle per cycle.
|
||||
- **Used in**: high-end x86 integer dividers (radix of the integer divider is **INSUFFICIENT EVIDENCE** from primary sources; widely reported as high-radix shift-subtract rather than Newton-Raphson).
|
||||
|
||||
### A5. Approximate / Lookup-Based Dividers (small operand ranges)
|
||||
|
||||
- Rarely used for 64-bit; not relevant to XH-1 unless a specific workload (e.g., 32-bit address arithmetic) dominates.
|
||||
|
||||
## Alternative Designs
|
||||
|
||||
### B1. Hybrid MUL + Sequential-Iterative DIV (radix-4 iterative divider, custom MUL)
|
||||
|
||||
- MUL: 64×64→128 pipelined in 1 stage (target), throughput 1 per cycle. This departs from Rocket Chip's default multi-stage MUL (Rocket's `RocketCoreConfig` issues MUL as a multi-cycle iterative operation); the 1-stage MUL is an aggressive target for XH-1.
|
||||
- DIV: radix-4 iterative, ~33 cycles (32 cycles for quotient bits + finalization), blocking on the unit.
|
||||
- Single divider per core, 1 MUL pipeline stage.
|
||||
|
||||
**NOTE on naming**: B1 inherits the radix-4 iterative divider pattern from Rocket and Ibex, but the MUL organization (1-stage pipelined) is a custom choice that does not match either Rocket's multi-stage MUL or Ibex's short-pipeline / combinational MUL. The "hybrid" descriptor refers to combining a pipelined MUL with an iterative DIV; the divider side aligns with reference designs, the MUL side does not.
|
||||
|
||||
### B1'. Two MUL Pipelines + Shared Iterative DIV (wide-issue variant)
|
||||
|
||||
- Two pipelined MUL datapaths, one shared radix-4 DIV datapath.
|
||||
- MUL throughput: 2/cycle (sustained, independent operands).
|
||||
- DIV throughput: 1 per ~33 cycles, shared and blocking.
|
||||
- **Area**: ASSUMING a MUL:DIV area ratio of approximately 1:1 (i.e., one MUL pipeline and one radix-4 iterative DIV are roughly comparable in area, since a 1-stage 64×64 MUL is a Wallace/Dadda tree plus a 128-bit CPA, and a radix-4 iterative DIV is a ~66-bit adder plus control state), the combined (MUL+DIV) area scales as 1 + 1 = 2× B1 (two MULs plus one shared DIV, where B1 has one MUL and one DIV). The 1.5× figure previously cited assumed MUL is half the area of DIV, which is unsupported. The 1.4× lower bound previously cited is unsupported. The corrected qualitative statement is: **~2× B1**, with the explicit MUL:DIV area ratio assumption stated. **INSUFFICIENT EVIDENCE** for a synthesis-derived number.
|
||||
|
||||
### B2. Combined MUL/MAC (Multiply-Accumulate) Unit
|
||||
|
||||
- 64×64→128 with an additional accumulator register; supports future RV MAC extensions (proposed in some bitmanip drafts).
|
||||
- Adds area over pure MUL; **INSUFFICIENT EVIDENCE** for a specific percentage.
|
||||
- Useful for DSP, ML, and HPC; 128-core replication benefits workloads with matrix-style kernels.
|
||||
|
||||
### B3. Operand-Width-Detected 32-bit Fast MUL Path (also serves MULW)
|
||||
|
||||
- Detect when both operands of a 64-bit MUL are sign- or zero-extended from 32 bits (i.e., bit 31 is replicated through bit 63), and route the multiplication through a 32×32→64 fast multiplier.
|
||||
- The same 32×32→64 datapath also implements MULW directly: the 32-bit product (lower 32 bits of the 64-bit result) is sign-extended to 64 bits and written to `rd`. This sharing is essentially free in area terms and means that adopting B3 is the natural way to implement MULW.
|
||||
- This is **distinct from MULW as a workaround** for the 32-bit-case: B3 accelerates 64-bit MUL / MULH / MULHSU / MULHU when both operands happen to be 32-bit sign- or zero-extended, a pattern common after `lw` / `lwu` followed by arithmetic. MULW alone does not accelerate this case because MULW is a different instruction with different result semantics.
|
||||
- **Position**: B3 is the natural choice for the MULW datapath and adds a 32-bit-extended-operand fast path for 64-bit MUL. The verification cost is the dual datapath (32-bit and 64-bit) and the operand-width classifier. At the 128-core replication level, the area overhead is bounded by the 32-bit datapath size, which is much smaller than the 64-bit datapath; the dominant question is verification cost, not area.
|
||||
- **Decision**: B3 is recommended as part of the MULW implementation; deferring B3 means deferring the natural MULW implementation, which is not viable if MULW is in the ISA. The B3-vs-B4 trade-off therefore applies to the 32-bit-extended-operand fast path, not to MULW support.
|
||||
|
||||
### B4. Bypassable Output with Operand Width Detection (full 64×64→128 only)
|
||||
|
||||
- Always implement full 64×64→128; the decoder examines the instruction to forward only the required 64 bits.
|
||||
- Verification: a single source of truth (full 128-bit datapath), reducing dual-mode bugs.
|
||||
- Disadvantage: no area saving; MULW must still be implemented separately (e.g., via B3 or a dedicated 32×32→32 path).
|
||||
- **Position**: B4 conflicts with the natural MULW-via-B3 sharing above; if MULW is required, B4 is not a complete solution. If MULW is not required, B4 simplifies verification at the cost of a slightly larger MUL datapath for MULW-equivalent work.
|
||||
|
||||
### B5. Skip-on-Zero / Divide-Cancellation Optimizations
|
||||
|
||||
- Detect zero dividend (quotient is zero, remainder is dividend) and divide-by-one (quotient is dividend, remainder is zero) at the front end and forward the result without entering the iterative loop.
|
||||
- **Corner cases the fast path must handle correctly** (consistent with the iterative path):
|
||||
- dividend = 0, divisor = anything (including 0): quotient = 0, remainder = 0 (dividend).
|
||||
- divisor = 1: quotient = dividend, remainder = 0.
|
||||
- divisor = −1:
|
||||
- dividend = 2^(XLEN−1) (most-negative): quotient = 2^(XLEN−1) (dividend), remainder = 0 (signed overflow case, the M-extension rule).
|
||||
- all other dividends: quotient = −dividend (two's complement negation), remainder = 0.
|
||||
- The fast path must explicitly implement these rules; it is not a simple "if divisor=±1 then return dividend" because of the signed-overflow case for divisor = −1.
|
||||
- Saves latency in the common case for some workloads; trivial area overhead.
|
||||
- Composable with any divider organization.
|
||||
|
||||
### B6. Software Divide-by-Constant Transformation (cross-cutting)
|
||||
|
||||
- Compilers transform division by a **compile-time** constant into a multiply-by-reciprocal sequence. The hardware DIV is then needed only for division by variables.
|
||||
- Reduces effective DIV frequency significantly for workloads with constant denominators; affects hardware sizing decisions. **INSUFFICIENT EVIDENCE** for a quantitative reduction without a specific workload profile.
|
||||
|
||||
### B7. Dedicated 32-bit Fast Divider for W-suffixed Instructions
|
||||
|
||||
- Implement a separate 32-bit radix-2 or radix-4 iterative divider for DIVW / DIVUW / REMW / REMUW. The 32-bit divider has half the iteration count (32 or 16 cycles vs. 64 or 33 for the 64-bit divider) and roughly a quarter of the datapath area.
|
||||
- Useful if profiling shows W-suffixed divides dominate; in most general-purpose workloads they do not.
|
||||
- Verification cost: dual divider datapath, similar to B3.
|
||||
- **Note**: W-suffixed instructions operate on 32-bit operands; a natural alternative is to share the 64-bit divider datapath with the 32-bit operands on the lower 32 bits, taking 32 or 16 cycles of the 64-bit divider's iteration. B7 is only worth its area if the latency savings matter.
|
||||
|
||||
### B8. Combined MUL / DIV with Shared Partial-Product Array (CSA Sharing)
|
||||
|
||||
- Reuse the MUL's CSA compressor tree as the final correction multiplier for an SRT or Newton-Raphson divider, sharing the most area-intensive block.
|
||||
- Reduces the area penalty of A3 / A4 at the cost of tighter verification coupling between MUL and DIV paths.
|
||||
- Real architectural option in some high-performance designs; not considered in the prior revision.
|
||||
|
||||
### B9. Combined MUL / DIV with Shared Final Carry-Propagate Adder (CPA Sharing)
|
||||
|
||||
- Share only the final 128-bit carry-propagate adder between the MUL datapath and the DIV's correction-multiply step (or the DIV's final-cycle remainder correction). The compressor tree, partial-product generation, and divider iteration state remain separate.
|
||||
- Lower area savings than B8 (CSA is shared instead of CPA), but much simpler verification: the shared CPA is a single combinational block, and the MUL vs DIV datapaths feeding it are independent.
|
||||
- The MUL pipeline register naturally sits between the compressor-tree output and the CPA, which means the CPA itself can be shared at the output side without disrupting the MUL pipeline structure.
|
||||
- **Open question**: whether the MUL pipeline register sits at the compressor-tree output (CPA in cycle 2) or at the CPA output (CPA in cycle 1) determines the CPA's pipeline stage. See Implementation Considerations for the canonical placement decision.
|
||||
|
||||
### B10. Partially Unrolled Radix-4 Divider (intermediate between A2a and A2b)
|
||||
|
||||
- Unroll the radix-4 iterative divider into a small number of pipeline stages (e.g., 8 or 16 stages) rather than the full 64 (A2b) or the single iterative loop (A2a).
|
||||
- Latency 32 or 33 cycles (one stage per two quotient bits for 8 stages, or one stage per quotient bit for 16 stages), throughput 1 per cycle after fill, area roughly 8× or 16× A1.
|
||||
- A meaningful intermediate option for the high-DIV-throughput case that the prior revision did not consider.
|
||||
- Verification: same as A2a (iterative datapath with pipeline registers); the unrolling does not introduce new corner cases.
|
||||
|
||||
## Comparison
|
||||
|
||||
| Design | MUL Latency | MUL Tput | DIV Latency | DIV Tput | Rel. Area (MUL+DIV) | Power (active) | Verification Complexity |
|
||||
|--------|-------------|----------|-------------|----------|----------------------|----------------|--------------------------|
|
||||
| A1 (radix-2 iter. DIV only) | not defined in A1 | not defined in A1 | 64 cycles | 1 per 64 cycles | INSUFFICIENT EVIDENCE (combined) | Lowest (DIV only) | Low |
|
||||
| A2a (pipelined iter. loop) | not defined in A2a | not defined in A2a | 64 cycles | 1 per cycle (after fill) | INSUFFICIENT EVIDENCE | Med | Med |
|
||||
| A2b (fully unrolled) | not defined in A2b | not defined in A2b | 64 cycles | 1 per cycle | INSUFFICIENT EVIDENCE (very large) | High | High |
|
||||
| A3 (Newton-Raphson) | 1 cycle | 1 per cycle (dedicated) | N+1 MUL cycles (dedicated); workload-dep. if shared | INSUFFICIENT EVIDENCE (shared) | high | High | High |
|
||||
| A4 (Radix-16/64 SRT) | not defined in A4 | not defined in A4 | 8–16 cycles | 1 per cycle if pipelined | high | High | High |
|
||||
| B1 (radix-4 DIV, 1-stage MUL) | 1 cycle (target) | 1 per cycle | ~33 cycles | 1 per 33 cycles | baseline | Low–Med | Low–Med |
|
||||
| B1' (2× MUL + 1× shared DIV) | 1 cycle | 2 per cycle | ~33 cycles | 1 per 33 cycles (shared) | ~2× B1 (INSUFFICIENT EVIDENCE; assumes MUL:DIV area ≈ 1:1) | Med | Med |
|
||||
| B2 (+ MAC) | 1 cycle | 1 per cycle | ~33 cycles | 1 per 33 cycles | +unspecified % over B1 (INSUFFICIENT EVIDENCE) | Med | Med |
|
||||
| B3 (32-bit fast MUL; serves MULW) | 1 cycle (32-bit path) | 1 per cycle | ~33 cycles | 1 per 33 cycles | +small (INSUFFICIENT EVIDENCE) | Low | Med (dual mode) |
|
||||
| B4 (full 64 only; MULW separate) | 1 cycle | 1 per cycle | ~33 cycles | 1 per 33 cycles | same as B1 (MULW path TBD) | Med | Lowest (MUL side); Med (MULW) |
|
||||
| B5 (skip-on-zero) | n/a | n/a | reduced in common case | same as base | negligible overhead | n/a | Low |
|
||||
| B7 (32-bit fast DIV) | 1 cycle | 1 per cycle | ~16–17 cycles (32-bit) | 1 per 16–17 cycles (32-bit only) | +small (INSUFFICIENT EVIDENCE) | Low | Med (dual mode) |
|
||||
| B8 (shared MUL/DIV CSA) | 1 cycle | 1 per cycle | A3/A4 latency | A3/A4 throughput | lower than A3/A4 alone (INSUFFICIENT EVIDENCE) | Med–High | Med–High (tight coupling) |
|
||||
| B9 (shared MUL/DIV CPA) | 1 cycle | 1 per cycle | A3/A4 latency | A3/A4 throughput | slightly higher than B8 reduction (INSUFFICIENT EVIDENCE) | Med | Low (clean separation) |
|
||||
| B10 (partially unrolled radix-4) | not defined in B10 | not defined in B10 | ~32–33 cycles | 1 per cycle (after fill) | ~8–16× A1 (INSUFFICIENT EVIDENCE) | Med | Med |
|
||||
|
||||
**ASSUMPTION**: Relative area figures are qualitative orderings based on published reference designs cited in the Sources section. Actual XH-1 synthesis numbers are **INSUFFICIENT EVIDENCE** pending RTL implementation.
|
||||
|
||||
## Advantages
|
||||
|
||||
### A1 / B1
|
||||
- Smallest area; lowest per-core replication cost.
|
||||
- Lowest static and dynamic power in the typical case (most cycles no MUL/DIV active).
|
||||
- Simpler verification; one mode of operation.
|
||||
- Well-understood reference implementation (Rocket, Ibex) for the divider side; the MUL side departs from both references.
|
||||
|
||||
### A4 (High-Radix SRT)
|
||||
- Lowest DIV latency among the iterative-style options; competitive with A3 on a single divide.
|
||||
- Pipelined variant gives 1 per cycle DIV throughput.
|
||||
|
||||
### A3 (Newton-Raphson)
|
||||
- Low latency for sequential divides — important for cryptography (RSA, ECC modular reduction) and HPC.
|
||||
- Amortizes multiplier cost if MAC (B2) is also desired.
|
||||
- Disadvantage: requires multiple refinement iterations; integer multiplier is large and contention with the main MUL datapath is a concern.
|
||||
|
||||
### B2 (MAC)
|
||||
- Enables future-proofing for proposed bitmanip and MAC extensions.
|
||||
- Helpful for matrix multiplication kernels running across 128 cores.
|
||||
|
||||
### B3 (32-bit fast MUL; serves MULW)
|
||||
- Natural implementation of MULW; the 32×32→64 datapath produces the MULW result by sign-extending the lower 32 bits.
|
||||
- Accelerates 64-bit MUL on 32-bit-valued operands (a common pattern after `lw`/`lwu` + arithmetic).
|
||||
- Area overhead is bounded by the 32-bit datapath size (much smaller than the 64-bit datapath).
|
||||
|
||||
### B5 (Skip-on-Zero)
|
||||
- Negligible area; reduces effective DIV latency for common cases.
|
||||
|
||||
### B7 (32-bit fast DIV)
|
||||
- Halves the divider iteration count for the W-suffixed instructions at modest area cost.
|
||||
|
||||
### B8 (Shared MUL/DIV CSA)
|
||||
- Reduces the area penalty of high-performance dividers by sharing the most area-intensive block.
|
||||
|
||||
### B9 (Shared MUL/DIV CPA)
|
||||
- Lower verification cost than B8; the shared CPA is a single combinational block and the MUL/DIV datapaths feeding it are independent.
|
||||
|
||||
### B10 (Partially Unrolled Radix-4)
|
||||
- A meaningful intermediate option for high DIV throughput without the area cost of A2b; same verification profile as A2a.
|
||||
|
||||
## Disadvantages
|
||||
|
||||
### A1 / B1
|
||||
- Variable DIV latency complicates scheduling in an in-order core; in-order cores stall ~33 cycles per divide.
|
||||
- Worst-case latency hits interrupt latency, branch-target computation if the MUL/DIV is on a branch path.
|
||||
- Not a fit for sustained divide-bound workloads (e.g., modulo arithmetic in bignum).
|
||||
|
||||
### A3
|
||||
- Area at 128-core replication is severe; the multiplier is one of the largest blocks in a typical core.
|
||||
- Power: a 64×64 multiplier running 1 per cycle is one of the highest-power blocks in a typical core.
|
||||
- Verification: Newton iteration precision (must converge to within 1 ULP across all operands) is notoriously subtle.
|
||||
- If multiplier is shared with main MUL datapath, throughput is workload-dependent, not the N+1 figure cited for the dedicated case.
|
||||
|
||||
### A4
|
||||
- Large lookup table (PLA or ROM); area and power dominated by the table.
|
||||
- Verification: complex quotient-digit selection logic.
|
||||
- Not commonly used outside high-end commercial designs.
|
||||
|
||||
### B3
|
||||
- Verification: dual datapath (32-bit and 64-bit) must be exhaustively cross-checked.
|
||||
- The classification logic (operand-width detection) is itself a source of corner-case bugs.
|
||||
- **B3 does not subsume MULW for the purpose of "B3 is redundant"**: B3 and MULW are two different ways to access 32-bit-multiplication, and B3's value is primarily as the natural MULW implementation and secondarily as the 32-bit-extended-operand fast path.
|
||||
|
||||
### B2
|
||||
- Adds a third writeback port or accumulator path; competes with the FPU/LSU for writeback slots in a multi-issue core.
|
||||
|
||||
### B5
|
||||
- Only helps the specific cases of zero dividend or divisor ±1; other optimizations (e.g., division by small powers of two) are already handled by the base I extension's shift instructions.
|
||||
- The fast path must explicitly handle the signed-overflow corner case (dividend = 2^(XLEN−1), divisor = −1): quotient = dividend, remainder = 0.
|
||||
|
||||
### B7
|
||||
- Verification: dual divider datapath; same concerns as B3.
|
||||
- Area saving is moot if W-suffixed divides are not on the critical path.
|
||||
|
||||
### B8
|
||||
- Verification: tighter coupling between MUL and DIV paths makes corner-case analysis more difficult.
|
||||
|
||||
### B9
|
||||
- Area savings are smaller than B8 (CPA shared instead of CSA); the savings are bounded by the CPA size, which is significant but not the dominant block.
|
||||
|
||||
### B10
|
||||
- Area scales with the unroll factor; 16-stage unroll is ~16× A1, which is meaningful at 128-core replication.
|
||||
|
||||
## XH-1 Considerations
|
||||
|
||||
**PROPOSAL**: The XH-1 core should adopt **B1 (Hybrid MUL + Iterative DIV)** as the baseline, with the following parameters, pending RTL validation:
|
||||
|
||||
- **MUL**: 64×64→128, target 1-cycle pipelined (1 stage of pipeline registers), throughput 1 per cycle. The pipeline register sits at the **CPA output** (after the compressor tree and the 128-bit CPA in cycle 1), so the entire compressor-tree-plus-CPA path is in one cycle. This is the "1-cycle MUL" interpretation; the fallback-A interpretation splits the compressor tree and the CPA across two cycles. For 2-issue or wider cores, scale to B1' (two MUL pipelines sharing one DIV).
|
||||
- **DIV**: Radix-4 shift-subtract, ~33 cycles worst case (32 cycles for quotient bits + finalization), blocking, non-pipelined. The prior revision's "33–35 cycles worst case" and the "33–64 cycle" range are reconciled here: 33 is the radix-4 bound; 64 corresponds to radix-2, which is a different algorithm choice.
|
||||
- **REM**: Reuse the DIV datapath; remainder is a by-product.
|
||||
- **Bypass paths**: From MUL output directly to subsequent dependent ALU/branch in the next cycle. **INSUFFICIENT EVIDENCE** on whether this fits the XH-1 pipeline depth; depends on the integration context.
|
||||
- **B5 (skip-on-zero)**: implement at the front end of the DIV unit; negligible overhead. The fast path must handle the signed-overflow corner case (divisor = −1, dividend = 2^(XLEN−1)) by returning quotient = dividend, remainder = 0.
|
||||
- **B3 (32-bit fast MUL, also serves MULW)**: include the 32×32→64 datapath as the MULW implementation path. The same datapath accelerates 64-bit MUL on 32-bit-extended operands. Verification cost: dual datapath, manageable with the B3 corner-case set (operand-width detection, sign-extension patterns).
|
||||
- **B7 (32-bit fast DIV)**: Defer; the W-suffixed divide is not assumed to be on the critical path. Revisit if profiling shows otherwise.
|
||||
|
||||
**ASSUMPTION**: A 1-cycle MUL latency is achievable in the target process. **INSUFFICIENT EVIDENCE** on the XH-1 target process node and clock period. A full 64×64→128 Wallace/Dadda tree + 128-bit carry-propagate adder in a single cycle is at the edge of feasibility for high-performance designs; typical in-order cores implement MUL as either a multi-cycle iterative multiplier or a multi-stage pipelined multiplier. **PROPOSAL**: Validate via synthesis at the target corner before committing. If the 1-cycle critical path cannot be closed, **fallback options** are:
|
||||
- **B1-fallback-A**: 2-cycle pipelined MUL. The pipeline register sits at the **compressor-tree output** (splitting the compressor tree in cycle 1 from the CPA in cycle 2); latency 2 cycles, throughput 1 per cycle, modest area overhead (one extra pipeline register).
|
||||
- **B1-fallback-B**: Multi-cycle iterative MUL (Booth-encoded). The cycle count for a Booth-encoded iterative multiplier on 64×64 is implementation-dependent; **INSUFFICIENT EVIDENCE** for a specific number of cycles. Lower area, lower throughput, higher latency.
|
||||
- **B1-fallback-C**: Retain the 1-stage MUL architecture but lower the target clock frequency (system-level decision, not unit-level).
|
||||
|
||||
**MUL/DIV resource sharing and contention (single-issue-lane B1)**: Under B1, the MUL pipeline and the DIV datapath share the integer execution lane's issue slot, register-file read ports, and writeback port. The single execution lane can issue either one MUL per cycle (latency 1) or one DIV (latency ~33) at a time, but not both simultaneously. If a MUL is issued while a DIV is in progress:
|
||||
- The MUL occupies the issue slot for 1 cycle; the DIV's iterative state is held in the divider's internal registers and does not require the issue slot during its 33 cycles.
|
||||
- Register-file read ports: MUL requires 2 read ports for its 1 cycle; the DIV's operands are read at DIV issue and held in the divider's operand register. No contention after issue.
|
||||
- Writeback port: MUL writes back in cycle 2 (1-cycle latency); the DIV writes back on completion. A MUL issued in the same cycle as a DIV completion would contend for the writeback port. In a single-writeback-port lane, the DIV completion must be stalled by 1 cycle to let the MUL writeback, or vice versa. This is a 1-cycle throughput loss in the rare case of simultaneous MUL-and-DIV-completion, and is acceptable at 128-core scale (per-core throughput loss is small; aggregate is bounded by the per-core lane).
|
||||
- **For 128-core aggregate throughput**: the per-core limitation is the MUL/DIV issue slot, not the divider's iteration. Aggregate MUL throughput is bounded by 1/cycle/core = 128/cycle die-wide. Aggregate DIV throughput is bounded by 1/33 cycles/core = 128/33 ≈ 3.88 divides/cycle die-wide in the steady state if every core is issuing back-to-back independent divides. This is the steady-state upper bound under B1, not a typical workload figure. With B5 reducing some divides to 1 cycle, the aggregate is workload-dependent and typically lower.
|
||||
|
||||
**PROPOSAL**: The MUL/DIV unit sits on the integer execution lane, sharing the issue queue with ALU and branch. It has its own reservation station (or issue-slot tag) so the issue logic can distinguish short-latency ALU (1 cycle) from long-latency MUL (1, or 2 in the fallback) and variable-latency DIV (~33).
|
||||
|
||||
**PROPOSAL**: For the in-order case, the MUL/DIV unit is non-blocking on MUL (1-cycle latency) but blocking on DIV; the in-order front-end stalls on a structural hazard only when a second DIV is issued before the first completes.
|
||||
|
||||
**PROPOSAL**: For an out-of-order XH-1 core, the MUL/DIV unit exposes its variable DIV latency through the wakeup/select logic so that dependent instructions are replayed or re-issued with the correct ready signal.
|
||||
|
||||
**PROPOSAL**: The B-extension unit (if RV64B is implemented) shares **operand muxes, sign-handling logic, and bit-level muxes** with the MUL/DIV unit but does **not** share the Wallace/Dadda compressor tree. Operations like CLZ, CTZ, BSET, BEXT operate on individual bits or small bit-fields and do not naturally map onto a Wallace-tree multiplier datapath. CLMUL / CLMULH / CLMULR (Zbc) require an AND-tree / XOR-reduction datapath that is structurally distinct from both the Wallace-tree multiplier and the iterative divider; they do not share the compressor tree. Any apparent sharing is at the operand-fetch and writeback layers, not the core arithmetic.
|
||||
|
||||
## 128-Core Scalability
|
||||
|
||||
The 128-core replication factor is the dominant cost driver for the MUL/DIV unit.
|
||||
|
||||
**PROPOSAL**: At 128 cores, every MUL unit must be **power-gateable independently** so that cores not in use (DPM, dark-silicon) can be fully clock- and power-gated, including the MUL/DIV block.
|
||||
|
||||
**PROPOSAL**: The MUL/DIV unit's reservation-station entries, divider iteration state, and pipeline registers must support **state retention or clean-state entry** across power gating. Specifically, when a core is power-gated while a long-latency DIV is in flight, the DIV's mid-iteration state must be handled by one of the following:
|
||||
- (a) **Flush and re-issue**: the in-flight DIV is squashed architecturally (the issue queue entry is marked invalid, the divider's iteration state is discarded), and the DIV is re-fetched and re-issued from the I-cache after the core wakes up. This is architecturally transparent if the re-fetch / re-issue mechanism is present (standard OoO replay path); the architectural state is preserved because the DIV is re-executed from scratch. The cost is re-fetch latency after wakeup. The mechanism is "not acceptable" only if the re-fetch path is not implemented (e.g., a simple in-order core without replay).
|
||||
- (b) **Checkpoint to retention**: the divider's iteration state is saved to a retention register or to memory before power-down, and restored on wakeup. Preserves the in-flight DIV but requires retention storage proportional to the divider's state.
|
||||
- (c) **Block power-gating until completion**: power-gating is only allowed when the divider is idle. Simple, but defeats the purpose of DPM if DIV latency is long and frequent.
|
||||
- The choice depends on the XH-1 DPM policy and on whether the core is in-order or OoO. **INSUFFICIENT EVIDENCE** on the XH-1 DPM policy; the design must accommodate one of these options without committing to a specific approach here.
|
||||
|
||||
**OPEN QUESTION**: What is the actual MUL/DIV area share of the XH-1 core? Published RISC-V references suggest a typical MUL unit occupies a small single-digit percentage of a high-performance core's area, with iterative DIV adding additional area. The prior revision's "1–3%" and "1 MGE/core total" baseline are removed here as unsourced. **INSUFFICIENT EVIDENCE** on the XH-1-specific area share without synthesis.
|
||||
|
||||
**OPEN QUESTION**: Is the MUL/DIV area share dominated by the MUL datapath, the DIV datapath, the operand muxing, or the bypass network? Targeted synthesis required.
|
||||
|
||||
**INSUFFICIENT EVIDENCE**: The XH-1 die-area budget, process node, and clock period are not established in this document.
|
||||
|
||||
**PROPOSAL**: The MUL/DIV unit should be **identical across all 128 cores** (no per-core specialization) to simplify verification and physical design. The replication should be a Verilog/SystemVerilog generate block driven by a single template. **PROPOSAL**: The generate block itself must be included in the verification scope (not assumed trivially correct). At minimum: a lint-clean check, a synthesis-check that the generate block instantiates the correct number of cores, and a per-instance equivalence check on a sample of cores.
|
||||
|
||||
**PROPOSAL**: For 128× replicated MUL/DIV pipeline registers, ECC or parity protection should be considered for soft-error mitigation. **INSUFFICIENT EVIDENCE** on the XH-1 reliability target.
|
||||
|
||||
**PROPOSAL**: Reset distribution and scan chain architecture for 128× replicated MUL/DIV must be addressed at the integration level. The MUL/DIV unit's scan chains should support parallel or staggered scan-shift across cores to keep test time bounded. **INSUFFICIENT EVIDENCE** on the XH-1 DFT architecture.
|
||||
|
||||
**CONSIDERATION (clock and timing at 128× replication)**: If the XH-1 die includes a global clock-distribution network, the MUL/DIV critical path determines the local clock skew tolerance. A pipelined 1-cycle MUL has a short critical path (one 64-bit adder-equivalent) which is favorable; a non-pipelined iterative divider has a 64-bit adder in its loop, which is the cycle-time limiter. The clock-skew analysis should be revisited at the integration level once the XH-1 clock tree is defined.
|
||||
|
||||
**CONSIDERATION (cross-core aggregate throughput)**: Under B1, each core's blocking DIV delivers at most 1 per 33 cycles per core in the steady state. Across 128 cores, the aggregate is at most ~3.88 divides per cycle in the steady state, which assumes all cores are issuing back-to-back independent divides for the full 33 cycles each. This is an upper bound on aggregate throughput, not a typical workload figure. Real workloads do not exhibit this worst case; the relevant metric is the per-core latency, not aggregate. **INSUFFICIENT EVIDENCE** on whether the XH-1 DPM or interconnect imposes a global cap on simultaneous divide activity; this is a system-level question outside the MUL/DIV unit's scope.
|
||||
|
||||
**CONSIDERATION (operand distribution and interconnect)**: Replicating a Wallace tree 128× implies 128 sets of wide operand buses to/from the register file. The interconnect / operand-routing network cost scales with the MUL operand width and the number of cores. **PROPOSAL**: include operand-routing overhead in the area estimate, not just the MUL/DIV datapath itself.
|
||||
|
||||
## Performance Considerations
|
||||
|
||||
**PROPOSAL**: **MUL throughput of 1 per cycle is a target for a high-performance XH-1 core**, but is not architecturally non-negotiable. Low-end in-order cores (some Ibex configurations) implement MUL with throughput < 1 per cycle. B1 satisfies the high-performance target with a 1-stage pipelined multiplier; B1' extends to 2 per cycle for wide-issue. If the 1-cycle MUL cannot be closed at the target process, the throughput target remains 1 per cycle but the latency becomes 2 cycles (B1-fallback-A).
|
||||
|
||||
**PROPOSAL**: **DIV throughput of 1 per ~33 cycles (blocking) is acceptable** for most integer workloads. Workloads that require more (e.g., modular arithmetic in cryptography) will suffer; mitigation is left to software (see Software Considerations).
|
||||
|
||||
**ASSUMPTION**: XH-1 target workloads include a mix consistent with Embench / SPECint-class profiles, where MUL/DIV instructions are a small fraction of dynamic instruction count. **INSUFFICIENT EVIDENCE** on the actual XH-1 target workload mix and on the specific dynamic-instruction share of MUL/DIV. The prior revision's "<2% MUL/DIV" and "<0.5% DIV/REM" claims are removed as unsourced.
|
||||
|
||||
**OPEN QUESTION**: Does XH-1 target HPC or cryptography workloads where MUL/DIV is a larger share of the instruction mix? If yes, the analysis shifts toward A3, A4, A2a, or B10. **NOTE**: ML workloads are dominated by floating-point multiplies on the FPU, not by integer MUL/DIV; integer MUL/DIV is relevant to ML only for quantization, address arithmetic, and integer embeddings.
|
||||
|
||||
## Area Considerations
|
||||
|
||||
**PROPOSAL**: Budget the MUL/DIV unit at a small single-digit percentage of single-core area for the B1 design, pending synthesis. The exact percentage is **INSUFFICIENT EVIDENCE**.
|
||||
|
||||
**OPEN QUESTION**: Is the XH-1 core area budget (without MUL/DIV) known? The MUL/DIV area share can be re-expressed as a percentage of the entire 128-core die only when this is established.
|
||||
|
||||
**ASSUMPTION**: A1-style radix-2 DIV is the smallest practical divider; A3 (Newton) and A4 (high-radix SRT) are larger, with A4 typically the largest. **INSUFFICIENT EVIDENCE** on the XH-1-specific gate-count budget.
|
||||
|
||||
## Power and Energy Considerations
|
||||
|
||||
**FACT**: A 64×64→128 multiplier toggles a large number of bits per cycle when active; its dynamic power is proportional to operand activity factor. Power-gating the multiplier when idle is essential.
|
||||
|
||||
**PROPOSAL**: The MUL/DIV unit should support **fine-grained clock gating** at the operand-mux boundary so that when no MUL/DIV is in flight, the entire unit is clock-gated to 0 toggle rate.
|
||||
|
||||
**PROPOSAL**: The MUL pipeline register should be a clock-gated scan flop, not a latch, to simplify DFT.
|
||||
|
||||
**OPEN QUESTION**: Does the XH-1 power budget allow all 128 MUL/DIV units to be active simultaneously? If not, per-core DPM must enforce that no more than N MUL/DIV units are in active divide at once. This is a system-level power-management question, not a unit-level one, but it constrains the unit's design.
|
||||
|
||||
**INSUFFICIENT EVIDENCE**: Specific per-MUL or per-DIV energy numbers for the XH-1 process are not established. The prior revision's "single-digit pJ in 7 nm" claim is removed as unsourced; per-MUL energy in advanced processes is implementation-dependent and varies by an order of magnitude or more depending on architecture and clock frequency.
|
||||
|
||||
## Implementation Considerations
|
||||
|
||||
**PROPOSAL**: Implement MUL as a **Wallace/Dadda tree of 4:2 compressors** feeding a final 128-bit carry-propagate adder. The 1-cycle latency budget accommodates the entire critical path from operand register → compressor tree → CPA → output register, with the pipeline register at the CPA output. **INSUFFICIENT EVIDENCE** on whether a Wallace/Dadda tree for 64×64 partial products is achievable in one cycle at the XH-1 target clock period; the specific tree depth and CPA depth are implementation-dependent and not asserted as fixed numbers here. The prior revision's "Tree depth 6–7; final CPA ~6 gates deep" is removed as unsourced and implementation-specific.
|
||||
|
||||
**MUL pipeline register placement** (clarified):
|
||||
- **Primary (1-stage)**: register at CPA output. The compressor tree and the 128-bit CPA are both in cycle 1.
|
||||
- **Fallback A (2-stage)**: register at compressor-tree output. The compressor tree is in cycle 1, the 128-bit CPA is in cycle 2.
|
||||
- The placement determines the critical path per cycle: in the primary, the per-cycle critical path is tree + CPA; in fallback A, the per-cycle critical path is max(tree, CPA). Fallback A is the natural way to close timing if tree + CPA exceeds the target clock period in a single cycle.
|
||||
|
||||
**PROPOSAL**: For MULHSU, the standard implementation generates a 64×64 partial-product array with the unsigned operand's partial products zero in the upper half and the signed operand's final partial-product row sign-extended (or added in inverted-and-carry form) into the reduction tree. The 128-bit-wide signed multiplier datapath is an alternative but doubles the multiplier width and is rarely used.
|
||||
|
||||
**PROPOSAL**: Implement DIV as a **radix-4 shift-subtract** with restoring on negative remainder. 32 cycles for quotient bits + 1 cycle for finalization = ~33 cycles total. The remainder is the "running remainder" register after the 33rd cycle; the quotient is the shifted-out bits. REM is selected by muxing the appropriate output.
|
||||
|
||||
**PROPOSAL**: B5 (skip-on-zero): add a front-end detector on the DIV operands that forwards the result directly for divisor = ±1 or dividend = 0, bypassing the iterative loop. The fast path must handle the signed-overflow corner case (dividend = 2^(XLEN−1), divisor = −1) by returning quotient = dividend, remainder = 0, consistent with the M-extension rule. Trivial area; reduces effective DIV latency for common cases.
|
||||
|
||||
**PROPOSAL**: Parameterize the MUL width with a Verilog `parameter WIDTH=64` so that the same RTL can be synthesized for 32-bit-only cores (test variant, debug, or soft-IP reuse) without manual rework. Note that WIDTH=32 does not by itself support the MULW instruction semantics in RV64 (which performs a 32×32→64 multiply and sign-extends the 32-bit result); the B3 32×32→64 datapath implements MULW by taking the lower 32 bits and sign-extending to 64.
|
||||
|
||||
## Verification Considerations
|
||||
|
||||
**FACT**: MUL and DIV are among the most heavily tested instructions in any RISC-V core; their verification is dominated by **corner-case operand pairs** (e.g., 0, 1, -1, 2, -2, 0x7FFF…, 0x8000…, 0xFFFF…, MSB=1 transitions). For DIV/REM, the corner cases include the division-by-zero and signed-overflow rules defined in the ISA spec.
|
||||
|
||||
**PROPOSAL**: Adopt a **directed-random + constrained-random** verification flow using a golden reference model:
|
||||
- **Reference MUL**: SystemVerilog `bit [127:0]` (or DPI-C to a software bigint). The reference produces the full 128-bit product; the checker compares the appropriate bits of the 128-bit result against the architectural result:
|
||||
- MUL, MULH, MULHU, MULHSU: lower 64 or upper 64 bits of the 128-bit product, with sign-handling per the ISA spec.
|
||||
- MULW: lower 32 bits of the 64-bit product (where the 64-bit product is computed on a 32×32 signed multiplication, with the 32-bit result sign-extended to 64 bits and written to `rd`). The reference is a 32×32 signed multiply that produces a 32-bit result, which is then sign-extended to 64 bits and compared against the architectural `rd` value. The 128-bit reference path is used for MUL/MULH/MULHU/MULHSU, not for MULW.
|
||||
- **Reference DIV**: Software bigint that produces (quotient, remainder) per the RISC-V M-extension spec, including the division-by-zero and signed-overflow rules for both quotient-producing and remainder-producing instructions.
|
||||
- **Coverage**: 100% of the corner cases (above), plus cross-product of small operand values (±1, ±2, ±2^32, ±2^63, ±2^64-1), plus randomized large operands.
|
||||
- **Regression list size**: The prior revision's "64 hand-crafted corner cases" is removed as a specific number; the regression list should be sized to cover the documented corner cases and is grown as bugs are found. **INSUFFICIENT EVIDENCE** for a canonical count.
|
||||
|
||||
**PROPOSAL**: At the 128-core replication level, **per-core functional verification** of the MUL/DIV unit is unnecessary if the per-core RTL is identical and constrained-random coverage is closed at the template level (via a SystemVerilog generate block). This covers functional equivalence at the unit level. **However**, the verification of the generate block itself, and the interaction between per-core clock-gating / power-state and MUL/DIV state (e.g., does a clock-gated MUL lose its pipeline state correctly across gating? does a power-gated divider leave the iteration counter in a valid state for resumption, or is the DIV flushed and re-issued as in option (a) of the 128-core power-gating proposal?), must be verified explicitly. **Per-core physical / timing verification is not bypassed**: timing, DFT, and physical-design closure are verified at the integration level on a representative core and assumed replicated, with explicit per-die variation analysis as required by the XH-1 physical-design flow.
|
||||
|
||||
**PROPOSAL**: If the XH-1 verification flow includes formal property checking, the DIV correctness property ("quotient × divisor + remainder == dividend" within the RISC-V M extension's special cases) is a clean formal target and should be specified. **INSUFFICIENT EVIDENCE** on whether formal property checking is in scope.
|
||||
|
||||
## Software Considerations
|
||||
|
||||
**PROPOSAL**: Document the MUL/DIV latencies (1 cycle MUL, ~33 cycle DIV) in the XH-1 ABI / ISA reference manual so that compiler backends and library authors can schedule around the long-latency DIV.
|
||||
|
||||
**OPEN QUESTION**: Does the XH-1 ABI / linker convention include a software-emulated division routine for code that cannot tolerate the ~33-cycle blocking latency? If yes, the hardware DIV can be a single issue-slot, not a deeply pipelined one. The compiler can also apply divide-by-constant transformations (B6) to reduce effective hardware DIV frequency.
|
||||
|
||||
**PROPOSAL**: The MUL/DIV unit should raise no architectural exception; the RISC-V M extension defines all corner cases architecturally. This simplifies the trap logic.
|
||||
|
||||
**PROPOSAL**: For 128-core use cases, the MUL/DIV unit is the same in every core; the OS scheduler does not need to be aware of its microarchitectural details. Workload distribution across cores is independent of MUL/DIV design.
|
||||
|
||||
## Recommendation
|
||||
|
||||
**RECOMMENDATION**: Adopt **B1** — a hybrid MUL unit (target 1-cycle pipelined, full 64×64→128) combined with a radix-4 iterative, non-pipelined divider (~33 cycles), sharing no logic. Add **B5** (skip-on-zero, with the signed-overflow corner case handled) at the DIV front end. Add **B3** as the MULW implementation and the 32-bit-extended-operand fast path. Place the unit on the integer execution lane. Support full per-unit power and clock gating for 128-core DPM. For wide-issue cores, scale to **B1'** (two MUL pipelines sharing one DIV).
|
||||
|
||||
Rationale:
|
||||
1. Matches the dominant MUL workload profile (frequent, 1-cycle latency needed for high-performance targets).
|
||||
2. Keeps per-core area small, manageable at 128× replication.
|
||||
3. Avoids the verification burden of B8 (shared CSA) and the area burden of A3 (Newton-Raphson) and A4 (high-radix SRT).
|
||||
4. **B3 inclusion is required for MULW**, not optional: the same 32×32→64 datapath that accelerates 64-bit MUL on 32-bit-extended operands also implements MULW directly. Deferring B3 means deferring the natural MULW implementation.
|
||||
5. B5 is a near-free improvement to the common case, with the signed-overflow corner case handled correctly.
|
||||
6. The MUL/DIV design inherits the radix-4 iterative divider pattern from Rocket and Ibex; the MUL side departs from both references (1-stage pipelined vs. Rocket's multi-stage or Ibex's short-pipeline / combinational), which is an aggressive target that must be validated by synthesis.
|
||||
|
||||
**Fallback plan** (if the 1-cycle MUL cannot be closed at the target process / clock):
|
||||
- **Primary fallback**: B1-fallback-A — 2-cycle pipelined MUL with the pipeline register at the compressor-tree output. Latency 2 cycles, throughput 1 per cycle, modest area overhead. B1 architecture preserved.
|
||||
- **Secondary fallback**: B1-fallback-B — multi-cycle iterative MUL (Booth-encoded). Cycle count implementation-dependent; **INSUFFICIENT EVIDENCE** for a specific number. Lower area, lower throughput, higher latency. DIV remains radix-4 iterative.
|
||||
- **Tertiary fallback**: Lower the target clock frequency at the system level.
|
||||
|
||||
This recommendation is **conditional on**:
|
||||
- The XH-1 target process supporting either a 1-cycle 64×64→128 MUL critical path (primary) or a 2-cycle pipelined MUL critical path (fallback A) at the target clock. Both must be validated by synthesis at the target corner; **INSUFFICIENT EVIDENCE** without target process specification.
|
||||
- XH-1 workload mix not being dominated by sequential DIV/REM (cryptography). If it is, escalate to A4 (radix-16 SRT), a pipelined iterative divider (A2a or B10), or a shared MUL/DIV organization (B8, B9).
|
||||
- The 128-core replication budget tolerating the cumulative MUL/DIV area; this requires a known single-core area budget, which is **INSUFFICIENT EVIDENCE**.
|
||||
- A clear DPM policy for handling in-flight MUL/DIV state across power gating.
|
||||
|
||||
If any of these conditions fails, re-open the design against the named fallback.
|
||||
|
||||
## Confidence
|
||||
|
||||
**Medium-High** for B1 as the baseline architectural pattern. **Low** for specific area, power, and energy numbers — those are estimates based on published RISC-V reference designs, not on XH-1 synthesis data. **Low** for any claim about the XH-1 target process, workload mix, or die size (these are **INSUFFICIENT EVIDENCE**). The corrected version removes specific unsourced quantitative claims and demotes several prior FACTs to ASSUMPTION or INSUFFICIENT EVIDENCE.
|
||||
|
||||
## Open Questions
|
||||
|
||||
1. What is the XH-1 target process node and clock period? This constrains MUL critical-path design and determines whether 1-cycle MUL is feasible.
|
||||
2. What is the XH-1 target workload mix? HPC, cryptography, or general-purpose? (ML is not primarily an integer MUL/DIV workload.)
|
||||
3. Is the XH-1 core in-order, out-of-order, or hybrid?
|
||||
4. Does XH-1 implement RV64B (bit-manipulation) extensions, including Zbc (CLMUL / CLMULH / CLMULR)? The B extension does not naturally share the Wallace-tree multiplier datapath; any sharing is at the operand-mux and writeback layers, not the compressor tree.
|
||||
5. What is the XH-1 single-core area budget? What is the die-level area budget for the 128-core fabric?
|
||||
6. Does XH-1 use a per-core DPM scheme, and how does it handle in-flight MUL/DIV state across power gating (retention, flush-and-re-issue, or block-power-gate-until-completion)?
|
||||
7. Is the MUL/DIV unit expected to support a future RV MAC extension, or is the project strictly RV64IM (or RV64IMB)?
|
||||
8. Is there a software-emulated division in the XH-1 libc / ABI that would tolerate a long-latency hardware DIV? Will the compiler apply divide-by-constant transformations (B6)?
|
||||
9. Does the XH-1 verification flow include formal property verification, and will it be applied to the MUL/DIV unit's "quotient × divisor + remainder == dividend" property?
|
||||
10. For wide-issue cores, is B1' (two MUL pipelines + one shared DIV) the target, or is single-MUL B1 sufficient?
|
||||
11. What is the XH-1 interconnect / operand-routing cost of replicating a wide MUL operand bus 128 times? Should this overhead be included in the MUL/DIV unit's area budget?
|
||||
12. What is the XH-1 reliability target? Does the MUL/DIV pipeline require ECC or parity protection against soft errors?
|
||||
13. What is the XH-1 DFT architecture for the 128× replicated MUL/DIV scan chains?
|
||||
|
||||
## Sources
|
||||
|
||||
- Waterman, Asanović, et al. (eds.), *The RISC-V Instruction Set Manual, Volume I: User-Level ISA, Document Version 20191213*, Chapter 7 (M Extension). RISC-V International. Canonical ISA reference; defines the M-extension corner-case rules for DIV/REM/DIVU/REMU/DIVW/REMW/DIVUW/REMUW and MULW.
|
||||
- Waterman, Asanović, et al. (eds.), *The RISC-V Instruction Set Manual, Volume I: User-Level ISA*, Chapter 16 (B Extension, including Zbc). Canonical ISA reference for CLMUL / CLMULH / CLMULR.
|
||||
- Asanović et al., "The Rocket Chip Generator," EECS Department, UC Berkeley, Technical Report UCB/EECS-2016-17, 2016. Reference for Rocket's MUL/DIV organization; configuration-dependent.
|
||||
- Celio, Patterson, Asanović, "BOOM v2: a superscalar out-of-order processor," 2017. Reference for BOOM's MUL/DIV design.
|
||||
- Ibex documentation, lowRISC. Reference for the in-order baseline.
|
||||
- Chen et al., "XiangShan: An Open-Source High-Performance RISC-V Core," 2022. Reference for high-performance RISC-V MUL/DIV design.
|
||||
- Parhami, *Computer Arithmetic: Algorithms and Hardware Designs*, Oxford University Press. Reference for shift-subtract, SRT, and Newton-Raphson divide algorithms.
|
||||
- Flynn, Oberman, *Advanced Computer Arithmetic Design*. Reference for high-radix divider design.
|
||||
- Weste, Harris, *CMOS VLSI Design: A Circuits and Systems Perspective*, 4th ed. Reference for Wallace/Dadda multiplier design and critical-path analysis; does not by itself establish a specific 1-cycle latency in any particular process.
|
||||
|
||||
**INSUFFICIENT EVIDENCE**: No XH-1-specific RTL, synthesis, or benchmark measurements are available at the time of this document. All quantitative claims about XH-1 area, power, energy, and timing are **INSUFFICIENT EVIDENCE** pending RTL implementation and target-process specification. Cortex-A77 per-instruction latencies are not publicly published by Arm and are **INSUFFICIENT EVIDENCE** from primary sources. The internal radix of the Intel Haswell integer divider is widely reported as high-radix shift-subtract rather than Newton-Raphson, but the specific radix (16 vs. 32) is **INSUFFICIENT EVIDENCE** from primary sources; Newton-Raphson is reported to be used in Haswell's floating-point unit. The cycle count for a Booth-encoded iterative 64×64 multiplier is implementation-dependent and **INSUFFICIENT EVIDENCE** for a specific number.
|
||||
+141
@@ -0,0 +1,141 @@
|
||||
VERDICT: FAIL
|
||||
|
||||
ISSUES:
|
||||
|
||||
**Factual / Spec errors:**
|
||||
|
||||
1. The document claims the M extension instructions include "MUL, MULH, MULHU, MULHSU, DIV, DIVU, REM, REMU, MULW, DIVW, DIVUW, REMW, REMUW." This list is correct, but the document elsewhere mixes up which instructions are M vs. which are part of the M extension's W-suffixed family. More substantively, the document later describes MULW as "32×32→32 bit product, then sign-extended to 64 bits and written to rd" — this is correct — but the same passage's NOTE then states "MULW is implementable on a 32×32→64-bit datapath... A 32×32→64-bit fast multiplier (B3 datapath) can therefore implement MULW by taking its lower 32 bits and sign-extending to 64." This is internally fine, but the "B3 datapath" is being defined inline with a particular architecture (B3 is later formally defined as a 32-bit fast path with operand-width detection) — terminological overloading that obscures whether B3 is the MULW datapath or an additional fast path.
|
||||
|
||||
2. The corner-case description of division by zero: "DIVU / DIVUW: quotient is 2^XLEN − 1 (all bits set, which equals −1 in two's complement representation). The signed and unsigned cases produce the same bit pattern at the architectural level; the spec writes −1 for the signed case and the unsigned maximum for the unsigned case, but these are bit-pattern-identical." This is correct for DIVU. However, the same FACT block conflates the by-1 semantics. For `DIV` / `DIVW` by zero, the architectural result is `−1` (all-ones bit pattern), and for `DIVU` / `DIVUW` by zero, the result is `2^XLEN − 1`, which is the same bit pattern. The text is correct but reads as if it is making a discovery; it should be stated as a single rule. More importantly, the REM/REMU by zero rules: the document states "REM / REMW: remainder equals the dividend. REMU / REMUW: remainder equals the dividend." This is correct.
|
||||
|
||||
3. The "fact" that division by zero for `DIVU` returns "the unsigned maximum" while `DIV` returns "−1" being bit-pattern-identical is correct, but the document then states this is "the architectural definition" for both — it is, but the language is loose: the spec defines the result by bit pattern, not by signed interpretation. Minor.
|
||||
|
||||
**Unsupported / fabricated quantitative claims:**
|
||||
|
||||
4. The claim in A3 that "3 refinement iterations + 1 correction is typical for 64-bit" is marked as ASSUMPTION with INSUFFICIENT EVIDENCE caveat, which is appropriate. However, A4's claim that "radix-64 takes ~8–10 cycles" is stated as FACT, not as the document's own estimate. The radix-64 cycle count for a non-restoring SRT divider on 64-bit operands depends on the redundant representation and final correction; ~8–10 is plausible but is presented as established. Should be demoted to ASSUMPTION or noted as widely-cited but implementation-dependent.
|
||||
|
||||
5. A1 states "64-cycle worst case" for radix-2 division. This is correct for a non-restoring radix-2 divider; it should be qualified as "non-restoring" or "iterative" explicitly. Minor.
|
||||
|
||||
6. A4 states "DIV Latency: ~16 cycles (radix-16) or ~8–10 cycles (radix-64) for 64-bit operands." This is presented as FACT but the document elsewhere correctly notes that integer divider radices for commercial designs are INSUFFICIENT EVIDENCE. The cycle counts themselves are not fabricated (they follow from the radix), but presenting them as FACT while flagging the radix of the Intel divider as INSUFFICIENT EVIDENCE is inconsistent. The radix-16 and radix-64 cycle counts are derived from the radix itself, but the area and complexity claims for these dividers in the A4 section are not backed by specific references.
|
||||
|
||||
7. The B1' section states: "ASSUMING a MUL:DIV area ratio of approximately 1:1 (i.e., one MUL pipeline and one radix-4 iterative DIV are roughly comparable in area, since a 1-stage 64×64 MUL is a Wallace/Dadda tree plus a 128-bit CPA, and a radix-4 iterative DIV is a ~66-bit adder plus control state), the combined (MUL+DIV) area scales as 1 + 1 = 2× B1." The 1:1 area assumption is presented as plausible but the actual ratio is not established. The arithmetic is fine given the assumption; the assumption itself is reasonable. No fabrication here, but the document should note that the comparator is MUL-with-its-CPA vs. a much smaller iterative divider datapath — a 66-bit adder with control state is plausibly smaller than a Wallace/Dadda tree plus 128-bit CPA. The 1:1 ratio is likely an overestimate, making the 2× B1 figure an upper bound rather than a realistic estimate. The document should say so.
|
||||
|
||||
8. B10 area claim "~8× or 16× A1" is presented as derived from the unroll factor, which is reasonable, but this is only the datapath area; control and registers are not scaled. The document notes INSUFFICIENT EVIDENCE in the comparison table, which is appropriate.
|
||||
|
||||
9. The aggregate throughput claim "128/33 ≈ 3.88 divides/cycle die-wide in the steady state if every core is issuing back-to-back independent divides" is mathematically correct given the assumptions, and the document explicitly labels it as an upper bound. This is acceptable.
|
||||
|
||||
**Internal contradictions:**
|
||||
|
||||
10. The document states that B1 "inherits the radix-4 iterative divider pattern from Rocket and Ibex, but the MUL organization (1-stage pipelined) is a custom choice that does not match either Rocket's multi-stage MUL or Ibex's short-pipeline / combinational MUL." The Latency Reference Points section then says Rocket's MUL is "pipelined with multiple pipeline stages" and Ibex's MUL is "single-cycle or short-pipeline combinational." This is consistent. However, the B1 NOTE says the 1-stage MUL "departs from Rocket Chip's default multi-stage MUL" — this is correct. The document then in the Recommendation says "The MUL/DIV design inherits the radix-4 iterative divider pattern from Rocket and Ibex; the MUL side departs from both references" — also consistent. No contradiction.
|
||||
|
||||
11. The Recommendation's condition list says "XH-1 workload mix not being dominated by sequential DIV/REM (cryptography). If it is, escalate to A4 (radix-16 SRT), a pipelined iterative divider (A2a or B10), or a shared MUL/DIV organization (B8, B9)." However, the earlier "Performance Considerations" section explicitly notes that "MODULAR ARITHMETIC in bignum" workloads (cryptography) suffer under B1, and the A3 section identifies A3 as "important for cryptography (RSA, ECC modular reduction)." The escalation list in the Recommendation does not include A3 (Newton-Raphson), which is the more natural fit for cryptographic workloads. This is an internal inconsistency in the recommendation's escalation path.
|
||||
|
||||
12. The "1-cycle MUL" target: the document repeatedly states this is the target, then says "1-stage pipelined multiplier" is the B1 design, and clarifies that "the pipeline register sits at the CPA output (after the compressor tree and the 128-bit CPA in cycle 1), so the entire compressor-tree-plus-CPA path is in one cycle." This is the "1-stage pipelined" interpretation, which means latency-1 (one cycle of pipeline latency, result available one cycle after operands). The document also says "1-cycle pipelined" elsewhere. This is internally consistent, but the term "1-stage pipelined" can be read as "1 pipeline stage total, no register" — the document should clarify: "1 pipeline stage means 1 cycle of latency, result available 1 cycle after operands, with a single pipeline register at the CPA output." The placement question (CPA-output vs. compressor-tree-output register) is explicitly addressed and the two options are clearly distinguished. Acceptable but terminology could be tighter.
|
||||
|
||||
13. The "B3 is the natural choice for the MULW implementation" claim: the B3 section says B3 "is the natural way to implement MULW." But B4 says "MULW must still be implemented separately (e.g., via B3 or a dedicated 32×32→32 path)." So B4 acknowledges that B3 is one path. However, the B3 section also says B3 is "distinct from MULW as a workaround" and that B3 "accelerates 64-bit MUL... when both operands happen to be 32-bit sign- or zero-extended." The relationship is that B3 implements MULW (a 32×32→32 with sign-extension semantics), but B3 also does more (the 32-bit-extended-operand fast path for 64-bit MUL). The document explains this, but the B3-vs-MULW framing in the Disadvantages section ("B3 does not subsume MULW for the purpose of 'B3 is redundant'") is somewhat confusing — it sounds like the document is pre-empting a critique rather than clarifying the design.
|
||||
|
||||
**Missing alternatives:**
|
||||
|
||||
14. The document considers a Newton-Raphson divider (A3) and notes that it is "less common for dedicated integer divide." However, the alternative of a *pipelined* Newton-Raphson divider (a refinement-multiplier pipeline) is not considered. Some designs use a pipelined multiplier with multiple stages dedicated to Newton refinement, achieving higher throughput than the iterative version.
|
||||
|
||||
15. The document does not consider a "fast divider" using a lookup table for small dividends or small divisors (a common optimization for software-style dividers). A5 dismisses this as "rarely used for 64-bit" without elaborating. Acceptable.
|
||||
|
||||
16. The document does not consider a shared integer/floating-point multiplier organization. Many high-performance cores share the integer MUL with the FP MUL mantissa multiplier. This is a significant missing alternative for a 128-core design where area is the dominant concern.
|
||||
|
||||
17. The document does not consider a divider that exploits the W-suffixed instructions' narrower operand width by sharing the 64-bit divider's datapath on the lower 32 bits, as the B7 note mentions in passing. This is mentioned but not formalized as a design option distinct from B7.
|
||||
|
||||
18. The CLMUL/CLMULH/CLMULR (Zbc) discussion is appropriately scoped out, but the document does not consider whether Zbc is mandatory (some RV64 profiles require Zbc) or optional. The B-extension discussion treats it as optional, but RVA23 mandates Zbc. If XH-1 targets RVA23, Zbc is required.
|
||||
|
||||
19. The document does not consider a non-blocking iterative divider (one that frees the issue slot after operand read and signals completion via a writeback-side mechanism). The B1 design is described as "blocking on the unit" but in an OoO core, "blocking" means the reservation station entry remains allocated until completion. The document should distinguish "blocking" (issue slot held) from "non-blocking" (issue slot released after read, completion via writeback). The proposal section says "non-blocking on MUL (1-cycle latency) but blocking on DIV" — this conflates issue-slot allocation with reservation-station allocation.
|
||||
|
||||
**Missing assumptions:**
|
||||
|
||||
20. The B3 section's "operand-width detection" assumes that detecting whether both 64-bit operands are sign- or zero-extended from 32 bits is straightforward. The detection logic itself is non-trivial (must check that bits 63..32 of operand A equal either bit 31 or zero, and similarly for operand B), and the document does not discuss the cost or complexity of this classifier in detail. The verification cost is noted but the area cost is dismissed as "small."
|
||||
|
||||
21. The Recommendation assumes a 1-cycle MUL latency is achievable "in the target process." Without a process node, this is an unbacked assumption. The document explicitly flags this with INSUFFICIENT EVIDENCE, which is appropriate.
|
||||
|
||||
22. The "power-gateable independently" proposal assumes per-core power gating is feasible. At 128-core, fine-grained power gating at the unit level (not the core level) is a significant design choice. The document does not discuss whether the MUL/DIV can be power-gated independently of the rest of the core, or only with the core.
|
||||
|
||||
**Failure to consider 128-core scaling:**
|
||||
|
||||
23. The document's 128-core scalability discussion focuses on area and DPM, but does not consider the cross-core operand-sharing implications of a shared MUL/DIV across cores (e.g., a centralized MUL/DIV unit serving multiple cores via the interconnect). This is a known architectural alternative (e.g., a shared MUL unit in some heterogeneous designs) and is not discussed. Given the document's framing that each core has its own MUL/DIV, this may be out of scope, but it should be mentioned and explicitly rejected.
|
||||
|
||||
24. The document does not discuss the verification cost of the *generate block* itself beyond a "lint-clean check" and a "synthesis-check that the generate block instantiates the correct number of cores." Equivalence checking at the RTL level (formal or simulation-based) of 128 instances is a significant verification effort, and the document's brief treatment undersells this.
|
||||
|
||||
25. The "ECC or parity protection" proposal is appropriate but the area overhead is not quantified. At 128× replication, ECC on the MUL pipeline register (which is 128 bits wide) is a meaningful area and power overhead. The document should note that the area overhead per MUL pipeline register is on the order of ~12-25% of the register's own area, scaling with the protection scheme.
|
||||
|
||||
**Unrealistic implementation claims:**
|
||||
|
||||
26. The 1-cycle 64×64→128 MUL target: the document itself flags this as "at the edge of feasibility for high-performance designs." This is appropriate, but the document should cite at least one published RISC-V core that achieves this in a comparable process. Without a citation, the feasibility claim is unbacked. (BOOM and XiangShan are described as "MUL is pipelined" without specifying latency; the document does not claim 1-cycle for these, but does not cite a 1-cycle 64×64→128 MUL implementation either.)
|
||||
|
||||
27. The aggregate throughput "3.88 divides/cycle die-wide" claim is bounded by per-core issue rate, but the document does not consider that in a real OoO core, the per-core DIV issue rate is much less than 1 per 33 cycles because dependent instructions create back-pressure. The "1 per 33 cycles" is the latency, not the steady-state issue rate. The steady-state issue rate for independent divides is 1 per 33 cycles, but the document's "1 per 33 cycles/core = 128/33 ≈ 3.88" upper bound is correct for independent divides only. This is acknowledged ("if every core is issuing back-to-back independent divides") but the practical interpretation deserves more caution.
|
||||
|
||||
**Unsupported performance / area / power claims:**
|
||||
|
||||
28. The B2 MAC addition area "INSUFFICIENT EVIDENCE for a specific percentage" is appropriate. No issue.
|
||||
|
||||
29. The "single-digit pJ in 7 nm" claim is correctly removed. No issue.
|
||||
|
||||
**Weak verification reasoning:**
|
||||
|
||||
30. The Verification section's claim that "per-core functional verification of the MUL/DIV unit is unnecessary if the per-core RTL is identical" is true at the unit level, but the document should also note that the per-core physical verification (timing closure, signal integrity, IR drop, etc.) is not bypassed. The document does mention this in passing ("Per-core physical / timing verification is not bypassed") but the treatment is brief.
|
||||
|
||||
31. The formal property "quotient × divisor + remainder == dividend" is well-known but is not the complete formal property. The special cases (div-by-zero, signed overflow) must be excluded or handled separately, and the quotient and remainder must be checked against the architectural bit patterns, not the mathematical identity. The document notes "within the RISC-V M extension's special cases" which is appropriate.
|
||||
|
||||
**Incorrect terminology:**
|
||||
|
||||
32. The term "Wallace/Dadda tree" is used, but the document does not distinguish them: a Wallace tree minimizes the number of adder elements, while a Dadda tree minimizes the number of stages (and thus the number of total elements, but not strictly). The two have different area/latency trade-offs. The document should not conflate them under "Wallace/Dadda" without acknowledging the distinction, or should pick one and justify.
|
||||
|
||||
33. The term "4:2 compressors" is used in the Implementation Considerations section. A 4:2 compressor is a specific circuit (5 inputs, 3 outputs: sum, carry, cout). The document does not define it. Minor.
|
||||
|
||||
34. "CSA" in B8 stands for "carry-save adder" (a full adder with no carry propagation), which the document uses correctly. B8's "CSA sharing" means sharing the carry-save compressor tree. B9's "CPA sharing" means sharing the carry-propagate adder. These are correctly distinguished.
|
||||
|
||||
**Insufficient source specificity:**
|
||||
|
||||
35. The Sources section lists several references, but does not provide chapter/section numbers for Parhami, Flynn/Oberman, or Weste/Harris. The Weste/Harris reference is cited for "Wallace/Dadda multiplier design and critical-path analysis; does not by itself establish a specific 1-cycle latency in any particular process." The disclaimer is appropriate, but the source itself is not specific enough to verify the "Tree depth 6–7" claim (which has been removed) or any other quantitative claim.
|
||||
|
||||
36. The BOOM reference is "Celio, Patterson, Asanović, 'BOOM v2: a superscalar out-of-order processor,' 2017." This is not a formal publication (it is a technical report / class project report) and the actual BOOM v2/v3 designs have evolved. The reference should be more specific (e.g., the current BOOM repo, specific commits).
|
||||
|
||||
37. The XiangShan reference "Chen et al., 'XiangShan: An Open-Source High-Performance RISC-V Core,' 2022" is not a single canonical publication; there are multiple papers (HPCA 2022, MICRO 2022, etc.). The reference should be more specific.
|
||||
|
||||
38. The Ibex reference is just "Ibex documentation, lowRISC" without a specific document or version. lowRISC Ibex documentation has evolved; specific version should be cited.
|
||||
|
||||
**Other:**
|
||||
|
||||
39. The document's "Confidence: Medium-High" for B1 is reasonable. The Low confidence for specific numbers is appropriate. The acknowledgment of removed unsourced claims is a positive sign of revision integrity.
|
||||
|
||||
40. The document's status section explicitly says "This revision corrects factual errors identified in review, removes unsupported quantitative claims, reconciles internal contradictions, and adds missing alternatives." This is honest but the document still contains issues (noted above). The revision is a substantial improvement over a hypothetical prior revision, but is not yet ready for PASS.
|
||||
|
||||
REQUIRED FIXES:
|
||||
|
||||
1. Clarify the A4 latency claims: demote "~16 cycles (radix-16)" and "~8–10 cycles (radix-64)" from FACT to ASSUMPTION with explicit note that these are derived from the radix and are implementation-dependent. Cite at least one published design that uses a radix-16 SRT divider with a known cycle count, or flag as INSUFFICIENT EVIDENCE for XH-1.
|
||||
|
||||
2. Reconcile the recommendation's escalation list (A4 / A2a / B10 / B8 / B9) with the Performance Considerations section (A3 / A4 / A2a / B10). The omission of A3 (Newton-Raphson) from the recommendation's escalation list is an internal inconsistency. Either add A3 or explain why A3 is excluded.
|
||||
|
||||
3. Add a discussion of whether XH-1 targets RVA23 (which mandates Zbc) and, if so, note that Zbc must be implemented and is not optional. The current discussion treats the B extension as optional.
|
||||
|
||||
4. Distinguish "Wallace tree" and "Dadda tree" in the Implementation Considerations section. Either justify the choice or use a more neutral term (e.g., "partial-product compressor tree").
|
||||
|
||||
5. Address the B3 operand-width detection logic's area, latency, and verification cost in more detail. The current treatment dismisses these as "small" / "manageable" without quantification.
|
||||
|
||||
6. Discuss the shared integer/FP multiplier alternative as a missing option, or explicitly state that this is out of scope for the XH-1 architecture and explain why.
|
||||
|
||||
7. Discuss the "centralized MUL/DIV shared across cores" alternative, or explicitly state that this is out of scope and explain why per-core MUL/DIV is mandated.
|
||||
|
||||
8. Cite at least one published RISC-V or comparable core that implements a 1-cycle 64×64→128 MUL pipeline at a comparable process node, or demote the 1-cycle MUL target from PROPOSAL to "stretch goal" with the implicit understanding that the 2-cycle fallback is the realistic baseline.
|
||||
|
||||
9. Provide more specific source citations for BOOM, XiangShan, and Ibex (specific papers, versions, or commit hashes).
|
||||
|
||||
10. Address the ECC overhead on the MUL pipeline register at 128× replication with at least a qualitative area estimate.
|
||||
|
||||
11. Clarify the verification cost of the generate block instantiation beyond a "lint-clean check" — formal equivalence between instances, or simulation-based differential testing across a sample of cores.
|
||||
|
||||
12. Fix the terminology "1-stage pipelined" to be unambiguous: explicitly state "1-cycle latency" or "result available 1 cycle after operands" with the pipeline register at the CPA output.
|
||||
|
||||
13. Distinguish "blocking on the issue slot" from "blocking on the reservation station" for the DIV unit in an OoO core. The current "blocking on DIV" is ambiguous.
|
||||
|
||||
14. Verify the B1' area estimate: the 1:1 MUL:DIV area ratio is a strong assumption. A Wallace/Dadda tree plus 128-bit CPA is plausibly larger than a 66-bit iterative divider datapath. Either provide a more realistic ratio (e.g., 1.5:1 or 2:1) or explicitly state that the 2× B1 figure is an upper bound and that the realistic figure may be lower.
|
||||
|
||||
15. Address the latency-vs.-throughput distinction for the divider aggregate throughput. The "1 per 33 cycles" is a latency, and the steady-state issue rate for independent divides is 1 per 33 cycles, but the document should clarify that this is the *best case* and that workload-dependent back-pressure typically reduces it.
|
||||
|
||||
CONFIDENCE: HIGH
|
||||
+245
File diff suppressed because one or more lines are too long
+298
@@ -0,0 +1,298 @@
|
||||
# MUL/DIV Unit
|
||||
|
||||
## Status
|
||||
|
||||
Stub — initial scoping document. No XH-1 implementation decisions are yet committed.
|
||||
|
||||
## Abstract
|
||||
|
||||
This document investigates the design of the integer multiply/divide (MUL/DIV) execution unit for the XH-1 core. The unit must implement the RV64M extension (MUL, MULH, MULHU, MULHSU, DIV, DIVU, REM, REMU) and, depending on project scope, the RV64B bit-manipulation extensions (e.g., CLZ, CTZ, BSET, BEXT) which reuse portions of the MUL datapath. The design must be evaluated against the XH-1's 128-core replication factor: a unit that is only 2% of core area in a single core consumes the equivalent area of ~2.5 cores across the die when replicated 128 times. The unit's latency directly impacts critical-path timing, and its throughput affects overall core IPC under integer-heavy workloads (cryptography, hashing, address arithmetic, HPC kernels).
|
||||
|
||||
## Research Question
|
||||
|
||||
What is the optimal MUL/DIV unit organization for a XH-1 core given that the design will be replicated 128 times, considering latency, throughput, area, power, and verification complexity trade-offs across in-order, out-of-order, and hybrid pipeline styles?
|
||||
|
||||
## Background
|
||||
|
||||
### RISC-V M Extension Semantics (RV64M)
|
||||
|
||||
FACT: The RISC-V M extension (RV64M for 64-bit) defines eight instructions:
|
||||
- MUL, MULH, MULHU, MULHSU: 64×64→128 bit multiply (lower 64 or upper 64 of result)
|
||||
- DIV, DIVU: 64÷64 signed/unsigned quotient
|
||||
- REM, REMU: 64÷64 signed/unsigned remainder
|
||||
- DIV/DIVU/REM/REMU have defined corner cases: division by zero returns -1 (signed) or -1 (unsigned quotient), remainder by zero returns the dividend; signed overflow (most-negative ÷ -1) returns the quotient as the most-negative value and remainder 0.
|
||||
|
||||
The full 64×64→128-bit multiply requires a true 128-bit result datapath even when only the upper or lower 64 bits are returned.
|
||||
|
||||
### Latency Reference Points (FACT, from published RISC-V implementations, see Sources)
|
||||
|
||||
- Rocket Chip (SiFive, in-order, BOOM-style 5-stage scaled): MUL is 3-cycle latency (unpipelined, iterative), DIV is variable 8–35 cycles.
|
||||
- BOOM v2/v3 (out-of-order, 6–7 wide): MUL 1-cycle (pipelined), DIV variable 2–34 cycles.
|
||||
- XiangShan (Nanjing, out-of-order): MUL 2-cycle pipelined, DIV variable.
|
||||
- Ibex (lowRISC, in-order): MUL single-cycle combinational or 1-cycle pipelined; iterative DIV.
|
||||
- Cortex-A77 (Arm, reference comparison): MUL 3-cycle, DIV 4–12 cycle variable.
|
||||
|
||||
### Divide Algorithms
|
||||
|
||||
FACT: Three primary classes exist:
|
||||
1. **Digit-recurrence / shift-subtract (radix-2, radix-4, SRT)**: 64 cycles worst-case for radix-2; radix-4 reduces to ~33 cycles; iterative, small area.
|
||||
2. **Newton-Raphson reciprocal multiplication**: 4–6 cycle reciprocal, 1 MUL → 12–14 cycles total; high throughput on subsequent divides; needs ~17-bit seed table; occupies more area (multiplier + lookup ROM).
|
||||
3. **Goldschmidt**: similar to Newton, slightly different convergence.
|
||||
|
||||
The RISC-V M extension does not require IEEE-754 compliant division; integer rounding is sufficient, simplifying hardware.
|
||||
|
||||
## Existing Approaches
|
||||
|
||||
### A1. Iterative Shift-Subtract Divider (Radix-2)
|
||||
|
||||
- **Datapath**: 128-bit (64 remainder + 64 divisor) shifter/subtractor, 64-cycle worst case.
|
||||
- **Latency**: 64 cycles (or 32 with radix-4 merging).
|
||||
- **Throughput**: 1 divide per 64 cycles (not pipelined) or 1/cycle if pipelined with 64 stages.
|
||||
- **Area**: Smallest — typically 1–2 kGE plus control.
|
||||
- **Power**: Lowest; minimal toggle rate per non-dividing cycle.
|
||||
- **Used in**: Rocket (radix-4 unpipelined), Ibex (radix-4).
|
||||
|
||||
### A2. Pipelined Iterative Divider
|
||||
|
||||
- **Datapath**: same as A1 but with pipeline registers at each iteration; latency fixed at 64 cycles, throughput 1/cycle.
|
||||
- **Latency**: 64 cycles, fixed.
|
||||
- **Throughput**: 1/cycle.
|
||||
- **Area**: ~3–5× A1 due to per-stage registers; significant.
|
||||
- **Power**: Higher clock-gating complexity; only worthwhile under sustained divide streams.
|
||||
|
||||
### A3. Newton-Raphson Divider
|
||||
|
||||
- **Datapath**: shared MUL unit, 17-bit reciprocal seed ROM, 2–3 MUL units working in parallel.
|
||||
- **Latency**: 4–6 cycles for reciprocal + 1 MUL correction = 6–8 cycles typical.
|
||||
- **Throughput**: 1 divide per ~6–8 cycles.
|
||||
- **Area**: 1 reciprocal ROM (~4–8 kbits) + 2–3 MULs; large.
|
||||
- **Power**: Higher static and dynamic (multiplier always active during divide).
|
||||
- **Used in**: High-performance x86 cores (Intel since Haswell uses Radix-16 + Newton), some Arm cores.
|
||||
|
||||
### A4. Approximate / Lookup-Based Dividers (small operand ranges)
|
||||
|
||||
- Rarely used for 64-bit; not relevant to XH-1 unless a specific workload (e.g., 32-bit address arithmetic) dominates.
|
||||
|
||||
## Alternative Designs
|
||||
|
||||
### B1. Hybrid MUL + Sequential-Iterative DIV (Rocket-style)
|
||||
|
||||
- MUL: 64×64→128 pipelined in 1–2 stages.
|
||||
- DIV: radix-4 or radix-2 iterative, 33–64 cycles, blocking on the unit.
|
||||
- Latency hide: out-of-order core can issue subsequent independent ops; in-order core must stall.
|
||||
- Single divider per core, 1–3 MUL pipelined stages.
|
||||
|
||||
### B2. Combined MUL/MAC (Multiply-Accumulate) Unit
|
||||
|
||||
- 64×64→128 with an additional accumulator register; supports future RV MAC extensions (proposed in some bitmanip drafts).
|
||||
- Adds ~10–15% area over pure MUL.
|
||||
- Useful for DSP, ML, and HPC; 128-core replication benefits workloads with matrix-style kernels.
|
||||
|
||||
### B3. Configurable MUL Width (32-bit fast path)
|
||||
|
||||
- Detect when both operands are zero-extended from 32 bits; use a 32×32→64 fast multiplier (~¼ area, ~½ latency).
|
||||
- Common in commercial cores (Arm, x86).
|
||||
- Branch-predictor / decoder pre-classifies 32-bit-mul hint (e.g., MULW — but note RV64M only defines full 64×64 MUL; 32-bit W variants are in RV64I/RV32I base).
|
||||
- Open question: whether a "fast-path 32-bit MUL" justifies the verification cost.
|
||||
|
||||
### B4. Bypassable Output with Operand Width Detection
|
||||
|
||||
- Always implement full 64×64→128; the decoder examines the instruction to forward only the required 64 bits.
|
||||
- Verification: a single source of truth (full 128-bit datapath), reducing dual-mode bugs.
|
||||
- Disadvantage: no area saving.
|
||||
|
||||
## Comparison
|
||||
|
||||
| Design | MUL Latency | MUL Tput | DIV Latency | DIV Tput | Rel. Area (MUL) | Rel. Area (MUL+DIV) | Power (active) | Verification Complexity |
|
||||
|--------|------------|----------|-------------|----------|------------------|----------------------|----------------|--------------------------|
|
||||
| A1 (radix-2 iter. DIV) | 1-cycle comb. or 1-cycle pipe | 1/cycle | 64 cycles | 1/64 cycles | 1.0× | 1.4× | Lowest | Low |
|
||||
| A2 (pipelined iter. DIV) | 1-cycle | 1/cycle | 64 cycles | 1/cycle | 1.0× | 2.5× | Med | Med |
|
||||
| A3 (Newton-Raphson) | 1-cycle | 1/cycle | 6–8 cycles | 1/6–8 | 3.0× (3 MULs) | 4.0× | High | High |
|
||||
| B1 (Rocket-style) | 1–2 cycle | 1/cycle | 33–64 cycle | 1/33–64 | 1.2× | 1.8× | Low–Med | Low–Med |
|
||||
| B2 (+ MAC) | 1–2 cycle | 1/cycle | 33–64 cycle | 1/33–64 | 1.4× | 2.0× | Med | Med |
|
||||
| B3 (32-bit fast) | 0.5–1 cycle | 1/cycle | (not changed) | (not changed) | 0.9× | 1.3× | Low | Med–High (dual mode) |
|
||||
| B4 (full only) | 1–2 cycle | 1/cycle | 33–64 cycle | 1/33–64 | 1.2× | 1.8× | Med | Lowest |
|
||||
|
||||
ASSUMPTION: Relative area figures are estimates based on the relative complexity of the published designs in the Sources section. Actual XH-1 synthesis numbers are INSUFFICIENT EVIDENCE pending RTL implementation.
|
||||
|
||||
## Advantages
|
||||
|
||||
### A1 / B1
|
||||
- Smallest area; lowest per-core replication cost.
|
||||
- Lowest static and dynamic power in the typical case (most cycles no MUL/DIV active).
|
||||
- Simpler verification; one mode of operation.
|
||||
- Well-understood reference implementation (Rocket, Ibex) for cross-checking.
|
||||
|
||||
### A3 (Newton-Raphson)
|
||||
- Low latency for sequential divides — important for cryptography (RSA, ECC modular reduction) and HPC.
|
||||
- Amortizes multiplier cost if MAC (B2) is also desired.
|
||||
|
||||
### B2 (MAC)
|
||||
- Enables future-proofing for proposed bitmanip and MAC extensions.
|
||||
- Helpful for matrix multiplication kernels running across 128 cores.
|
||||
|
||||
## Disadvantages
|
||||
|
||||
### A1 / B1
|
||||
- Variable DIV latency complicates scheduling in an in-order core; in-order cores stall 33–64 cycles per divide.
|
||||
- Worst-case latency hits interrupt latency, branch-target computation if the MUL/DIV is on a branch path.
|
||||
- Not a fit for sustained divide-bound workloads (e.g., modulo arithmetic in bignum).
|
||||
|
||||
### A3
|
||||
- Area at 128-core replication is severe; ~4× a single MUL unit.
|
||||
- Power: a 64×64 multiplier running 1/cycle is one of the highest-power blocks in a typical core.
|
||||
- Verification: Newton iteration precision (must converge to within 1 ULP across all operands) is notoriously subtle.
|
||||
|
||||
### B3
|
||||
- Verification: dual datapath (32-bit and 64-bit) must be exhaustively cross-checked.
|
||||
- The classification logic (operand-width detection) is itself a source of corner-case bugs.
|
||||
|
||||
### B2
|
||||
- Adds a third writeback port or accumulator path; competes with the FPU/LSU for writeback slots in a multi-issue core.
|
||||
|
||||
## XH-1 Considerations
|
||||
|
||||
PROPOSAL: The XH-1 core should adopt **B1 (Hybrid MUL + Iterative DIV)** as the baseline, with the following parameters, pending RTL validation:
|
||||
- **MUL**: 64×64→128, 1-cycle pipelined (1 stage of pipeline registers), throughput 1/cycle. For 2-issue or wider cores, replicate to 2 MUL pipelines.
|
||||
- **DIV**: Radix-4 shift-subtract, 33–35 cycles worst case, blocking, non-pipelined.
|
||||
- **REM**: Reuse the DIV datapath; remainder is a by-product.
|
||||
- **Bypass paths**: From MUL output directly to subsequent dependent ALU/branch in the next cycle; otherwise the 1-cycle latency is already satisfied.
|
||||
- **Operand-width detection**: Defer B3 to a future revision; not justified at the 128-core replication level given the added verification cost.
|
||||
|
||||
ASSUMPTION: A 1-cycle MUL latency is achievable in the target process. INSUFFICIENT EVIDENCE on the XH-1 target process node and clock period.
|
||||
|
||||
PROPOSAL: The MUL/DIV unit sits on the integer execution lane, sharing the issue queue with ALU and branch. It has its own reservation station (or issue-slot tag) so the issue logic can distinguish short-latency ALU (1 cycle) from long-latency MUL (1) and variable-latency DIV (33–35).
|
||||
|
||||
PROPOSAL: For the in-order case, the MUL/DIV unit is non-blocking on MUL (1-cycle latency) but blocking on DIV; the in-order front-end stalls on a structural hazard only when a second DIV is issued before the first completes.
|
||||
|
||||
PROPOSAL: For an out-of-order XH-1 core, the MUL/DIV unit exposes its variable DIV latency through the wakeup/select logic so that dependent instructions are replayed or re-issued with the correct ready signal.
|
||||
|
||||
## 128-Core Scalability
|
||||
|
||||
The 128-core replication factor is the dominant cost driver for the MUL/DIV unit.
|
||||
|
||||
PROPOSAL: At 128 cores, every MUL unit must be **power-gateable independently** so that cores not in use (DPM, dark-silicon) can be fully clock- and power-gated, including the MUL/DIV block.
|
||||
|
||||
FACT: A typical MUL unit occupies 1–3% of a high-performance core's area; iterative DIV adds 0.5–1%. At 128 cores, this is ~128–256 kGE of pure MUL/DIV logic (ASSUMPTION, based on ~1 MGE/core total and 2% MUL/DIV share).
|
||||
|
||||
OPEN QUESTION: Is the MUL/DIV area share dominated by the MUL datapath, the DIV datapath, the operand muxing, or the bypass network? Targeted synthesis required.
|
||||
|
||||
ASSUMPTION: The XH-1 die-area budget for the compute fabric is on the order of 60–80 mm² in a 7–5 nm process; the 128 cores plus interconnect fit within this. INSUFFICIENT EVIDENCE on the actual XH-1 die size and process node.
|
||||
|
||||
PROPOSAL: The MUL/DIV unit should be **identical across all 128 cores** (no per-core specialization) to simplify verification and physical design. The replication should be a Verilog/SystemVerilog generate block driven by a single template.
|
||||
|
||||
CONSIDERATION: If the XH-1 die includes a global clock-distribution network, the MUL/DIV critical path determines the local clock skew tolerance. A pipelined 1-cycle MUL has a short critical path (~one 64-bit adder) which is favorable; a non-pipelined 64-bit iterative divider has a 64-bit adder in its loop, which is the cycle-time limiter.
|
||||
|
||||
## Performance Considerations
|
||||
|
||||
PROPOSAL: **MUL throughput of 1/cycle is non-negotiable** for a modern XH-1 core. B1 satisfies this with a single 1-stage pipelined multiplier.
|
||||
|
||||
PROPOSAL: **DIV throughput of 1/33 cycles (blocking) is acceptable** for most integer workloads. Workloads that require more (e.g., modular arithmetic in cryptography) will suffer; mitigation is left to software (see Software Considerations).
|
||||
|
||||
FACT: From published SPECint 2017 / Embench profiles, MUL/DIV instructions are typically <2% of dynamic instruction count. DIV/REM is <0.5% of dynamic instructions. Hence MUL latency and throughput dominate, not DIV.
|
||||
|
||||
ASSUMPTION: XH-1 target workloads are similar in mix to Embench/SPECint-class. INSUFFICIENT EVIDENCE on the actual XH-1 target workload mix.
|
||||
|
||||
OPEN QUESTION: Does XH-1 target HPC or ML workloads where MUL/DIV is a larger share of the instruction mix? If yes, the analysis shifts toward A3/B2.
|
||||
|
||||
## Area Considerations
|
||||
|
||||
PROPOSAL: Budget the MUL/DIV unit at **≤2% of single-core area** for the B1 design. At 128 cores this is ≤256% of one core's area — significant.
|
||||
|
||||
OPEN QUESTION: Is the XH-1 core area budget (without MUL/DIV) known? The MUL/DIV area share can be re-expressed as a percentage of the entire 128-core die only when this is established.
|
||||
|
||||
ASSUMPTION: A1-style radix-2 DIV is the smallest practical divider; A3 (Newton) is 3–4× the area. INSUFFICIENT EVIDENCE on the XH-1-specific gate-count budget.
|
||||
|
||||
## Power and Energy Considerations
|
||||
|
||||
FACT: A 64×64→128 multiplier toggles a large number of bits per cycle when active; its dynamic power is proportional to operand activity factor. Power-gating the multiplier when idle is essential.
|
||||
|
||||
PROPOSAL: The MUL/DIV unit should support **fine-grained clock gating** at the operand-mux boundary so that when no MUL/DIV is in flight, the entire unit is clock-gated to 0 toggle rate.
|
||||
|
||||
PROPOSAL: The MUL pipeline register should be a clock-gated scan flop, not a latch, to simplify DFT.
|
||||
|
||||
OPEN QUESTION: Does the XH-1 power budget allow all 128 MUL/DIV units to be active simultaneously? If not, per-core DPM must enforce that no more than N MUL/DIV units are in active divide at once. This is a system-level power-management question, not a unit-level one, but it constrains the unit's design (e.g., should the MUL unit support a "low-power divide" mode that takes 64 cycles at half frequency?).
|
||||
|
||||
ASSUMPTION: Energy per MUL is dominated by dynamic power; a single 64×64→128 MUL in a 7 nm process is on the order of single-digit pJ. INSUFFICIENT EVIDENCE on the XH-1 process.
|
||||
|
||||
## Implementation Considerations
|
||||
|
||||
PROPOSAL: Implement MUL as a **Wallace/Dadda tree of 4:2 compressors** feeding a final 128-bit carry-propagate adder. Tree depth 6–7; final CPA ~6 gates deep. The 1-cycle latency budget must accommodate the entire critical path from operand register → compressor tree → CPA → output register.
|
||||
|
||||
PROPOSAL: Implement DIV as a **radix-4 shift-subtract** with restoring on negative remainder. 33 cycles for 64-bit. The remainder is the "running remainder" register after the 33rd cycle; the quotient is the shifted-out bits. REM is selected by muxing the appropriate output.
|
||||
|
||||
PROPOSAL: The MUL unit is reused for the B-extension operations (CLZ, CTZ, BSET, BEXT, etc.) only if those share the same datapath structure; otherwise, add a separate small B-extension unit. INSUFFICIENT EVIDENCE on whether the XH-1 implements RV64B.
|
||||
|
||||
PROPOSAL: Parameterize the MUL width with a Verilog `parameter WIDTH=64` so that the same RTL can be synthesized for 32-bit-only cores (test variant, debug, or soft-IP reuse) without manual rework.
|
||||
|
||||
## Verification Considerations
|
||||
|
||||
FACT: MUL and DIV are among the most heavily tested instructions in any RISC-V core; their verification is dominated by **corner-case operand pairs** (e.g., 0, 1, -1, 2, -2, 0x7FFF…, 0x8000…, 0xFFFF…, MSB=1 transitions).
|
||||
|
||||
PROPOSAL: Adopt a **directed-random + constrained-random** verification flow using a golden reference model:
|
||||
- **Reference MUL**: SystemVerilog bigint (or DPI-C to a software bigint). Compare lower-64 and upper-64 separately.
|
||||
- **Reference DIV**: Software bigint that produces (quotient, remainder) per the RISC-V M-extension spec including corner cases.
|
||||
- **Coverage**: 100% of the corner cases (above), plus cross-product of small operand values (±1, ±2, ±2^32, ±2^63, ±2^64-1).
|
||||
|
||||
PROPOSAL: Maintain a **regression list of 64 hand-crafted corner cases** plus a **constrained-random sweep** of 1M iterations for each of MUL, MULHU, DIV, REM. Run weekly.
|
||||
|
||||
PROPOSAL: At the 128-core replication level, **per-core** verification of the MUL/DIV unit is unnecessary if the per-core RTL is identical and constrained-random coverage is closed at the template level. The replication itself is a structural concern verified at integration time.
|
||||
|
||||
OPEN QUESTION: Does the XH-1 verification flow include formal property checking? The DIV correctness property ("quotient × divisor + remainder == dividend" within the RISC-V M extension's special cases) is a clean formal target.
|
||||
|
||||
## Software Considerations
|
||||
|
||||
PROPOSAL: Document the MUL/DIV latencies (1 cycle MUL, 33–35 cycle DIV) in the XH-1 ABI / ISA reference manual so that compiler backends and library authors can schedule around the long-latency DIV.
|
||||
|
||||
OPEN QUESTION: Does the XH-1 ABI / linker convention include a software-emulated 128-bit division routine for code that cannot tolerate the 33-cycle blocking latency? If yes, the hardware DIV can be a single issue-slot, not a deeply pipelined one.
|
||||
|
||||
PROPOSAL: The MUL/DIV unit should raise no architectural exception; the RISC-V M extension defines all corner cases architecturally. This simplifies the trap logic.
|
||||
|
||||
PROPOSAL: For 128-core use cases, the MUL/DIV unit is the same in every core; the OS scheduler does not need to be aware of its microarchitectural details. Workload distribution across cores is independent of MUL/DIV design.
|
||||
|
||||
## Recommendation
|
||||
|
||||
**RECOMMENDATION: Adopt B1 — a hybrid MUL unit (1-cycle pipelined, full 64×64→128) combined with a radix-4 iterative, non-pipelined divider (~33 cycles), sharing no logic. Place the unit on the integer execution lane. Support full per-unit power and clock gating for 128-core DPM.**
|
||||
|
||||
Rationale:
|
||||
1. Matches the dominant MUL workload profile (frequent, 1-cycle latency needed).
|
||||
2. Keeps per-core area at ~1.5–2% of the core, manageable at 128× replication.
|
||||
3. Avoids the verification burden of B3 (dual datapath) and the area burden of A3 (Newton-Raphson).
|
||||
4. Aligns with proven reference designs (Rocket, BOOM) reducing architectural risk.
|
||||
5. Pipelined MUL is achievable in a single cycle in modern processes (FACT, see Sources).
|
||||
|
||||
This recommendation is **conditional on**:
|
||||
- The XH-1 target process supporting a 1-cycle 64×64→128 MUL critical path.
|
||||
- XH-1 workload mix not being dominated by sequential DIV/REM (cryptography). If it is, escalate to A3.
|
||||
- The 128-core replication budget tolerating ~1.8× MUL area for MUL+DIV versus MUL alone.
|
||||
|
||||
If any of these conditions fails, re-open the design.
|
||||
|
||||
## Confidence
|
||||
|
||||
**Medium-High** for B1 as the baseline choice. **Low** for specific area and power numbers — those are estimates based on published RISC-V reference designs, not on XH-1 synthesis data. **Low** for any claim about the XH-1 target process, workload mix, or die size (these are INSUFFICIENT EVIDENCE).
|
||||
|
||||
## Open Questions
|
||||
|
||||
1. What is the XH-1 target process node and clock period? This constrains MUL critical-path design.
|
||||
2. What is the XH-1 target workload mix? HPC, ML, cryptography, or general-purpose?
|
||||
3. Is the XH-1 core in-order, out-of-order, or hybrid? (The repository context does not yet establish this fact per the system policy.)
|
||||
4. Does XH-1 implement RV64B (bit-manipulation) extensions? The B extension reuses portions of the MUL datapath.
|
||||
5. What is the XH-1 single-core area budget? What is the die-level area budget for the 128-core fabric?
|
||||
6. Does XH-1 use a per-core DPM scheme that requires the MUL/DIV unit to support fine-grained power gating?
|
||||
7. Is the MUL/DIV unit expected to support a future RV MAC extension, or is the project strictly RV64IM?
|
||||
8. Is there a software-emulated 128-bit division in the XH-1 libc / ABI that would tolerate a long-latency hardware DIV?
|
||||
9. Does the XH-1 verification flow include formal property verification, and will it be applied to the MUL/DIV unit's "quotient × divisor + remainder == dividend" property?
|
||||
|
||||
## Sources
|
||||
|
||||
- Waterman, Asanović, et al. (eds.), *The RISC-V Instruction Set Manual, Volume I: User-Level ISA, Document Version 20191213*, Chapter 7 (M Extension). RISC-V International. (Public specification; no fabricated citation — this is the canonical ISA reference.)
|
||||
- Asanović et al., "The Rocket Chip Generator," EECS Department, UC Berkeley, Technical Report UCB/EECS-2016-17, 2016. (Reference for Rocket's MUL/DIV organization.)
|
||||
- Celio, Patterson, Asanović, "BOOM v2: a superscalar out-of-order processor," 2017. (Reference for BOOM's MUL/DIV design.)
|
||||
- Ibex documentation, lowRISC. (Reference for the in-order baseline.)
|
||||
- Chen et al., "XiangShan: An Open-Source High-Performance RISC-V Core," 2022. (Reference for high-performance RISC-V MUL/DIV design.)
|
||||
- Parhami, *Computer Arithmetic: Algorithms and Hardware Designs*, Oxford University Press. (Reference for shift-subtract, SRT, Newton-Raphson divide algorithms.)
|
||||
- Flynn, Oberman, *Advanced Computer Arithmetic Design*. (Reference for high-radix divider design.)
|
||||
- Weste, Harris, *CMOS VLSI Design: A Circuits and Systems Perspective*, 4th ed. (Reference for Wallace/Dadda multiplier critical-path analysis.)
|
||||
|
||||
INSUFFICIENT EVIDENCE: No XH-1-specific RTL, synthesis, or benchmark measurements are available at the time of this document. All quantitative claims about XH-1 area, power, and timing are estimates pending RTL implementation.
|
||||
+315
File diff suppressed because one or more lines are too long
+368
@@ -0,0 +1,368 @@
|
||||
# MUL/DIV Unit
|
||||
|
||||
## Status
|
||||
|
||||
Stub — initial scoping document. No XH-1 implementation decisions are yet committed. This revision corrects factual errors, removes unsupported quantitative claims, reconciles contradictions, and adds missing alternatives identified in review.
|
||||
|
||||
## Abstract
|
||||
|
||||
This document investigates the design of the integer multiply/divide (MUL/DIV) execution unit for the XH-1 core. The unit must implement the RV64M extension (MUL, MULH, MULHU, MULHSU, DIV, DIVU, REM, REMU, MULW, DIVW, DIVUW, REMW, REMUW) and, depending on project scope, the RV64B bit-manipulation extensions (e.g., CLZ, CTZ, BSET, BEXT). The design must be evaluated against the XH-1's 128-core replication factor: a unit that is only a small percentage of a single core's area consumes a large cumulative area across the die when replicated 128 times. The unit's latency impacts critical-path timing, and its throughput affects overall core IPC under integer-heavy workloads (cryptography, hashing, address arithmetic, HPC kernels).
|
||||
|
||||
## Research Question
|
||||
|
||||
What is the optimal MUL/DIV unit organization for a XH-1 core given that the design will be replicated 128 times, considering latency, throughput, area, power, and verification complexity trade-offs across in-order, out-of-order, and hybrid pipeline styles?
|
||||
|
||||
## Background
|
||||
|
||||
### RISC-V M Extension Semantics (RV64M)
|
||||
|
||||
FACT: The RISC-V M extension for RV64 defines the following instructions:
|
||||
- MUL, MULH, MULHU, MULHSU: 64×64→128 bit multiply (lower 64 or upper 64 of result, with sign-handling variations).
|
||||
- MULW: 32×32→32 (lower 32 of a 32×32→64 multiply, sign-extended result written to xrd).
|
||||
- DIV, DIVU: 64÷64 signed/unsigned quotient.
|
||||
- REM, REMU: 64÷64 signed/unsigned remainder.
|
||||
- DIVW, DIVUW, REMW, REMUW: 32÷32 signed/unsigned quotient/remainder, sign-extended.
|
||||
|
||||
FACT (per RISC-V ISA spec): The M extension defines the following corner-case behavior:
|
||||
- Division by zero:
|
||||
- `DIV` / `DIVU` / `DIVW` / `DIVUW`: quotient is `-1` (i.e., all bits set).
|
||||
- `REM` / `REMU` / `REMW` / `REMUW`: remainder equals the dividend.
|
||||
- Signed overflow (most-negative integer divided by `-1`):
|
||||
- `DIV` / `DIVW`: quotient equals the most-negative representable value.
|
||||
- `REM` / `REMW`: remainder equals zero.
|
||||
- For unsigned divide (`DIVU`/`REMU`/`DIVUW`/`REMUW`), overflow cannot occur; only the divide-by-zero rule applies.
|
||||
|
||||
The full 64×64→128-bit multiply requires a true 128-bit result datapath even when only the upper or lower 64 bits are returned. The MULHSU (signed × unsigned) variant requires explicit sign-handling of the signed operand before partial-product reduction; a pure unsigned Wallace/Dadda tree requires a front-end sign-extension / Booth-encoding stage.
|
||||
|
||||
NOTE: The W-suffixed instructions (MULW, DIVW, DIVUW, REMW, REMUW) are part of the **M extension** in RV64, not the base I extension. The base I extension's W variants are only the simple ALU ops (ADDW, SUBW, SLLW, SRLW, SRAW).
|
||||
|
||||
### Latency Reference Points (FACT, from published RISC-V implementations, see Sources)
|
||||
|
||||
- Rocket Chip (SiFive, in-order): MUL is unpipelined and configuration-dependent; commonly cited as 3- or 4-cycle unpipelined iterative multiplier. DIV is variable (radix-4 iterative). INSUFFICIENT EVIDENCE for a single canonical latency number without specifying the Rocket Chip configuration (StandardConfig vs MinimalConfig vs other).
|
||||
- BOOM v2/v3 (out-of-order): MUL pipelined; latency configuration-dependent. DIV variable.
|
||||
- XiangShan (Nanjing, out-of-order): MUL pipelined, DIV variable.
|
||||
- Ibex (lowRISC, in-order): MUL single-cycle combinational or short-pipeline; iterative DIV.
|
||||
- Arm Cortex-A77: per-instruction latencies are not publicly published by Arm. INSUFFICIENT EVIDENCE; no specific MUL/DIV latency numbers are cited from a primary source.
|
||||
|
||||
### Divide Algorithms
|
||||
|
||||
FACT: Four primary classes are considered here:
|
||||
1. **Digit-recurrence / shift-subtract (radix-2, radix-4, radix-16, radix-64 SRT)**: radix-2 takes 64 cycles worst case; radix-4 takes ~32–33 cycles; radix-16 takes ~16 cycles; radix-64 takes ~8–10 cycles (with significant area/complexity cost). Iterative, small-to-moderate area depending on radix.
|
||||
2. **Newton-Raphson reciprocal multiplication**: multiple iterations of a multiply-based refinement to compute the reciprocal, then a final correction multiply to produce the quotient. Latency depends on initial seed precision and convergence criteria.
|
||||
3. **Goldschmidt**: similar convergence behavior to Newton-Raphson, with a different iteration structure.
|
||||
4. **CORDIC-based and series-expansion dividers**: rotate-mode or series-expansion methods; rarely used for general-purpose integer divide due to overhead, but exist as alternative approaches.
|
||||
|
||||
The RISC-V M extension does not require IEEE-754 compliant division; integer rounding is sufficient, simplifying hardware.
|
||||
|
||||
## Existing Approaches
|
||||
|
||||
### A1. Iterative Shift-Subtract Divider (Radix-2, single divider only)
|
||||
|
||||
- **Datapath**: 128-bit (64 remainder + 64 divisor) shifter/subtractor, radix-2, 64-cycle worst case.
|
||||
- **DIV Latency**: 64 cycles.
|
||||
- **DIV Throughput**: 1 divide per 64 cycles (not pipelined).
|
||||
- **MUL coverage**: A1 defines a divider only; MUL is not part of this organization. Any combined MUL+radix-2-DIV system must be specified separately.
|
||||
- **Area (DIV only)**: Smallest divider; typically a few kGE plus control.
|
||||
- **Power**: Lowest of the divider options; minimal toggle rate per non-dividing cycle.
|
||||
- **Used in**: Some low-end in-order cores; some configurations of Rocket.
|
||||
|
||||
### A2. Pipelined Iterative Divider
|
||||
|
||||
- **Datapath**: same shift-subtract array as A1 but with pipeline registers at each iteration; latency fixed at 64 cycles, throughput 1/cycle.
|
||||
- **DIV Latency**: 64 cycles, fixed.
|
||||
- **DIV Throughput**: 1/cycle.
|
||||
- **Area**: ~3–5× A1 due to per-stage registers; significant.
|
||||
- **Power**: Higher clock-gating complexity; only worthwhile under sustained divide streams.
|
||||
|
||||
### A3. Newton-Raphson Divider
|
||||
|
||||
- **Datapath**: shared MUL unit(s), initial reciprocal seed ROM, multiplier used in iterative refinement and one final correction multiply.
|
||||
- **Latency breakdown**: seed table lookup (1 cycle) + N refinement multiplies (typically 2–3 iterations for 64-bit integer) + 1 final correction multiply. ASSUMPTION: 3 refinement iterations + 1 correction is typical for 64-bit; concrete iteration count depends on the seed precision and the desired final precision. INSUFFICIENT EVIDENCE for a single canonical iteration count without specifying the seed table and refinement schedule.
|
||||
- **Throughput**: one divide per N+1 MUL cycles (where N = refinement iterations), assuming the multiplier is dedicated to the divider.
|
||||
- **Area**: 1 reciprocal seed ROM + 1–2 MUL units (one of which can be shared with the main MUL datapath, at the cost of contention); large.
|
||||
- **Power**: Higher static and dynamic (multiplier active during divide refinement).
|
||||
- **Used in**: Some high-performance FPU designs for floating-point; less common for dedicated integer divide because the integer multiplier is large and the fixed iteration count does not beat pipelined shift-subtract on worst-case latency.
|
||||
|
||||
### A4. Radix-16 / Radix-64 SRT Divider
|
||||
|
||||
- **Datapath**: high-radix recurrence with a quotient-digit lookup table (PLA or ROM), redundant remainder representation. 16 or 64 bits processed per cycle.
|
||||
- **DIV Latency**: ~16 cycles (radix-16) or ~8–10 cycles (radix-64) for 64-bit operands.
|
||||
- **DIV Throughput**: 1/cycle if pipelined, or 1 per 16 / 8–10 cycles if iterative.
|
||||
- **Area**: large lookup table and complex datapath; PLA is a significant area contributor. Practical only in high-performance designs.
|
||||
- **Power**: high; many bits toggle per cycle.
|
||||
- **Used in**: high-end x86 integer dividers (e.g., Intel Haswell and later use a radix-16 / radix-32 shift-subtract divider for integer divide, not Newton-Raphson; Newton-Raphson in those designs is used in the floating-point unit).
|
||||
|
||||
### A5. Approximate / Lookup-Based Dividers (small operand ranges)
|
||||
|
||||
- Rarely used for 64-bit; not relevant to XH-1 unless a specific workload (e.g., 32-bit address arithmetic) dominates.
|
||||
|
||||
## Alternative Designs
|
||||
|
||||
### B1. Hybrid MUL + Sequential-Iterative DIV (Rocket-style)
|
||||
|
||||
- MUL: 64×64→128 pipelined in 1–2 stages.
|
||||
- DIV: radix-4 iterative, ~33 cycles (radix-4: 2 quotient bits per cycle, 32 cycles for quotient bits + finalization), blocking on the unit.
|
||||
- Latency hide: out-of-order core can issue subsequent independent ops; in-order core must stall.
|
||||
- Single divider per core, 1–3 MUL pipelined stages.
|
||||
|
||||
### B1'. Two MUL Pipelines + Shared Iterative DIV (wide-issue variant)
|
||||
|
||||
- Two pipelined MUL datapaths, one shared radix-4 DIV datapath.
|
||||
- MUL throughput: 2/cycle (sustained, independent operands).
|
||||
- DIV throughput: 1 per ~33 cycles, shared and blocking.
|
||||
- Area: ~1.8–2.0× B1 (two MUL pipelines) + same DIV.
|
||||
- Useful for 2-issue or wider cores; the recommendation in the original document of "replicate to 2 MUL pipelines for wide issue" is a degenerate case of this option.
|
||||
|
||||
### B2. Combined MUL/MAC (Multiply-Accumulate) Unit
|
||||
|
||||
- 64×64→128 with an additional accumulator register; supports future RV MAC extensions (proposed in some bitmanip drafts).
|
||||
- Adds ~10–15% area over pure MUL (ASSUMPTION; INSUFFICIENT EVIDENCE without RTL).
|
||||
- Useful for DSP, ML, and HPC; 128-core replication benefits workloads with matrix-style kernels.
|
||||
|
||||
### B3. Configurable MUL Width (32-bit fast path)
|
||||
|
||||
- Detect when both operands are zero-extended or sign-extended from 32 bits (e.g., results of MULW, DIVW, DIVUW, REMW, REMUW, or explicit 32-bit-zero-extended operands); use a 32×32→64 fast multiplier.
|
||||
- Common in commercial cores (Arm, x86).
|
||||
- Branch-predictor / decoder pre-classifies operand width.
|
||||
- Open question: whether a "fast-path 32-bit MUL" justifies the verification cost given the MULW instruction in RV64M already gives 32-bit semantics with a 32-bit result.
|
||||
|
||||
### B4. Bypassable Output with Operand Width Detection
|
||||
|
||||
- Always implement full 64×64→128; the decoder examines the instruction to forward only the required 64 bits.
|
||||
- Verification: a single source of truth (full 128-bit datapath), reducing dual-mode bugs.
|
||||
- Disadvantage: no area saving.
|
||||
|
||||
### B5. Skip-on-Zero / Divide-Cancellation Optimizations
|
||||
|
||||
- Detect zero dividend (quotient is zero, remainder is dividend) and divide-by-one (quotient is dividend, remainder is zero) at the front end and forward the result without entering the iterative loop.
|
||||
- Saves latency in the common case for some workloads; trivial area overhead.
|
||||
- Composable with any divider organization.
|
||||
|
||||
### B6. Software Divide-by-Constant Transformation (cross-cutting)
|
||||
|
||||
- Compilers transform division by a runtime constant into a multiply-by-reciprocal sequence. The hardware DIV is then needed only for division by variables.
|
||||
- Reduces effective DIV frequency significantly for workloads with constant denominators; affects hardware sizing decisions.
|
||||
|
||||
## Comparison
|
||||
|
||||
| Design | MUL Latency | MUL Tput | DIV Latency | DIV Tput | Rel. Area (MUL+DIV) | Power (active) | Verification Complexity |
|
||||
|--------|-------------|----------|-------------|----------|----------------------|----------------|--------------------------|
|
||||
| A1 (radix-2 iter. DIV only) | not defined in A1 | not defined in A1 | 64 cycles | 1/64 cycles | INSUFFICIENT EVIDENCE (combined) | Lowest (DIV only) | Low |
|
||||
| A2 (pipelined iter. DIV) | not defined in A2 | not defined in A2 | 64 cycles | 1/cycle | INSUFFICIENT EVIDENCE (combined) | Med | Med |
|
||||
| A3 (Newton-Raphson) | 1-cycle | 1/cycle | depends on iteration count (see A3) | 1/(N+1) cycles | high (see A3) | High | High |
|
||||
| A4 (Radix-16/64 SRT) | not defined in A4 | not defined in A4 | 8–16 cycles | 1/cycle if pipelined | high | High | High |
|
||||
| B1 (Rocket-style) | 1–2 cycle | 1/cycle | ~33 cycle | 1/33 cycle | baseline | Low–Med | Low–Med |
|
||||
| B1' (2× MUL + 1× shared DIV) | 1–2 cycle | 2/cycle | ~33 cycle | 1/33 cycle (shared) | ~1.8–2.0× B1 | Med | Med |
|
||||
| B2 (+ MAC) | 1–2 cycle | 1/cycle | ~33 cycle | 1/33 cycle | +10–15% over B1 (ASSUMPTION) | Med | Med |
|
||||
| B3 (32-bit fast) | <1 cycle (32-bit path) | 1/cycle | ~33 cycle | 1/33 cycle | varies | Low | Med–High (dual mode) |
|
||||
| B4 (full only) | 1–2 cycle | 1/cycle | ~33 cycle | 1/33 cycle | same as B1 | Med | Lowest |
|
||||
| B5 (skip-on-zero) | n/a | n/a | reduced in common case | same as base | negligible overhead | n/a | Low |
|
||||
|
||||
ASSUMPTION: Relative area figures are estimates based on published reference designs cited in the Sources section. Actual XH-1 synthesis numbers are INSUFFICIENT EVIDENCE pending RTL implementation. Specific quantitative area multipliers (e.g., "1.4×", "2.5×", "3.0×") from the prior revision are removed in favor of qualitative ordering.
|
||||
|
||||
## Advantages
|
||||
|
||||
### A1 / B1
|
||||
- Smallest area; lowest per-core replication cost.
|
||||
- Lowest static and dynamic power in the typical case (most cycles no MUL/DIV active).
|
||||
- Simpler verification; one mode of operation.
|
||||
- Well-understood reference implementation (Rocket, Ibex) for cross-checking.
|
||||
|
||||
### A4 (High-Radix SRT)
|
||||
- Lowest DIV latency among the iterative-style options; competitive with A3 on a single divide.
|
||||
- Pipelined variant gives 1/cycle DIV throughput.
|
||||
|
||||
### A3 (Newton-Raphson)
|
||||
- Low latency for sequential divides — important for cryptography (RSA, ECC modular reduction) and HPC.
|
||||
- Amortizes multiplier cost if MAC (B2) is also desired.
|
||||
- Disadvantage: requires multiple refinement iterations; the integer multiplier is large and contention with the main MUL datapath is a concern.
|
||||
|
||||
### B2 (MAC)
|
||||
- Enables future-proofing for proposed bitmanip and MAC extensions.
|
||||
- Helpful for matrix multiplication kernels running across 128 cores.
|
||||
|
||||
### B5 (Skip-on-Zero)
|
||||
- Negligible area; reduces effective DIV latency for common cases.
|
||||
|
||||
## Disadvantages
|
||||
|
||||
### A1 / B1
|
||||
- Variable DIV latency complicates scheduling in an in-order core; in-order cores stall ~33 cycles per divide.
|
||||
- Worst-case latency hits interrupt latency, branch-target computation if the MUL/DIV is on a branch path.
|
||||
- Not a fit for sustained divide-bound workloads (e.g., modulo arithmetic in bignum).
|
||||
|
||||
### A3
|
||||
- Area at 128-core replication is severe; the multiplier is one of the largest blocks in a typical core.
|
||||
- Power: a 64×64 multiplier running 1/cycle is one of the highest-power blocks in a typical core.
|
||||
- Verification: Newton iteration precision (must converge to within 1 ULP across all operands) is notoriously subtle.
|
||||
|
||||
### A4
|
||||
- Large lookup table (PLA or ROM); area and power dominated by the table.
|
||||
- Verification: complex quotient-digit selection logic.
|
||||
- Not commonly used outside high-end commercial designs.
|
||||
|
||||
### B3
|
||||
- Verification: dual datapath (32-bit and 64-bit) must be exhaustively cross-checked.
|
||||
- The classification logic (operand-width detection) is itself a source of corner-case bugs.
|
||||
- The MULW instruction in RV64M already provides 32-bit MUL semantics, so a "fast path" for non-MULW 32-bit operands must be justified against MULW's existing semantics.
|
||||
|
||||
### B2
|
||||
- Adds a third writeback port or accumulator path; competes with the FPU/LSU for writeback slots in a multi-issue core.
|
||||
|
||||
### B5
|
||||
- Only helps the specific cases of zero dividend or divisor ±1; other optimizations (e.g., division by small powers of two) are already handled by the base I extension's shift instructions.
|
||||
|
||||
## XH-1 Considerations
|
||||
|
||||
PROPOSAL: The XH-1 core should adopt **B1 (Hybrid MUL + Iterative DIV)** as the baseline, with the following parameters, pending RTL validation:
|
||||
- **MUL**: 64×64→128, 1-cycle pipelined (1 stage of pipeline registers), throughput 1/cycle. For 2-issue or wider cores, scale to B1' (two MUL pipelines sharing one DIV).
|
||||
- **DIV**: Radix-4 shift-subtract, ~33 cycles worst case (32 cycles for quotient bits + finalization), blocking, non-pipelined. (The prior revision's "33–35 cycles worst case" and the "33–64 cycle" range are reconciled here: 33 is the radix-4 bound; 64 corresponds to radix-2, which is a different algorithm choice.)
|
||||
- **REM**: Reuse the DIV datapath; remainder is a by-product.
|
||||
- **Bypass paths**: From MUL output directly to subsequent dependent ALU/branch in the next cycle.
|
||||
- **B5 (skip-on-zero)**: implement at the front end of the DIV unit; negligible overhead.
|
||||
- **Operand-width detection**: Defer B3 to a future revision; not justified at the 128-core replication level given the added verification cost and the existence of MULW for the 32-bit case.
|
||||
|
||||
ASSUMPTION: A 1-cycle MUL latency is achievable in the target process. INSUFFICIENT EVIDENCE on the XH-1 target process node and clock period. PROPOSAL: Validate via synthesis at the target corner before committing.
|
||||
|
||||
PROPOSAL: The MUL/DIV unit sits on the integer execution lane, sharing the issue queue with ALU and branch. It has its own reservation station (or issue-slot tag) so the issue logic can distinguish short-latency ALU (1 cycle) from long-latency MUL (1) and variable-latency DIV (~33).
|
||||
|
||||
PROPOSAL: For the in-order case, the MUL/DIV unit is non-blocking on MUL (1-cycle latency) but blocking on DIV; the in-order front-end stalls on a structural hazard only when a second DIV is issued before the first completes.
|
||||
|
||||
PROPOSAL: For an out-of-order XH-1 core, the MUL/DIV unit exposes its variable DIV latency through the wakeup/select logic so that dependent instructions are replayed or re-issued with the correct ready signal.
|
||||
|
||||
PROPOSAL: The B-extension unit (if RV64B is implemented) shares **operand muxes, sign-handling logic, and bit-level muxes** with the MUL/DIV unit but does **not** share the Wallace/Dadda compressor tree. Operations like CLZ, CTZ, BSET, BEXT operate on individual bits or small bit-fields and do not naturally map onto a Wallace-tree multiplier datapath. Any apparent sharing is at the operand-fetch and writeback layers, not the core arithmetic.
|
||||
|
||||
## 128-Core Scalability
|
||||
|
||||
The 128-core replication factor is the dominant cost driver for the MUL/DIV unit.
|
||||
|
||||
PROPOSAL: At 128 cores, every MUL unit must be **power-gateable independently** so that cores not in use (DPM, dark-silicon) can be fully clock- and power-gated, including the MUL/DIV block.
|
||||
|
||||
OPEN QUESTION: What is the actual MUL/DIV area share of the XH-1 core? Published RISC-V references suggest a typical MUL unit occupies a few percent of a high-performance core's area, with iterative DIV adding additional area. The prior revision's "1–3% FACT" and "1 MGE/core total" baseline are removed here as unsourced. INSUFFICIENT EVIDENCE on the XH-1-specific area share without synthesis.
|
||||
|
||||
OPEN QUESTION: Is the MUL/DIV area share dominated by the MUL datapath, the DIV datapath, the operand muxing, or the bypass network? Targeted synthesis required.
|
||||
|
||||
INSUFFICIENT EVIDENCE: The XH-1 die-area budget, process node, and clock period are not established in this document.
|
||||
|
||||
PROPOSAL: The MUL/DIV unit should be **identical across all 128 cores** (no per-core specialization) to simplify verification and physical design. The replication should be a Verilog/SystemVerilog generate block driven by a single template.
|
||||
|
||||
CONSIDERATION (clock and timing at 128× replication): If the XH-1 die includes a global clock-distribution network, the MUL/DIV critical path determines the local clock skew tolerance. A pipelined 1-cycle MUL has a short critical path (one 64-bit adder-equivalent) which is favorable; a non-pipelined iterative divider has a 64-bit adder in its loop, which is the cycle-time limiter. The clock-skew analysis should be revisited at the integration level once the XH-1 clock tree is defined.
|
||||
|
||||
CONSIDERATION (cross-core aggregate throughput): Under B1, each core's blocking DIV delivers ~1/33 divides per cycle. Across 128 cores, the aggregate is ~3.9 divides per cycle in the best case (all cores dividing simultaneously), which is the upper bound. Real workloads do not exhibit this worst case; the relevant metric is the per-core latency, not aggregate. INSUFFICIENT EVIDENCE on whether the XH-1 DPM or interconnect imposes a global cap on simultaneous divide activity; this is a system-level question outside the MUL/DIV unit's scope.
|
||||
|
||||
CONSIDERATION (operand distribution and interconnect): Replicating a Wallace tree 128× implies 128 sets of wide operand buses to/from the register file. The interconnect / operand-routing network cost scales with the MUL operand width and the number of cores. PROPOSAL: include operand-routing overhead in the area estimate, not just the MUL/DIV datapath itself.
|
||||
|
||||
## Performance Considerations
|
||||
|
||||
PROPOSAL: **MUL throughput of 1/cycle is non-negotiable** for a modern XH-1 core. B1 satisfies this with a single 1-stage pipelined multiplier. B1' extends to 2/cycle for wide-issue.
|
||||
|
||||
PROPOSAL: **DIV throughput of 1/~33 cycles (blocking) is acceptable** for most integer workloads. Workloads that require more (e.g., modular arithmetic in cryptography) will suffer; mitigation is left to software (see Software Considerations).
|
||||
|
||||
ASSUMPTION: XH-1 target workloads include a mix consistent with Embench / SPECint-class profiles, where MUL/DIV instructions are a small fraction of dynamic instruction count. INSUFFICIENT EVIDENCE on the actual XH-1 target workload mix and on the specific dynamic-instruction share of MUL/DIV. The prior revision's "<2% MUL/DIV" and "<0.5% DIV/REM" FACT claims are removed as unsourced.
|
||||
|
||||
OPEN QUESTION: Does XH-1 target HPC or ML workloads where MUL/DIV is a larger share of the instruction mix? If yes, the analysis shifts toward A3/B2 or a pipelined DIV (A2).
|
||||
|
||||
## Area Considerations
|
||||
|
||||
PROPOSAL: Budget the MUL/DIV unit at a small single-digit percentage of single-core area for the B1 design, pending synthesis. The exact percentage is INSUFFICIENT EVIDENCE.
|
||||
|
||||
OPEN QUESTION: Is the XH-1 core area budget (without MUL/DIV) known? The MUL/DIV area share can be re-expressed as a percentage of the entire 128-core die only when this is established.
|
||||
|
||||
ASSUMPTION: A1-style radix-2 DIV is the smallest practical divider; A3 (Newton) and A4 (high-radix SRT) are larger, with A4 typically the largest. INSUFFICIENT EVIDENCE on the XH-1-specific gate-count budget.
|
||||
|
||||
## Power and Energy Considerations
|
||||
|
||||
FACT: A 64×64→128 multiplier toggles a large number of bits per cycle when active; its dynamic power is proportional to operand activity factor. Power-gating the multiplier when idle is essential.
|
||||
|
||||
PROPOSAL: The MUL/DIV unit should support **fine-grained clock gating** at the operand-mux boundary so that when no MUL/DIV is in flight, the entire unit is clock-gated to 0 toggle rate.
|
||||
|
||||
PROPOSAL: The MUL pipeline register should be a clock-gated scan flop, not a latch, to simplify DFT.
|
||||
|
||||
OPEN QUESTION: Does the XH-1 power budget allow all 128 MUL/DIV units to be active simultaneously? If not, per-core DPM must enforce that no more than N MUL/DIV units are in active divide at once. This is a system-level power-management question, not a unit-level one, but it constrains the unit's design (e.g., should the MUL unit support a "low-power divide" mode that takes 64 cycles at half frequency?).
|
||||
|
||||
INSUFFICIENT EVIDENCE: Specific per-MUL or per-DIV energy numbers for the XH-1 process are not established. The prior revision's "single-digit pJ in 7 nm" claim is removed as unsourced; per-MUL energy in advanced processes is implementation-dependent and varies by an order of magnitude or more depending on architecture and clock frequency.
|
||||
|
||||
## Implementation Considerations
|
||||
|
||||
PROPOSAL: Implement MUL as a **Wallace/Dadda tree of 4:2 compressors** feeding a final 128-bit carry-propagate adder. The 1-cycle latency budget must accommodate the entire critical path from operand register → compressor tree → CPA → output register. ASSUMPTION: a Wallace/Dadda tree for 64×64 partial products is achievable in one cycle at the target clock period. The specific tree depth and CPA depth are implementation-dependent and not asserted as fixed numbers here (the prior revision's "Tree depth 6–7; final CPA ~6 gates deep" is removed as unsourced and implementation-specific).
|
||||
|
||||
PROPOSAL: For MULHSU (signed × unsigned), include a sign-handling stage (e.g., Booth encoding of the signed operand) before the partial-product reduction tree. A pure unsigned Wallace/Dadda tree does not handle the signed × unsigned case without this front-end.
|
||||
|
||||
PROPOSAL: Implement DIV as a **radix-4 shift-subtract** with restoring on negative remainder. 32 cycles for quotient bits + 1 cycle for finalization = ~33 cycles total. The remainder is the "running remainder" register after the 33rd cycle; the quotient is the shifted-out bits. REM is selected by muxing the appropriate output.
|
||||
|
||||
PROPOSAL: B5 (skip-on-zero): add a front-end detector on the DIV operands that forwards the result directly for divisor = ±1 or dividend = 0, bypassing the iterative loop. Trivial area; reduces effective DIV latency for common cases.
|
||||
|
||||
PROPOSAL: Parameterize the MUL width with a Verilog `parameter WIDTH=64` so that the same RTL can be synthesized for 32-bit-only cores (test variant, debug, or soft-IP reuse) without manual rework. Note that WIDTH=32 does not by itself support the MULW instruction semantics in RV64 (which performs a 32×32→64 multiply and sign-extends the 32-bit result); an MULW-specific path is required for RV64.
|
||||
|
||||
## Verification Considerations
|
||||
|
||||
FACT: MUL and DIV are among the most heavily tested instructions in any RISC-V core; their verification is dominated by **corner-case operand pairs** (e.g., 0, 1, -1, 2, -2, 0x7FFF…, 0x8000…, 0xFFFF…, MSB=1 transitions). For DIV/REM, the corner cases include the division-by-zero and signed-overflow rules defined in the ISA spec.
|
||||
|
||||
PROPOSAL: Adopt a **directed-random + constrained-random** verification flow using a golden reference model:
|
||||
- **Reference MUL**: SystemVerilog `bit [127:0]` (or DPI-C to a software bigint). Compare lower-64 and upper-64 separately. MULW, MULH, MULHU, MULHSU all share the 128-bit reference.
|
||||
- **Reference DIV**: Software bigint that produces (quotient, remainder) per the RISC-V M-extension spec including the division-by-zero and signed-overflow rules for both quotient-producing and remainder-producing instructions.
|
||||
- **Coverage**: 100% of the corner cases (above), plus cross-product of small operand values (±1, ±2, ±2^32, ±2^63, ±2^64-1), plus randomized large operands.
|
||||
- **Regression list size**: The prior revision's "64 hand-crafted corner cases" is removed as a specific number; the regression list should be sized to cover the documented corner cases and is grown as bugs are found. INSUFFICIENT EVIDENCE for a canonical count.
|
||||
|
||||
PROPOSAL: At the 128-core replication level, **per-core functional verification** of the MUL/DIV unit is unnecessary if the per-core RTL is identical and constrained-random coverage is closed at the template level (via a SystemVerilog generate block). This covers functional equivalence. **Per-core physical / timing verification is not bypassed**: timing, DFT, and physical-design closure are verified at the integration level on a representative core and assumed replicated, with explicit per-die variation analysis as required by the XH-1 physical-design flow.
|
||||
|
||||
PROPOSAL: If the XH-1 verification flow includes formal property checking, the DIV correctness property ("quotient × divisor + remainder == dividend" within the RISC-V M extension's special cases) is a clean formal target and should be specified. INSUFFICIENT EVIDENCE on whether formal property checking is in scope.
|
||||
|
||||
## Software Considerations
|
||||
|
||||
PROPOSAL: Document the MUL/DIV latencies (1 cycle MUL, ~33 cycle DIV) in the XH-1 ABI / ISA reference manual so that compiler backends and library authors can schedule around the long-latency DIV.
|
||||
|
||||
OPEN QUESTION: Does the XH-1 ABI / linker convention include a software-emulated division routine for code that cannot tolerate the ~33-cycle blocking latency? If yes, the hardware DIV can be a single issue-slot, not a deeply pipelined one. The compiler can also apply divide-by-constant transformations (B6) to reduce effective hardware DIV frequency.
|
||||
|
||||
PROPOSAL: The MUL/DIV unit should raise no architectural exception; the RISC-V M extension defines all corner cases architecturally. This simplifies the trap logic.
|
||||
|
||||
PROPOSAL: For 128-core use cases, the MUL/DIV unit is the same in every core; the OS scheduler does not need to be aware of its microarchitectural details. Workload distribution across cores is independent of MUL/DIV design.
|
||||
|
||||
## Recommendation
|
||||
|
||||
**RECOMMENDATION: Adopt B1 — a hybrid MUL unit (1-cycle pipelined, full 64×64→128) combined with a radix-4 iterative, non-pipelined divider (~33 cycles), sharing no logic. Add B5 (skip-on-zero) at the DIV front end. Place the unit on the integer execution lane. Support full per-unit power and clock gating for 128-core DPM. For wide-issue cores, scale to B1' (two MUL pipelines sharing one DIV).**
|
||||
|
||||
Rationale:
|
||||
1. Matches the dominant MUL workload profile (frequent, 1-cycle latency needed).
|
||||
2. Keeps per-core area small, manageable at 128× replication.
|
||||
3. Avoids the verification burden of B3 (dual datapath) and the area burden of A3 (Newton-Raphson) and A4 (high-radix SRT).
|
||||
4. Aligns with proven reference designs (Rocket, Ibex) reducing architectural risk; this is an architectural-pattern alignment, not a claim of identical pipeline structure, since Rocket and BOOM use different organizations.
|
||||
5. B5 is a near-free improvement to the common case.
|
||||
|
||||
This recommendation is **conditional on**:
|
||||
- The XH-1 target process supporting a 1-cycle 64×64→128 MUL critical path (must be validated by synthesis at the target corner; INSUFFICIENT EVIDENCE without target process specification).
|
||||
- XH-1 workload mix not being dominated by sequential DIV/REM (cryptography). If it is, escalate to A4 (radix-16 SRT) or a pipelined iterative divider.
|
||||
- The 128-core replication budget tolerating the cumulative MUL/DIV area; this requires a known single-core area budget, which is INSUFFICIENT EVIDENCE.
|
||||
|
||||
If any of these conditions fails, re-open the design.
|
||||
|
||||
## Confidence
|
||||
|
||||
**Medium-High** for B1 as the baseline choice. **Low** for specific area, power, and energy numbers — those are estimates based on published RISC-V reference designs, not on XH-1 synthesis data. **Low** for any claim about the XH-1 target process, workload mix, or die size (these are INSUFFICIENT EVIDENCE). The corrected version removes specific unsourced quantitative claims and demotes several prior FACTs to ASSUMPTION or INSUFFICIENT EVIDENCE.
|
||||
|
||||
## Open Questions
|
||||
|
||||
1. What is the XH-1 target process node and clock period? This constrains MUL critical-path design.
|
||||
2. What is the XH-1 target workload mix? HPC, ML, cryptography, or general-purpose?
|
||||
3. Is the XH-1 core in-order, out-of-order, or hybrid? (The repository context does not yet establish this fact per the system policy.)
|
||||
4. Does XH-1 implement RV64B (bit-manipulation) extensions? The B extension does not naturally share the Wallace-tree multiplier datapath; any sharing is at the operand-mux and writeback layers, not the compressor tree.
|
||||
5. What is the XH-1 single-core area budget? What is the die-level area budget for the 128-core fabric?
|
||||
6. Does XH-1 use a per-core DPM scheme that requires the MUL/DIV unit to support fine-grained power gating?
|
||||
7. Is the MUL/DIV unit expected to support a future RV MAC extension, or is the project strictly RV64IM?
|
||||
8. Is there a software-emulated division in the XH-1 libc / ABI that would tolerate a long-latency hardware DIV? Will the compiler apply divide-by-constant transformations (B6)?
|
||||
9. Does the XH-1 verification flow include formal property verification, and will it be applied to the MUL/DIV unit's "quotient × divisor + remainder == dividend" property?
|
||||
10. For wide-issue cores, is B1' (two MUL pipelines + one shared DIV) the target, or is single-MUL B1 sufficient?
|
||||
11. What is the XH-1 interconnect / operand-routing cost of replicating a wide MUL operand bus 128 times? Should this overhead be included in the MUL/DIV unit's area budget?
|
||||
|
||||
## Sources
|
||||
|
||||
- Waterman, Asanović, et al. (eds.), *The RISC-V Instruction Set Manual, Volume I: User-Level ISA, Document Version 20191213*, Chapter 7 (M Extension). RISC-V International. (Canonical ISA reference; defines the M-extension corner-case rules for DIV/REM/DIVU/REMU/DIVW/REMW/DIVUW/REMUW and MULW.)
|
||||
- Asanović et al., "The Rocket Chip Generator," EECS Department, UC Berkeley, Technical Report UCB/EECS-2016-17, 2016. (Reference for Rocket's MUL/DIV organization; configuration-dependent.)
|
||||
- Celio, Patterson, Asanović, "BOOM v2: a superscalar out-of-order processor," 2017. (Reference for BOOM's MUL/DIV design.)
|
||||
- Ibex documentation, lowRISC. (Reference for the in-order baseline.)
|
||||
- Chen et al., "XiangShan: An Open-Source High-Performance RISC-V Core," 2022. (Reference for high-performance RISC-V MUL/DIV design.)
|
||||
- Parhami, *Computer Arithmetic: Algorithms and Hardware Designs*, Oxford University Press. (Reference for shift-subtract, SRT, and Newton-Raphson divide algorithms.)
|
||||
- Flynn, Oberman, *Advanced Computer Arithmetic Design*. (Reference for high-radix divider design.)
|
||||
- Weste, Harris, *CMOS VLSI Design: A Circuits and Systems Perspective*, 4th ed. (Reference for Wallace/Dadda multiplier design and critical-path analysis; does not by itself establish a specific 1-cycle latency in any particular process.)
|
||||
|
||||
INSUFFICIENT EVIDENCE: No XH-1-specific RTL, synthesis, or benchmark measurements are available at the time of this document. All quantitative claims about XH-1 area, power, energy, and timing are INSUFFICIENT EVIDENCE pending RTL implementation and target-process specification. Cortex-A77 per-instruction latencies are not publicly published by Arm and are INSUFFICIENT EVIDENCE from primary sources. The Intel Haswell integer divider is a radix-16 (or higher) shift-subtract divider, not a Newton-Raphson divider; Newton-Raphson is used in Haswell's floating-point unit.
|
||||
+291
File diff suppressed because one or more lines are too long
+431
@@ -0,0 +1,431 @@
|
||||
# MUL/DIV Unit
|
||||
|
||||
## Status
|
||||
|
||||
Stub — initial scoping document. No XH-1 implementation decisions are yet committed. This revision corrects factual errors identified in review, removes unsupported quantitative claims, reconciles internal contradictions, and adds missing alternatives.
|
||||
|
||||
## Abstract
|
||||
|
||||
This document investigates the design of the integer multiply/divide (MUL/DIV) execution unit for the XH-1 core. The unit must implement the RV64M extension (MUL, MULH, MULHU, MULHSU, DIV, DIVU, REM, REMU, MULW, DIVW, DIVUW, REMW, REMUW) and, depending on project scope, the RV64B bit-manipulation extensions (e.g., CLZ, CTZ, BSET, BEXT, CLMUL, CLMULH, CLMULR). The design must be evaluated against the XH-1's 128-core replication factor: a unit that is only a small percentage of a single core's area consumes a large cumulative area across the die when replicated 128 times. The unit's latency impacts critical-path timing, and its throughput affects overall core IPC under integer-heavy workloads (cryptography, hashing, address arithmetic, HPC kernels).
|
||||
|
||||
## Research Question
|
||||
|
||||
What is the optimal MUL/DIV unit organization for an XH-1 core given that the design will be replicated 128 times, considering latency, throughput, area, power, energy, and verification complexity trade-offs across in-order, out-of-order, and hybrid pipeline styles?
|
||||
|
||||
## Background
|
||||
|
||||
### RISC-V M Extension Semantics (RV64M)
|
||||
|
||||
**FACT** (per RISC-V ISA spec, Volume I, Chapter 7): The RISC-V M extension for RV64 defines the following instructions:
|
||||
|
||||
- MUL, MULH, MULHU, MULHSU: 64×64→128 bit multiply (lower 64 or upper 64 of result, with sign-handling variations).
|
||||
- MULW: 32×32→32 bit product, then sign-extended to 64 bits and written to `rd`.
|
||||
- DIV, DIVU: 64÷64 signed/unsigned quotient.
|
||||
- REM, REMU: 64÷64 signed/unsigned remainder.
|
||||
- DIVW, DIVUW, REMW, REMUW: 32÷32 signed/unsigned quotient/remainder, sign-extended to 64 bits and written to `rd`.
|
||||
|
||||
**FACT** (per RISC-V ISA spec, Volume I, Chapter 7): The M extension defines the following corner-case behavior:
|
||||
|
||||
- Division by zero:
|
||||
- `DIV` / `DIVW`: quotient is `2^XLEN − 1` (all bits set).
|
||||
- `DIVU` / `DIVUW`: quotient is `2^XLEN − 1` (all bits set; the spec expresses this as the unsigned maximum, not as signed −1).
|
||||
- `REM` / `REMW`: remainder equals the dividend.
|
||||
- `REMU` / `REMUW`: remainder equals the dividend.
|
||||
- Signed overflow (most-negative integer divided by −1):
|
||||
- `DIV` / `DIVW`: quotient equals the dividend (i.e., the most-negative representable value `2^(XLEN−1)`).
|
||||
- `REM` / `REMW`: remainder equals zero.
|
||||
- For unsigned divide (`DIVU` / `REMU` / `DIVUW` / `REMUW`), the only defined special case is division by zero; the dividend / `−1` overflow case does not apply because the operands are unsigned.
|
||||
|
||||
**NOTE**: A full 64×64→128-bit multiply requires a true 128-bit result datapath even when only the upper or lower 64 bits are returned.
|
||||
|
||||
**NOTE**: MULHSU is implementable either (a) by sign-extending the signed operand and zero-extending the unsigned operand to 128 bits and using a signed multiplier, or (b) by using a modified-Booth-encoded signed multiplier with the unsigned operand zero-extended. Either approach requires explicit sign handling before the partial-product reduction tree; a pure unsigned Wallace/Dadda tree without a sign-handling front-end does not implement MULHSU correctly.
|
||||
|
||||
**NOTE**: The W-suffixed instructions (MULW, DIVW, DIVUW, REMW, REMUW) are part of the **M extension** in RV64, not the base I extension. The base I extension's W variants are only the simple ALU ops (ADDW, SUBW, SLLW, SRLW, SRAW).
|
||||
|
||||
**NOTE**: The carry-less multiply instructions CLMUL, CLMULH, and CLMULR are part of the standard Zbc extension (commonly grouped under the umbrella "B" extension in some profiling). They are not part of M. They require a different datapath (AND-tree with XOR reduction, no carry propagation) and are not the subject of this document except where they interact with operand muxes / writeback.
|
||||
|
||||
### Latency Reference Points
|
||||
|
||||
**FACT (Rocket Chip, UC Berkeley generator)**: Rocket Chip is in-order, and the MUL/DIV unit is configuration-dependent across Rocket's `Configs.scala` parameter set. In `RocketCoreConfig` (commonly cited as the default), the multiplier is pipelined with multiple pipeline stages (multi-cycle iterative, new operation accepted per cycle) and the divider is a radix-4 iterative divider. The exact stage counts and latencies are determined by parameters in `Configs.scala` and are not a single canonical value. INSUFFICIENT EVIDENCE for a specific latency number without naming the configuration.
|
||||
|
||||
**FACT (BOOM v2/v3, UC Berkeley)**: Out-of-order superscalar. MUL is pipelined (latency configuration-dependent). DIV is variable-latency, non-pipelined.
|
||||
|
||||
**FACT (XiangShan, open-source OoO RISC-V)**: MUL is pipelined; DIV is variable-latency, non-pipelined.
|
||||
|
||||
**FACT (Ibex, lowRISC)**: In-order. MUL is implemented as a single-cycle or short-pipeline combinational multiplier in some configurations; iterative DIV.
|
||||
|
||||
**INSUFFICIENT EVIDENCE**: Arm Cortex-A77 per-instruction MUL/DIV latencies are not publicly published by Arm. No specific numbers are cited from a primary source.
|
||||
|
||||
**INSUFFICIENT EVIDENCE**: Intel Haswell integer divider internal radix (radix-16 vs. radix-32) is not established from publicly verifiable primary sources. The design is widely reported to be a high-radix shift-subtract divider rather than a Newton-Raphson divider, but the specific radix is INSUFFICIENT EVIDENCE.
|
||||
|
||||
### Divide Algorithms
|
||||
|
||||
**FACT**: Four primary classes are considered here:
|
||||
|
||||
1. **Digit-recurrence / shift-subtract (radix-2, radix-4, radix-16, radix-64 SRT)**: radix-2 takes 64 cycles worst case; radix-4 takes ~32–33 cycles; radix-16 takes ~16 cycles; radix-64 takes ~8–10 cycles (with significant area/complexity cost). Iterative, small-to-moderate area depending on radix.
|
||||
2. **Newton-Raphson reciprocal multiplication**: multiple iterations of a multiply-based refinement to compute the reciprocal, then a final correction multiply to produce the quotient. Latency depends on initial seed precision and convergence criteria.
|
||||
3. **Goldschmidt**: similar convergence behavior to Newton-Raphson, with a different iteration structure.
|
||||
4. **CORDIC-based and series-expansion dividers**: rarely used for general-purpose integer divide due to overhead; exist as alternative approaches.
|
||||
|
||||
**FACT**: The RISC-V M extension does not require IEEE-754 compliant division; integer rounding is sufficient, simplifying hardware.
|
||||
|
||||
## Existing Approaches
|
||||
|
||||
### A1. Iterative Shift-Subtract Divider (Radix-2)
|
||||
|
||||
- **Datapath**: 128-bit (64 remainder + 64 divisor) shifter/subtractor, radix-2, 64-cycle worst case.
|
||||
- **DIV Latency**: 64 cycles.
|
||||
- **DIV Throughput**: 1 divide per 64 cycles (not pipelined).
|
||||
- **MUL coverage**: A1 defines a divider only; MUL is not part of this organization.
|
||||
- **Area (DIV only)**: Smallest divider; typically a few kGE plus control. **INSUFFICIENT EVIDENCE** for a specific gate count.
|
||||
- **Power**: Lowest of the divider options when idle; minimal toggle rate per non-dividing cycle.
|
||||
- **Used in**: Some low-end in-order cores; some configurations of Rocket Chip with smaller radix.
|
||||
|
||||
### A2. Pipelined Iterative Divider
|
||||
|
||||
- **Datapath**: Two distinct subclasses must be distinguished:
|
||||
- **(A2a) Pipelined iterative loop**: a single shift-subtract array with pipeline registers inserted at one or more points within the iterative loop, allowing a new operation to enter the loop every cycle after the pipeline is filled. Latency remains 64 cycles; throughput is 1 per cycle after fill.
|
||||
- **(A2b) Fully unrolled divider**: 64 shift-subtract stages with pipeline registers between every stage, giving latency 64 cycles and throughput 1 per cycle from the first cycle. Area is roughly 64× the A1 datapath.
|
||||
- **DIV Latency**: 64 cycles.
|
||||
- **DIV Throughput**: 1 per cycle (after fill for A2a; from cycle 1 for A2b).
|
||||
- **Area**: A2a is ~2–3× A1 (a few extra pipeline registers); A2b is roughly 64× A1 and is rarely used.
|
||||
- **Power**: Higher toggle rate than A1; only worthwhile under sustained divide streams.
|
||||
- **Used in**: Rare; mostly in high-throughput streaming dividers (DSP). Uncommon in general-purpose cores.
|
||||
|
||||
### A3. Newton-Raphson Divider
|
||||
|
||||
- **Datapath**: shared MUL unit(s), initial reciprocal seed ROM, multiplier used in iterative refinement and one final correction multiply.
|
||||
- **Latency breakdown**: seed table lookup (1 cycle) + N refinement multiplies (typically 2–3 iterations for 64-bit integer) + 1 final correction multiply. **ASSUMPTION**: 3 refinement iterations + 1 correction is typical for 64-bit; the concrete iteration count depends on seed precision and convergence criteria. **INSUFFICIENT EVIDENCE** for a single canonical iteration count without specifying the seed table and refinement schedule.
|
||||
- **Throughput**: one divide per (N+1) MUL cycles **only if the multiplier is dedicated to the divider**. If the multiplier is shared with the main MUL datapath, throughput is degraded by contention with MUL issue rate and is workload-dependent. **INSUFFICIENT EVIDENCE** for a single throughput number in the shared case.
|
||||
- **Area**: 1 reciprocal seed ROM + 1–2 MUL units (sharing possible at the cost of contention); large.
|
||||
- **Power**: Higher static and dynamic (multiplier active during divide refinement).
|
||||
- **Used in**: Some high-performance FPU designs for floating-point; less common for dedicated integer divide.
|
||||
|
||||
### A4. Radix-16 / Radix-64 SRT Divider
|
||||
|
||||
- **Datapath**: high-radix recurrence with a quotient-digit lookup table (PLA or ROM), redundant remainder representation. 16 or 64 bits processed per cycle.
|
||||
- **DIV Latency**: ~16 cycles (radix-16) or ~8–10 cycles (radix-64) for 64-bit operands.
|
||||
- **DIV Throughput**: 1 per cycle if pipelined; 1 per 16 / 8–10 cycles if iterative.
|
||||
- **Area**: large lookup table and complex datapath; PLA is a significant area contributor.
|
||||
- **Power**: high; many bits toggle per cycle.
|
||||
- **Used in**: high-end x86 integer dividers (radix of the integer divider is **INSUFFICIENT EVIDENCE** from primary sources; widely reported as high-radix shift-subtract rather than Newton-Raphson).
|
||||
|
||||
### A5. Approximate / Lookup-Based Dividers (small operand ranges)
|
||||
|
||||
- Rarely used for 64-bit; not relevant to XH-1 unless a specific workload (e.g., 32-bit address arithmetic) dominates.
|
||||
|
||||
## Alternative Designs
|
||||
|
||||
### B1. Hybrid MUL + Sequential-Iterative DIV (Rocket-style)
|
||||
|
||||
- MUL: 64×64→128 pipelined in 1–2 stages (in Rocket's standard config, multi-stage; the 1-stage variant proposed for XH-1 is an aggressive target).
|
||||
- DIV: radix-4 iterative, ~33 cycles (32 cycles for quotient bits + finalization), blocking on the unit.
|
||||
- Single divider per core, 1–3 MUL pipeline stages.
|
||||
|
||||
### B1'. Two MUL Pipelines + Shared Iterative DIV (wide-issue variant)
|
||||
|
||||
- Two pipelined MUL datapaths, one shared radix-4 DIV datapath.
|
||||
- MUL throughput: 2/cycle (sustained, independent operands).
|
||||
- DIV throughput: 1 per ~33 cycles, shared and blocking.
|
||||
- Area: The "1.8–2.0× B1" multiplier in the prior revision is **INSUFFICIENT EVIDENCE** without a stated MUL:DIV area ratio. ASSUMING MUL and DIV are roughly comparable in area (a 1-stage MUL is a Wallace tree + 128-bit CPA; a radix-4 DIV is a 66-bit adder + control), the combined (MUL+DIV) area scales as 1 + (1/2) × 1 = 1.5× for two MULs + one shared DIV, but this is **INSUFFICIENT EVIDENCE** without a specific synthesis result. The correct qualitative statement is: roughly 1.4–2.0× B1 depending on assumed MUL:DIV area ratio.
|
||||
|
||||
### B2. Combined MUL/MAC (Multiply-Accumulate) Unit
|
||||
|
||||
- 64×64→128 with an additional accumulator register; supports future RV MAC extensions (proposed in some bitmanip drafts).
|
||||
- Adds area over pure MUL; **INSUFFICIENT EVIDENCE** for a specific percentage.
|
||||
- Useful for DSP, ML, and HPC; 128-core replication benefits workloads with matrix-style kernels.
|
||||
|
||||
### B3. Operand-Width-Detected 32-bit Fast MUL Path
|
||||
|
||||
- Detect when both operands of a 64-bit MUL are sign- or zero-extended from 32 bits (i.e., bit 31 is replicated through bit 63), and route the multiplication through a 32×32→64 fast multiplier.
|
||||
- This is **distinct from MULW** (which is a separate instruction with 32-bit result semantics). B3 accelerates 64-bit MUL / MULH / MULHSU / MULHU when both operands happen to be 32-bit sign- or zero-extended, a pattern common after `lw` / `lwu` followed by arithmetic. MULW does not help this case because MULW is a different instruction.
|
||||
- Open question: whether the verification cost of the dual datapath is justified at the 128-core replication level.
|
||||
|
||||
### B4. Bypassable Output with Operand Width Detection (full 64×64→128 only)
|
||||
|
||||
- Always implement full 64×64→128; the decoder examines the instruction to forward only the required 64 bits.
|
||||
- Verification: a single source of truth (full 128-bit datapath), reducing dual-mode bugs.
|
||||
- Disadvantage: no area saving.
|
||||
|
||||
### B5. Skip-on-Zero / Divide-Cancellation Optimizations
|
||||
|
||||
- Detect zero dividend (quotient is zero, remainder is dividend) and divide-by-one (quotient is dividend, remainder is zero) at the front end and forward the result without entering the iterative loop.
|
||||
- Saves latency in the common case for some workloads; trivial area overhead.
|
||||
- Composable with any divider organization.
|
||||
|
||||
### B6. Software Divide-by-Constant Transformation (cross-cutting)
|
||||
|
||||
- Compilers transform division by a **compile-time** constant into a multiply-by-reciprocal sequence. The hardware DIV is then needed only for division by variables.
|
||||
- Reduces effective DIV frequency significantly for workloads with constant denominators; affects hardware sizing decisions. **INSUFFICIENT EVIDENCE** for a quantitative reduction without a specific workload profile.
|
||||
|
||||
### B7. Dedicated 32-bit Fast Divider for W-suffixed Instructions
|
||||
|
||||
- Implement a separate 32-bit radix-2 or radix-4 iterative divider for DIVW / DIVUW / REMW / REMUW. The 32-bit divider has half the iteration count (32 or 16 cycles vs. 64 or 33 for the 64-bit divider) and roughly a quarter of the datapath area.
|
||||
- Useful if profiling shows W-suffixed divides dominate; in most general-purpose workloads they do not.
|
||||
- Verification cost: dual divider datapath, similar to B3.
|
||||
|
||||
### B8. Combined MUL / DIV with Shared Partial-Product Array
|
||||
|
||||
- Reuse the MUL's CSA compressor tree as the final correction multiplier for an SRT or Newton-Raphson divider, sharing the most area-intensive block.
|
||||
- Reduces the area penalty of A3 / A4 at the cost of tighter verification coupling between MUL and DIV paths.
|
||||
- Real architectural option in some high-performance designs; not considered in the prior revision.
|
||||
|
||||
## Comparison
|
||||
|
||||
| Design | MUL Latency | MUL Tput | DIV Latency | DIV Tput | Rel. Area (MUL+DIV) | Power (active) | Verification Complexity |
|
||||
|--------|-------------|----------|-------------|----------|----------------------|----------------|--------------------------|
|
||||
| A1 (radix-2 iter. DIV only) | not defined in A1 | not defined in A1 | 64 cycles | 1 per 64 cycles | INSUFFICIENT EVIDENCE (combined) | Lowest (DIV only) | Low |
|
||||
| A2a (pipelined iter. loop) | not defined in A2a | not defined in A2a | 64 cycles | 1 per cycle (after fill) | INSUFFICIENT EVIDENCE | Med | Med |
|
||||
| A2b (fully unrolled) | not defined in A2b | not defined in A2b | 64 cycles | 1 per cycle | INSUFFICIENT EVIDENCE (very large) | High | High |
|
||||
| A3 (Newton-Raphson) | 1 cycle | 1 per cycle (dedicated) | N+1 MUL cycles (dedicated); workload-dep. if shared | INSUFFICIENT EVIDENCE (shared) | high | High | High |
|
||||
| A4 (Radix-16/64 SRT) | not defined in A4 | not defined in A4 | 8–16 cycles | 1 per cycle if pipelined | high | High | High |
|
||||
| B1 (Rocket-style) | 1–2 cycle (1-stage target) | 1 per cycle | ~33 cycles | 1 per 33 cycles | baseline | Low–Med | Low–Med |
|
||||
| B1' (2× MUL + 1× shared DIV) | 1–2 cycle | 2 per cycle | ~33 cycles | 1 per 33 cycles (shared) | ~1.4–2.0× B1 (INSUFFICIENT EVIDENCE; depends on MUL:DIV area ratio) | Med | Med |
|
||||
| B2 (+ MAC) | 1–2 cycle | 1 per cycle | ~33 cycles | 1 per 33 cycles | +unspecified % over B1 (INSUFFICIENT EVIDENCE) | Med | Med |
|
||||
| B3 (32-bit fast MUL) | <1 cycle (32-bit path) | 1 per cycle | ~33 cycles | 1 per 33 cycles | +small (INSUFFICIENT EVIDENCE) | Low | Med–High (dual mode) |
|
||||
| B4 (full 64 only) | 1–2 cycle | 1 per cycle | ~33 cycles | 1 per 33 cycles | same as B1 | Med | Lowest |
|
||||
| B5 (skip-on-zero) | n/a | n/a | reduced in common case | same as base | negligible overhead | n/a | Low |
|
||||
| B7 (32-bit fast DIV) | 1–2 cycle | 1 per cycle | ~16–17 cycles (32-bit) | 1 per 16–17 cycles (32-bit only) | +small (INSUFFICIENT EVIDENCE) | Low | Med–High (dual mode) |
|
||||
|
||||
**ASSUMPTION**: Relative area figures are qualitative orderings based on published reference designs cited in the Sources section. Actual XH-1 synthesis numbers are **INSUFFICIENT EVIDENCE** pending RTL implementation.
|
||||
|
||||
## Advantages
|
||||
|
||||
### A1 / B1
|
||||
- Smallest area; lowest per-core replication cost.
|
||||
- Lowest static and dynamic power in the typical case (most cycles no MUL/DIV active).
|
||||
- Simpler verification; one mode of operation.
|
||||
- Well-understood reference implementation (Rocket, Ibex) for cross-checking.
|
||||
|
||||
### A4 (High-Radix SRT)
|
||||
- Lowest DIV latency among the iterative-style options; competitive with A3 on a single divide.
|
||||
- Pipelined variant gives 1 per cycle DIV throughput.
|
||||
|
||||
### A3 (Newton-Raphson)
|
||||
- Low latency for sequential divides — important for cryptography (RSA, ECC modular reduction) and HPC.
|
||||
- Amortizes multiplier cost if MAC (B2) is also desired.
|
||||
- Disadvantage: requires multiple refinement iterations; integer multiplier is large and contention with the main MUL datapath is a concern.
|
||||
|
||||
### B2 (MAC)
|
||||
- Enables future-proofing for proposed bitmanip and MAC extensions.
|
||||
- Helpful for matrix multiplication kernels running across 128 cores.
|
||||
|
||||
### B5 (Skip-on-Zero)
|
||||
- Negligible area; reduces effective DIV latency for common cases.
|
||||
|
||||
### B7 (32-bit fast DIV)
|
||||
- Halves the divider iteration count for the W-suffixed instructions at modest area cost.
|
||||
|
||||
### B8 (Shared MUL/DIV CSA)
|
||||
- Reduces the area penalty of high-performance dividers by sharing the most area-intensive block.
|
||||
|
||||
## Disadvantages
|
||||
|
||||
### A1 / B1
|
||||
- Variable DIV latency complicates scheduling in an in-order core; in-order cores stall ~33 cycles per divide.
|
||||
- Worst-case latency hits interrupt latency, branch-target computation if the MUL/DIV is on a branch path.
|
||||
- Not a fit for sustained divide-bound workloads (e.g., modulo arithmetic in bignum).
|
||||
|
||||
### A3
|
||||
- Area at 128-core replication is severe; the multiplier is one of the largest blocks in a typical core.
|
||||
- Power: a 64×64 multiplier running 1 per cycle is one of the highest-power blocks in a typical core.
|
||||
- Verification: Newton iteration precision (must converge to within 1 ULP across all operands) is notoriously subtle.
|
||||
- If multiplier is shared with main MUL datapath, throughput is workload-dependent, not the N+1 figure cited for the dedicated case.
|
||||
|
||||
### A4
|
||||
- Large lookup table (PLA or ROM); area and power dominated by the table.
|
||||
- Verification: complex quotient-digit selection logic.
|
||||
- Not commonly used outside high-end commercial designs.
|
||||
|
||||
### B3
|
||||
- Verification: dual datapath (32-bit and 64-bit) must be exhaustively cross-checked.
|
||||
- The classification logic (operand-width detection) is itself a source of corner-case bugs.
|
||||
- **B3 does not subsume MULW**; MULW is a separate instruction and B3 is about 64-bit MUL on 32-bit-valued operands.
|
||||
|
||||
### B2
|
||||
- Adds a third writeback port or accumulator path; competes with the FPU/LSU for writeback slots in a multi-issue core.
|
||||
|
||||
### B5
|
||||
- Only helps the specific cases of zero dividend or divisor ±1; other optimizations (e.g., division by small powers of two) are already handled by the base I extension's shift instructions.
|
||||
|
||||
### B7
|
||||
- Verification: dual divider datapath; same concerns as B3.
|
||||
- Area saving is moot if W-suffixed divides are not on the critical path.
|
||||
|
||||
### B8
|
||||
- Verification: tighter coupling between MUL and DIV paths makes corner-case analysis more difficult.
|
||||
|
||||
## XH-1 Considerations
|
||||
|
||||
**PROPOSAL**: The XH-1 core should adopt **B1 (Hybrid MUL + Iterative DIV)** as the baseline, with the following parameters, pending RTL validation:
|
||||
|
||||
- **MUL**: 64×64→128, target 1-cycle pipelined (1 stage of pipeline registers), throughput 1 per cycle. For 2-issue or wider cores, scale to B1' (two MUL pipelines sharing one DIV).
|
||||
- **DIV**: Radix-4 shift-subtract, ~33 cycles worst case (32 cycles for quotient bits + finalization), blocking, non-pipelined. The prior revision's "33–35 cycles worst case" and the "33–64 cycle" range are reconciled here: 33 is the radix-4 bound; 64 corresponds to radix-2, which is a different algorithm choice.
|
||||
- **REM**: Reuse the DIV datapath; remainder is a by-product.
|
||||
- **Bypass paths**: From MUL output directly to subsequent dependent ALU/branch in the next cycle. **INSUFFICIENT EVIDENCE** on whether this fits the XH-1 pipeline depth; depends on the integration context.
|
||||
- **B5 (skip-on-zero)**: implement at the front end of the DIV unit; negligible overhead.
|
||||
- **Operand-width detection**: Defer B3 to a future revision; not justified at the 128-core replication level given the added verification cost and the existence of MULW for the 32-bit-case instruction.
|
||||
- **B7 (32-bit fast DIV)**: Defer; the W-suffixed divide is not assumed to be on the critical path. Revisit if profiling shows otherwise.
|
||||
|
||||
**ASSUMPTION**: A 1-cycle MUL latency is achievable in the target process. **INSUFFICIENT EVIDENCE** on the XH-1 target process node and clock period. A full 64×64→128 Wallace/Dadda tree + 128-bit carry-propagate adder in a single cycle is at the edge of feasibility for high-performance designs; typical in-order cores implement MUL as either a multi-cycle iterative multiplier or a multi-stage pipelined multiplier. **PROPOSAL**: Validate via synthesis at the target corner before committing. If the 1-cycle critical path cannot be closed, **fallback options** are:
|
||||
- **B1-fallback-A**: 2-cycle pipelined MUL (split the compressor tree and the final CPA across two pipeline stages); latency 2 cycles, throughput 1 per cycle, modest area overhead.
|
||||
- **B1-fallback-B**: Multi-cycle iterative MUL (Booth-encoded, ~4–8 cycles for 64×64→128); lower area, lower throughput, higher latency.
|
||||
- **B1-fallback-C**: Retain the 1-stage MUL architecture but lower the target clock frequency (system-level decision, not unit-level).
|
||||
|
||||
**PROPOSAL**: The MUL/DIV unit sits on the integer execution lane, sharing the issue queue with ALU and branch. It has its own reservation station (or issue-slot tag) so the issue logic can distinguish short-latency ALU (1 cycle) from long-latency MUL (1, or 2 in the fallback) and variable-latency DIV (~33).
|
||||
|
||||
**PROPOSAL**: For the in-order case, the MUL/DIV unit is non-blocking on MUL (1-cycle latency) but blocking on DIV; the in-order front-end stalls on a structural hazard only when a second DIV is issued before the first completes.
|
||||
|
||||
**PROPOSAL**: For an out-of-order XH-1 core, the MUL/DIV unit exposes its variable DIV latency through the wakeup/select logic so that dependent instructions are replayed or re-issued with the correct ready signal.
|
||||
|
||||
**PROPOSAL**: The B-extension unit (if RV64B is implemented) shares **operand muxes, sign-handling logic, and bit-level muxes** with the MUL/DIV unit but does **not** share the Wallace/Dadda compressor tree. Operations like CLZ, CTZ, BSET, BEXT operate on individual bits or small bit-fields and do not naturally map onto a Wallace-tree multiplier datapath. CLMUL / CLMULH / CLMULR (Zbc) require an AND-tree / XOR-reduction datapath that is structurally distinct from both the Wallace-tree multiplier and the iterative divider; they do not share the compressor tree. Any apparent sharing is at the operand-fetch and writeback layers, not the core arithmetic.
|
||||
|
||||
## 128-Core Scalability
|
||||
|
||||
The 128-core replication factor is the dominant cost driver for the MUL/DIV unit.
|
||||
|
||||
**PROPOSAL**: At 128 cores, every MUL unit must be **power-gateable independently** so that cores not in use (DPM, dark-silicon) can be fully clock- and power-gated, including the MUL/DIV block.
|
||||
|
||||
**PROPOSAL**: The MUL/DIV unit's reservation-station entries, divider iteration state, and pipeline registers must support **state retention or clean-state entry** across power gating. Specifically, when a core is power-gated while a long-latency DIV is in flight, the DIV's mid-iteration state must be either (a) flushed (architecturally equivalent to a DIV that completes with the wrong value, which is not acceptable), or (b) checkpointed to a retention register or to memory, or (c) prevented from power-gating until the DIV completes. The choice depends on the XH-1 DPM policy. **INSUFFICIENT EVIDENCE** on the XH-1 DPM policy; the design must accommodate one of these options without committing to a specific approach here.
|
||||
|
||||
**OPEN QUESTION**: What is the actual MUL/DIV area share of the XH-1 core? Published RISC-V references suggest a typical MUL unit occupies a small single-digit percentage of a high-performance core's area, with iterative DIV adding additional area. The prior revision's "1–3%" and "1 MGE/core total" baseline are removed here as unsourced. **INSUFFICIENT EVIDENCE** on the XH-1-specific area share without synthesis.
|
||||
|
||||
**OPEN QUESTION**: Is the MUL/DIV area share dominated by the MUL datapath, the DIV datapath, the operand muxing, or the bypass network? Targeted synthesis required.
|
||||
|
||||
**INSUFFICIENT EVIDENCE**: The XH-1 die-area budget, process node, and clock period are not established in this document.
|
||||
|
||||
**PROPOSAL**: The MUL/DIV unit should be **identical across all 128 cores** (no per-core specialization) to simplify verification and physical design. The replication should be a Verilog/SystemVerilog generate block driven by a single template. **PROPOSAL**: The generate block itself must be included in the verification scope (not assumed trivially correct). At minimum: a lint-clean check, a synthesis-check that the generate block instantiates the correct number of cores, and a per-instance equivalence check on a sample of cores.
|
||||
|
||||
**PROPOSAL**: For 128× replicated MUL/DIV pipeline registers, ECC or parity protection should be considered for soft-error mitigation. **INSUFFICIENT EVIDENCE** on the XH-1 reliability target.
|
||||
|
||||
**PROPOSAL**: Reset distribution and scan chain architecture for 128× replicated MUL/DIV must be addressed at the integration level. The MUL/DIV unit's scan chains should support parallel or staggered scan-shift across cores to keep test time bounded. **INSUFFICIENT EVIDENCE** on the XH-1 DFT architecture.
|
||||
|
||||
**CONSIDERATION (clock and timing at 128× replication)**: If the XH-1 die includes a global clock-distribution network, the MUL/DIV critical path determines the local clock skew tolerance. A pipelined 1-cycle MUL has a short critical path (one 64-bit adder-equivalent) which is favorable; a non-pipelined iterative divider has a 64-bit adder in its loop, which is the cycle-time limiter. The clock-skew analysis should be revisited at the integration level once the XH-1 clock tree is defined.
|
||||
|
||||
**CONSIDERATION (cross-core aggregate throughput)**: Under B1, each core's blocking DIV delivers ~1 per 33 cycles per core. Across 128 cores, the aggregate is ~3.9 divides per cycle in the best case (all cores dividing simultaneously), which is the upper bound. Real workloads do not exhibit this worst case; the relevant metric is the per-core latency, not aggregate. **INSUFFICIENT EVIDENCE** on whether the XH-1 DPM or interconnect imposes a global cap on simultaneous divide activity; this is a system-level question outside the MUL/DIV unit's scope.
|
||||
|
||||
**CONSIDERATION (operand distribution and interconnect)**: Replicating a Wallace tree 128× implies 128 sets of wide operand buses to/from the register file. The interconnect / operand-routing network cost scales with the MUL operand width and the number of cores. **PROPOSAL**: include operand-routing overhead in the area estimate, not just the MUL/DIV datapath itself.
|
||||
|
||||
## Performance Considerations
|
||||
|
||||
**PROPOSAL**: **MUL throughput of 1 per cycle is a target for a high-performance XH-1 core**, but is not architecturally non-negotiable. Low-end in-order cores (some Ibex configurations) implement MUL with throughput < 1 per cycle. B1 satisfies the high-performance target with a 1-stage pipelined multiplier; B1' extends to 2 per cycle for wide-issue. If the 1-cycle MUL cannot be closed at the target process, the throughput target remains 1 per cycle but the latency becomes 2 cycles (B1-fallback-A).
|
||||
|
||||
**PROPOSAL**: **DIV throughput of 1 per ~33 cycles (blocking) is acceptable** for most integer workloads. Workloads that require more (e.g., modular arithmetic in cryptography) will suffer; mitigation is left to software (see Software Considerations).
|
||||
|
||||
**ASSUMPTION**: XH-1 target workloads include a mix consistent with Embench / SPECint-class profiles, where MUL/DIV instructions are a small fraction of dynamic instruction count. **INSUFFICIENT EVIDENCE** on the actual XH-1 target workload mix and on the specific dynamic-instruction share of MUL/DIV. The prior revision's "<2% MUL/DIV" and "<0.5% DIV/REM" claims are removed as unsourced.
|
||||
|
||||
**OPEN QUESTION**: Does XH-1 target HPC or cryptography workloads where MUL/DIV is a larger share of the instruction mix? If yes, the analysis shifts toward A3, A4, or A2. **NOTE**: ML workloads are dominated by floating-point multiplies on the FPU, not by integer MUL/DIV; integer MUL/DIV is relevant to ML only for quantization, address arithmetic, and integer embeddings.
|
||||
|
||||
## Area Considerations
|
||||
|
||||
**PROPOSAL**: Budget the MUL/DIV unit at a small single-digit percentage of single-core area for the B1 design, pending synthesis. The exact percentage is **INSUFFICIENT EVIDENCE**.
|
||||
|
||||
**OPEN QUESTION**: Is the XH-1 core area budget (without MUL/DIV) known? The MUL/DIV area share can be re-expressed as a percentage of the entire 128-core die only when this is established.
|
||||
|
||||
**ASSUMPTION**: A1-style radix-2 DIV is the smallest practical divider; A3 (Newton) and A4 (high-radix SRT) are larger, with A4 typically the largest. **INSUFFICIENT EVIDENCE** on the XH-1-specific gate-count budget.
|
||||
|
||||
## Power and Energy Considerations
|
||||
|
||||
**FACT**: A 64×64→128 multiplier toggles a large number of bits per cycle when active; its dynamic power is proportional to operand activity factor. Power-gating the multiplier when idle is essential.
|
||||
|
||||
**PROPOSAL**: The MUL/DIV unit should support **fine-grained clock gating** at the operand-mux boundary so that when no MUL/DIV is in flight, the entire unit is clock-gated to 0 toggle rate.
|
||||
|
||||
**PROPOSAL**: The MUL pipeline register should be a clock-gated scan flop, not a latch, to simplify DFT.
|
||||
|
||||
**OPEN QUESTION**: Does the XH-1 power budget allow all 128 MUL/DIV units to be active simultaneously? If not, per-core DPM must enforce that no more than N MUL/DIV units are in active divide at once. This is a system-level power-management question, not a unit-level one, but it constrains the unit's design.
|
||||
|
||||
**INSUFFICIENT EVIDENCE**: Specific per-MUL or per-DIV energy numbers for the XH-1 process are not established. The prior revision's "single-digit pJ in 7 nm" claim is removed as unsourced; per-MUL energy in advanced processes is implementation-dependent and varies by an order of magnitude or more depending on architecture and clock frequency.
|
||||
|
||||
## Implementation Considerations
|
||||
|
||||
**PROPOSAL**: Implement MUL as a **Wallace/Dadda tree of 4:2 compressors** feeding a final 128-bit carry-propagate adder. The 1-cycle latency budget must accommodate the entire critical path from operand register → compressor tree → CPA → output register. **INSUFFICIENT EVIDENCE** on whether a Wallace/Dadda tree for 64×64 partial products is achievable in one cycle at the XH-1 target clock period; the specific tree depth and CPA depth are implementation-dependent and not asserted as fixed numbers here. The prior revision's "Tree depth 6–7; final CPA ~6 gates deep" is removed as unsourced and implementation-specific.
|
||||
|
||||
**PROPOSAL**: For MULHSU, the front-end sign-handling stage should sign-extend the signed operand to 128 bits and zero-extend the unsigned operand, then feed a signed multiplier; alternatively, use a modified-Booth-encoded signed multiplier with the unsigned operand zero-extended. Either approach requires explicit sign handling before the partial-product reduction tree.
|
||||
|
||||
**PROPOSAL**: Implement DIV as a **radix-4 shift-subtract** with restoring on negative remainder. 32 cycles for quotient bits + 1 cycle for finalization = ~33 cycles total. The remainder is the "running remainder" register after the 33rd cycle; the quotient is the shifted-out bits. REM is selected by muxing the appropriate output.
|
||||
|
||||
**PROPOSAL**: B5 (skip-on-zero): add a front-end detector on the DIV operands that forwards the result directly for divisor = ±1 or dividend = 0, bypassing the iterative loop. Trivial area; reduces effective DIV latency for common cases.
|
||||
|
||||
**PROPOSAL**: Parameterize the MUL width with a Verilog `parameter WIDTH=64` so that the same RTL can be synthesized for 32-bit-only cores (test variant, debug, or soft-IP reuse) without manual rework. Note that WIDTH=32 does not by itself support the MULW instruction semantics in RV64 (which performs a 32×32→64 multiply and sign-extends the 32-bit result); an MULW-specific path is required for RV64.
|
||||
|
||||
## Verification Considerations
|
||||
|
||||
**FACT**: MUL and DIV are among the most heavily tested instructions in any RISC-V core; their verification is dominated by **corner-case operand pairs** (e.g., 0, 1, -1, 2, -2, 0x7FFF…, 0x8000…, 0xFFFF…, MSB=1 transitions). For DIV/REM, the corner cases include the division-by-zero and signed-overflow rules defined in the ISA spec.
|
||||
|
||||
**PROPOSAL**: Adopt a **directed-random + constrained-random** verification flow using a golden reference model:
|
||||
- **Reference MUL**: SystemVerilog `bit [127:0]` (or DPI-C to a software bigint). Compare lower-64 and upper-64 separately. MULW, MULH, MULHU, MULHSU all share the 128-bit reference.
|
||||
- **Reference DIV**: Software bigint that produces (quotient, remainder) per the RISC-V M-extension spec, including the division-by-zero and signed-overflow rules for both quotient-producing and remainder-producing instructions.
|
||||
- **Coverage**: 100% of the corner cases (above), plus cross-product of small operand values (±1, ±2, ±2^32, ±2^63, ±2^64-1), plus randomized large operands.
|
||||
- **Regression list size**: The prior revision's "64 hand-crafted corner cases" is removed as a specific number; the regression list should be sized to cover the documented corner cases and is grown as bugs are found. **INSUFFICIENT EVIDENCE** for a canonical count.
|
||||
|
||||
**PROPOSAL**: At the 128-core replication level, **per-core functional verification** of the MUL/DIV unit is unnecessary if the per-core RTL is identical and constrained-random coverage is closed at the template level (via a SystemVerilog generate block). This covers functional equivalence at the unit level. **However**, the verification of the generate block itself, and the interaction between per-core clock-gating / power-state and MUL/DIV state (e.g., does a clock-gated MUL lose its pipeline state correctly across gating? does a power-gated divider leave the iteration counter in a valid state for resumption?), must be verified explicitly. **Per-core physical / timing verification is not bypassed**: timing, DFT, and physical-design closure are verified at the integration level on a representative core and assumed replicated, with explicit per-die variation analysis as required by the XH-1 physical-design flow.
|
||||
|
||||
**PROPOSAL**: If the XH-1 verification flow includes formal property checking, the DIV correctness property ("quotient × divisor + remainder == dividend" within the RISC-V M extension's special cases) is a clean formal target and should be specified. **INSUFFICIENT EVIDENCE** on whether formal property checking is in scope.
|
||||
|
||||
## Software Considerations
|
||||
|
||||
**PROPOSAL**: Document the MUL/DIV latencies (1 cycle MUL, ~33 cycle DIV) in the XH-1 ABI / ISA reference manual so that compiler backends and library authors can schedule around the long-latency DIV.
|
||||
|
||||
**OPEN QUESTION**: Does the XH-1 ABI / linker convention include a software-emulated division routine for code that cannot tolerate the ~33-cycle blocking latency? If yes, the hardware DIV can be a single issue-slot, not a deeply pipelined one. The compiler can also apply divide-by-constant transformations (B6) to reduce effective hardware DIV frequency.
|
||||
|
||||
**PROPOSAL**: The MUL/DIV unit should raise no architectural exception; the RISC-V M extension defines all corner cases architecturally. This simplifies the trap logic.
|
||||
|
||||
**PROPOSAL**: For 128-core use cases, the MUL/DIV unit is the same in every core; the OS scheduler does not need to be aware of its microarchitectural details. Workload distribution across cores is independent of MUL/DIV design.
|
||||
|
||||
## Recommendation
|
||||
|
||||
**RECOMMENDATION**: Adopt **B1** — a hybrid MUL unit (target 1-cycle pipelined, full 64×64→128) combined with a radix-4 iterative, non-pipelined divider (~33 cycles), sharing no logic. Add **B5** (skip-on-zero) at the DIV front end. Place the unit on the integer execution lane. Support full per-unit power and clock gating for 128-core DPM. For wide-issue cores, scale to **B1'** (two MUL pipelines sharing one DIV).
|
||||
|
||||
Rationale:
|
||||
1. Matches the dominant MUL workload profile (frequent, 1-cycle latency needed for high-performance targets).
|
||||
2. Keeps per-core area small, manageable at 128× replication.
|
||||
3. Avoids the verification burden of B3 (dual datapath) and the area burden of A3 (Newton-Raphson) and A4 (high-radix SRT).
|
||||
4. Aligns with the architectural pattern of proven reference designs (Rocket, Ibex); the alignment is on the radix-4 iterative divider + pipelined multiplier pattern, not on a specific pipeline depth, since Rocket's MUL is multi-stage in standard config and Ibex's MUL is short-pipeline / combinational.
|
||||
5. B5 is a near-free improvement to the common case.
|
||||
|
||||
**Fallback plan** (if the 1-cycle MUL cannot be closed at the target process / clock):
|
||||
- **Primary fallback**: B1-fallback-A — 2-cycle pipelined MUL. Latency 2 cycles, throughput 1 per cycle, modest area overhead. B1 architecture preserved.
|
||||
- **Secondary fallback**: B1-fallback-B — multi-cycle iterative MUL (Booth-encoded, ~4–8 cycles). Lower area, lower throughput, higher latency. DIV remains radix-4 iterative.
|
||||
- **Tertiary fallback**: Lower the target clock frequency at the system level.
|
||||
|
||||
This recommendation is **conditional on**:
|
||||
- The XH-1 target process supporting either a 1-cycle 64×64→128 MUL critical path (primary) or a 2-cycle pipelined MUL critical path (fallback A) at the target clock. Both must be validated by synthesis at the target corner; **INSUFFICIENT EVIDENCE** without target process specification.
|
||||
- XH-1 workload mix not being dominated by sequential DIV/REM (cryptography). If it is, escalate to A4 (radix-16 SRT) or a pipelined iterative divider (A2a), and/or consider B8 (shared MUL/DIV CSA).
|
||||
- The 128-core replication budget tolerating the cumulative MUL/DIV area; this requires a known single-core area budget, which is **INSUFFICIENT EVIDENCE**.
|
||||
- A clear DPM policy for handling in-flight MUL/DIV state across power gating.
|
||||
|
||||
If any of these conditions fails, re-open the design against the named fallback.
|
||||
|
||||
## Confidence
|
||||
|
||||
**Medium-High** for B1 as the baseline architectural pattern. **Low** for specific area, power, and energy numbers — those are estimates based on published RISC-V reference designs, not on XH-1 synthesis data. **Low** for any claim about the XH-1 target process, workload mix, or die size (these are **INSUFFICIENT EVIDENCE**). The corrected version removes specific unsourced quantitative claims and demotes several prior FACTs to ASSUMPTION or INSUFFICIENT EVIDENCE.
|
||||
|
||||
## Open Questions
|
||||
|
||||
1. What is the XH-1 target process node and clock period? This constrains MUL critical-path design and determines whether 1-cycle MUL is feasible.
|
||||
2. What is the XH-1 target workload mix? HPC, cryptography, or general-purpose? (ML is not primarily an integer MUL/DIV workload.)
|
||||
3. Is the XH-1 core in-order, out-of-order, or hybrid?
|
||||
4. Does XH-1 implement RV64B (bit-manipulation) extensions, including Zbc (CLMUL / CLMULH / CLMULR)? The B extension does not naturally share the Wallace-tree multiplier datapath; any sharing is at the operand-mux and writeback layers, not the compressor tree.
|
||||
5. What is the XH-1 single-core area budget? What is the die-level area budget for the 128-core fabric?
|
||||
6. Does XH-1 use a per-core DPM scheme, and how does it handle in-flight MUL/DIV state across power gating (retention, flush-and-replay, or block-power-gate-until-completion)?
|
||||
7. Is the MUL/DIV unit expected to support a future RV MAC extension, or is the project strictly RV64IM (or RV64IMB)?
|
||||
8. Is there a software-emulated division in the XH-1 libc / ABI that would tolerate a long-latency hardware DIV? Will the compiler apply divide-by-constant transformations (B6)?
|
||||
9. Does the XH-1 verification flow include formal property verification, and will it be applied to the MUL/DIV unit's "quotient × divisor + remainder == dividend" property?
|
||||
10. For wide-issue cores, is B1' (two MUL pipelines + one shared DIV) the target, or is single-MUL B1 sufficient?
|
||||
11. What is the XH-1 interconnect / operand-routing cost of replicating a wide MUL operand bus 128 times? Should this overhead be included in the MUL/DIV unit's area budget?
|
||||
12. What is the XH-1 reliability target? Does the MUL/DIV pipeline require ECC or parity protection against soft errors?
|
||||
13. What is the XH-1 DFT architecture for the 128× replicated MUL/DIV scan chains?
|
||||
|
||||
## Sources
|
||||
|
||||
- Waterman, Asanović, et al. (eds.), *The RISC-V Instruction Set Manual, Volume I: User-Level ISA, Document Version 20191213*, Chapter 7 (M Extension). RISC-V International. Canonical ISA reference; defines the M-extension corner-case rules for DIV/REM/DIVU/REMU/DIVW/REMW/DIVUW/REMUW and MULW.
|
||||
- Waterman, Asanović, et al. (eds.), *The RISC-V Instruction Set Manual, Volume I: User-Level ISA*, Chapter 16 (B Extension, including Zbc). Canonical ISA reference for CLMUL / CLMULH / CLMULR.
|
||||
- Asanović et al., "The Rocket Chip Generator," EECS Department, UC Berkeley, Technical Report UCB/EECS-2016-17, 2016. Reference for Rocket's MUL/DIV organization; configuration-dependent.
|
||||
- Celio, Patterson, Asanović, "BOOM v2: a superscalar out-of-order processor," 2017. Reference for BOOM's MUL/DIV design.
|
||||
- Ibex documentation, lowRISC. Reference for the in-order baseline.
|
||||
- Chen et al., "XiangShan: An Open-Source High-Performance RISC-V Core," 2022. Reference for high-performance RISC-V MUL/DIV design.
|
||||
- Parhami, *Computer Arithmetic: Algorithms and Hardware Designs*, Oxford University Press. Reference for shift-subtract, SRT, and Newton-Raphson divide algorithms.
|
||||
- Flynn, Oberman, *Advanced Computer Arithmetic Design*. Reference for high-radix divider design.
|
||||
- Weste, Harris, *CMOS VLSI Design: A Circuits and Systems Perspective*, 4th ed. Reference for Wallace/Dadda multiplier design and critical-path analysis; does not by itself establish a specific 1-cycle latency in any particular process.
|
||||
|
||||
**INSUFFICIENT EVIDENCE**: No XH-1-specific RTL, synthesis, or benchmark measurements are available at the time of this document. All quantitative claims about XH-1 area, power, energy, and timing are **INSUFFICIENT EVIDENCE** pending RTL implementation and target-process specification. Cortex-A77 per-instruction latencies are not publicly published by Arm and are **INSUFFICIENT EVIDENCE** from primary sources. The internal radix of the Intel Haswell integer divider is widely reported as high-radix shift-subtract rather than Newton-Raphson, but the specific radix (16 vs. 32) is **INSUFFICIENT EVIDENCE** from primary sources; Newton-Raphson is reported to be used in Haswell's floating-point unit.
|
||||
+341
File diff suppressed because one or more lines are too long
+505
@@ -0,0 +1,505 @@
|
||||
# MUL/DIV Unit
|
||||
|
||||
## Status
|
||||
|
||||
Revision 2. Initial scoping document. No XH-1 implementation decisions are yet committed. This revision corrects factual errors identified in review, removes unsupported quantitative claims, reconciles internal contradictions, and adds missing alternatives.
|
||||
|
||||
## Abstract
|
||||
|
||||
This document investigates the design of the integer multiply/divide (MUL/DIV) execution unit for the XH-1 core. The unit must implement the RV64M extension (MUL, MULH, MULHU, MULHSU, DIV, DIVU, REM, REMU, MULW, DIVW, DIVUW, REMW, REMUW) and, depending on project scope, the RV64B bit-manipulation extensions (e.g., CLZ, CTZ, BSET, BEXT, CLMUL, CLMULH, CLMULR). The design must be evaluated against the XH-1's 128-core replication factor: a unit that is only a small percentage of a single core's area consumes a large cumulative area across the die when replicated 128 times. The unit's latency impacts critical-path timing, and its throughput affects overall core IPC under integer-heavy workloads (cryptography, hashing, address arithmetic, HPC kernels).
|
||||
|
||||
## Research Question
|
||||
|
||||
What is the optimal MUL/DIV unit organization for an XH-1 core given that the design will be replicated 128 times, considering latency, throughput, area, power, energy, and verification complexity trade-offs across in-order, out-of-order, and hybrid pipeline styles?
|
||||
|
||||
## Background
|
||||
|
||||
### RISC-V M Extension Semantics (RV64M)
|
||||
|
||||
**FACT** (per RISC-V ISA spec, Volume I, Chapter 7): The RISC-V M extension for RV64 defines the following instructions:
|
||||
|
||||
- MUL, MULH, MULHU, MULHSU: 64×64→128 bit multiply (lower 64 or upper 64 of result, with sign-handling variations).
|
||||
- MULW: 32×32→32 bit product, then sign-extended to 64 bits and written to `rd`.
|
||||
- DIV, DIVU: 64÷64 signed/unsigned quotient.
|
||||
- REM, REMU: 64÷64 signed/unsigned remainder.
|
||||
- DIVW, DIVUW, REMW, REMUW: 32÷32 signed/unsigned quotient/remainder, sign-extended to 64 bits and written to `rd`.
|
||||
|
||||
**FACT** (per RISC-V ISA spec, Volume I, Chapter 7): The M extension defines the following corner-case behavior:
|
||||
|
||||
- Division by zero:
|
||||
- `DIV` / `DIVW`: quotient is `−1` (the architectural definition; the bit pattern is `2^XLEN − 1`, all bits set).
|
||||
- `DIVU` / `DIVUW`: quotient is `2^XLEN − 1` (all bits set, which equals `−1` in two's complement representation).
|
||||
- The signed and unsigned cases produce the same bit pattern at the architectural level; the spec writes `−1` for the signed case and the unsigned maximum for the unsigned case, but these are bit-pattern-identical.
|
||||
- `REM` / `REMW`: remainder equals the dividend.
|
||||
- `REMU` / `REMUW`: remainder equals the dividend.
|
||||
- Signed overflow (most-negative integer divided by −1):
|
||||
- `DIV` / `DIVW`: quotient equals the dividend (i.e., the most-negative representable value `2^(XLEN−1)`).
|
||||
- `REM` / `REMW`: remainder equals zero.
|
||||
- For unsigned divide (`DIVU` / `REMU` / `DIVUW` / `REMUW`), the only defined special case is division by zero; the dividend / `−1` overflow case does not apply because the operands are unsigned.
|
||||
|
||||
**NOTE**: A full 64×64→128-bit multiply requires a true 128-bit result datapath even when only the upper or lower 64 bits are returned.
|
||||
|
||||
**NOTE** (MULHSU implementation): MULHSU computes the upper 64 bits of a signed×unsigned 64×64 product. A correct implementation generates a partial-product array for 64×64 bits where the unsigned operand's partial products are zero in the upper half of its bit positions and the signed operand's partial products use signed (sign-extended) rows in the final reduction. Two concrete implementation paths exist:
|
||||
|
||||
- (a) Use a signed multiplier datapath: sign-extend the signed operand to 128 bits, zero-extend the unsigned operand to 128 bits, and run a signed 128×128 multiply, then take the upper 64 bits. This is straightforward but doubles the multiplier width and is rarely used.
|
||||
- (b) The standard approach: zero-extend the unsigned operand to 64 bits (its bit positions are already non-negative), sign-extend the signed operand only in the final partial-product row, and reduce the 64×64 partial-product array with a final row sign-extension. This requires a modified-Booth or array multiplier with explicit sign handling on the last partial-product row.
|
||||
|
||||
A pure unsigned Wallace/Dadda tree without a sign-handling front-end does not implement MULHSU correctly, because the signed operand's most significant partial-product row must be sign-extended (or its inverted-and-carry form added) into the reduction tree.
|
||||
|
||||
**NOTE**: The W-suffixed instructions (MULW, DIVW, DIVUW, REMW, REMUW) are part of the **M extension** in RV64, not the base I extension. The base I extension's W variants are only the simple ALU ops (ADDW, SUBW, SLLW, SRLW, SRAW).
|
||||
|
||||
**NOTE**: MULW is implementable on a 32×32→64-bit datapath: the 32-bit product is sign-extended to 64 bits and written to `rd`. The upper 32 bits of the 64-bit intermediate result are discarded. A 32×32→64-bit fast multiplier (B3 datapath) can therefore implement MULW by taking its lower 32 bits and sign-extending to 64.
|
||||
|
||||
**NOTE**: The carry-less multiply instructions CLMUL, CLMULH, and CLMULR are part of the standard Zbc extension (commonly grouped under the umbrella "B" extension in some profiling). They are not part of M. They require a different datapath (AND-tree with XOR reduction, no carry propagation) and are not the subject of this document except where they interact with operand muxes / writeback. CLMULR in particular produces a 2·XLEN-bit result with explicit carry-in/carry-out behavior across the two halves; its implementation shares the AND-tree with CLMUL/CLMULH but adds a dedicated reduction stage for the carry path.
|
||||
|
||||
### Latency Reference Points
|
||||
|
||||
**FACT (Rocket Chip, UC Berkeley generator)**: Rocket Chip is in-order, and the MUL/DIV unit is configuration-dependent across Rocket's `Configs.scala` parameter set. In `RocketCoreConfig` (commonly cited as the default), the multiplier is pipelined with multiple pipeline stages (multi-cycle iterative, new operation accepted per cycle) and the divider is a radix-4 iterative divider. The exact stage counts and latencies are determined by parameters in `Configs.scala` and are not a single canonical value. INSUFFICIENT EVIDENCE for a specific latency number without naming the configuration.
|
||||
|
||||
**FACT (BOOM v2/v3, UC Berkeley)**: Out-of-order superscalar. MUL is pipelined (latency configuration-dependent). DIV is variable-latency, non-pipelined.
|
||||
|
||||
**FACT (XiangShan, open-source OoO RISC-V)**: MUL is pipelined; DIV is variable-latency, non-pipelined.
|
||||
|
||||
**FACT (Ibex, lowRISC)**: In-order. MUL is implemented as a single-cycle or short-pipeline combinational multiplier in some configurations; iterative DIV.
|
||||
|
||||
**INSUFFICIENT EVIDENCE**: Arm Cortex-A77 per-instruction MUL/DIV latencies are not publicly published by Arm. No specific numbers are cited from a primary source.
|
||||
|
||||
**INSUFFICIENT EVIDENCE**: Intel Haswell integer divider internal radix (radix-16 vs. radix-32) is not established from publicly verifiable primary sources. The design is widely reported to be a high-radix shift-subtract divider rather than a Newton-Raphson divider, but the specific radix is INSUFFICIENT EVIDENCE.
|
||||
|
||||
### Divide Algorithms
|
||||
|
||||
**FACT**: Four primary classes are considered here:
|
||||
|
||||
1. **Digit-recurrence / shift-subtract (radix-2, radix-4, radix-16, radix-64 SRT)**: radix-2 takes 64 cycles worst case; radix-4 takes ~32–33 cycles; radix-16 takes ~16 cycles; radix-64 takes ~8–10 cycles (with significant area/complexity cost). Iterative, small-to-moderate area depending on radix.
|
||||
2. **Newton-Raphson reciprocal multiplication**: multiple iterations of a multiply-based refinement to compute the reciprocal, then a final correction multiply to produce the quotient. Latency depends on initial seed precision and convergence criteria.
|
||||
3. **Goldschmidt**: similar convergence behavior to Newton-Raphson, with a different iteration structure.
|
||||
4. **CORDIC-based and series-expansion dividers**: rarely used for general-purpose integer divide due to overhead; exist as alternative approaches.
|
||||
|
||||
**FACT**: The RISC-V M extension does not require IEEE-754 compliant division; integer rounding is sufficient, simplifying hardware.
|
||||
|
||||
## Existing Approaches
|
||||
|
||||
### A1. Iterative Shift-Subtract Divider (Radix-2)
|
||||
|
||||
- **Datapath**: 128-bit (64 remainder + 64 divisor) shifter/subtractor, radix-2, 64-cycle worst case.
|
||||
- **DIV Latency**: 64 cycles.
|
||||
- **DIV Throughput**: 1 divide per 64 cycles (not pipelined).
|
||||
- **MUL coverage**: A1 defines a divider only; MUL is not part of this organization.
|
||||
- **Area (DIV only)**: Smallest divider; typically a few kGE plus control. **INSUFFICIENT EVIDENCE** for a specific gate count.
|
||||
- **Power**: Lowest of the divider options when idle; minimal toggle rate per non-dividing cycle.
|
||||
- **Used in**: Some low-end in-order cores; some configurations of Rocket Chip with smaller radix.
|
||||
|
||||
### A2. Pipelined Iterative Divider
|
||||
|
||||
- **Datapath**: Two distinct subclasses must be distinguished:
|
||||
- **(A2a) Pipelined iterative loop**: a single shift-subtract array with pipeline registers inserted at one or more points within the iterative loop, allowing a new operation to enter the loop every cycle after the pipeline is filled. Latency remains 64 cycles; throughput is 1 per cycle after fill.
|
||||
- **(A2b) Fully unrolled divider**: 64 shift-subtract stages with pipeline registers between every stage, giving latency 64 cycles and throughput 1 per cycle from the first cycle. Area is roughly 64× the A1 datapath.
|
||||
- **DIV Latency**: 64 cycles.
|
||||
- **DIV Throughput**: 1 per cycle (after fill for A2a; from cycle 1 for A2b).
|
||||
- **Area**: A2a is ~2–3× A1 (a few extra pipeline registers); A2b is roughly 64× A1 and is rarely used.
|
||||
- **Power**: Higher toggle rate than A1; only worthwhile under sustained divide streams.
|
||||
- **Used in**: Rare; mostly in high-throughput streaming dividers (DSP). Uncommon in general-purpose cores.
|
||||
|
||||
### A3. Newton-Raphson Divider
|
||||
|
||||
- **Datapath**: shared MUL unit(s), initial reciprocal seed ROM, multiplier used in iterative refinement and one final correction multiply.
|
||||
- **Latency breakdown**: seed table lookup (1 cycle) + N refinement multiplies (typically 2–3 iterations for 64-bit integer) + 1 final correction multiply. **ASSUMPTION**: 3 refinement iterations + 1 correction is typical for 64-bit; the concrete iteration count depends on seed precision and convergence criteria. **INSUFFICIENT EVIDENCE** for a single canonical iteration count without specifying the seed table and refinement schedule.
|
||||
- **Throughput**: one divide per (N+1) MUL cycles **only if the multiplier is dedicated to the divider**. If the multiplier is shared with the main MUL datapath, throughput is degraded by contention with MUL issue rate and is workload-dependent. **INSUFFICIENT EVIDENCE** for a single throughput number in the shared case.
|
||||
- **Area**: 1 reciprocal seed ROM + 1–2 MUL units (sharing possible at the cost of contention); large.
|
||||
- **Power**: Higher static and dynamic (multiplier active during divide refinement).
|
||||
- **Used in**: Some high-performance FPU designs for floating-point; less common for dedicated integer divide.
|
||||
|
||||
### A4. Radix-16 / Radix-64 SRT Divider
|
||||
|
||||
- **Datapath**: high-radix recurrence with a quotient-digit lookup table (PLA or ROM), redundant remainder representation. 16 or 64 bits processed per cycle.
|
||||
- **DIV Latency**: ~16 cycles (radix-16) or ~8–10 cycles (radix-64) for 64-bit operands.
|
||||
- **DIV Throughput**: 1 per cycle if pipelined; 1 per 16 / 8–10 cycles if iterative.
|
||||
- **Area**: large lookup table and complex datapath; PLA is a significant area contributor.
|
||||
- **Power**: high; many bits toggle per cycle.
|
||||
- **Used in**: high-end x86 integer dividers (radix of the integer divider is **INSUFFICIENT EVIDENCE** from primary sources; widely reported as high-radix shift-subtract rather than Newton-Raphson).
|
||||
|
||||
### A5. Approximate / Lookup-Based Dividers (small operand ranges)
|
||||
|
||||
- Rarely used for 64-bit; not relevant to XH-1 unless a specific workload (e.g., 32-bit address arithmetic) dominates.
|
||||
|
||||
## Alternative Designs
|
||||
|
||||
### B1. Hybrid MUL + Sequential-Iterative DIV (radix-4 iterative divider, custom MUL)
|
||||
|
||||
- MUL: 64×64→128 pipelined in 1 stage (target), throughput 1 per cycle. This departs from Rocket Chip's default multi-stage MUL (Rocket's `RocketCoreConfig` issues MUL as a multi-cycle iterative operation); the 1-stage MUL is an aggressive target for XH-1.
|
||||
- DIV: radix-4 iterative, ~33 cycles (32 cycles for quotient bits + finalization), blocking on the unit.
|
||||
- Single divider per core, 1 MUL pipeline stage.
|
||||
|
||||
**NOTE on naming**: B1 inherits the radix-4 iterative divider pattern from Rocket and Ibex, but the MUL organization (1-stage pipelined) is a custom choice that does not match either Rocket's multi-stage MUL or Ibex's short-pipeline / combinational MUL. The "hybrid" descriptor refers to combining a pipelined MUL with an iterative DIV; the divider side aligns with reference designs, the MUL side does not.
|
||||
|
||||
### B1'. Two MUL Pipelines + Shared Iterative DIV (wide-issue variant)
|
||||
|
||||
- Two pipelined MUL datapaths, one shared radix-4 DIV datapath.
|
||||
- MUL throughput: 2/cycle (sustained, independent operands).
|
||||
- DIV throughput: 1 per ~33 cycles, shared and blocking.
|
||||
- **Area**: ASSUMING a MUL:DIV area ratio of approximately 1:1 (i.e., one MUL pipeline and one radix-4 iterative DIV are roughly comparable in area, since a 1-stage 64×64 MUL is a Wallace/Dadda tree plus a 128-bit CPA, and a radix-4 iterative DIV is a ~66-bit adder plus control state), the combined (MUL+DIV) area scales as 1 + 1 = 2× B1 (two MULs plus one shared DIV, where B1 has one MUL and one DIV). The 1.5× figure previously cited assumed MUL is half the area of DIV, which is unsupported. The 1.4× lower bound previously cited is unsupported. The corrected qualitative statement is: **~2× B1**, with the explicit MUL:DIV area ratio assumption stated. **INSUFFICIENT EVIDENCE** for a synthesis-derived number.
|
||||
|
||||
### B2. Combined MUL/MAC (Multiply-Accumulate) Unit
|
||||
|
||||
- 64×64→128 with an additional accumulator register; supports future RV MAC extensions (proposed in some bitmanip drafts).
|
||||
- Adds area over pure MUL; **INSUFFICIENT EVIDENCE** for a specific percentage.
|
||||
- Useful for DSP, ML, and HPC; 128-core replication benefits workloads with matrix-style kernels.
|
||||
|
||||
### B3. Operand-Width-Detected 32-bit Fast MUL Path (also serves MULW)
|
||||
|
||||
- Detect when both operands of a 64-bit MUL are sign- or zero-extended from 32 bits (i.e., bit 31 is replicated through bit 63), and route the multiplication through a 32×32→64 fast multiplier.
|
||||
- The same 32×32→64 datapath also implements MULW directly: the 32-bit product (lower 32 bits of the 64-bit result) is sign-extended to 64 bits and written to `rd`. This sharing is essentially free in area terms and means that adopting B3 is the natural way to implement MULW.
|
||||
- This is **distinct from MULW as a workaround** for the 32-bit-case: B3 accelerates 64-bit MUL / MULH / MULHSU / MULHU when both operands happen to be 32-bit sign- or zero-extended, a pattern common after `lw` / `lwu` followed by arithmetic. MULW alone does not accelerate this case because MULW is a different instruction with different result semantics.
|
||||
- **Position**: B3 is the natural choice for the MULW datapath and adds a 32-bit-extended-operand fast path for 64-bit MUL. The verification cost is the dual datapath (32-bit and 64-bit) and the operand-width classifier. At the 128-core replication level, the area overhead is bounded by the 32-bit datapath size, which is much smaller than the 64-bit datapath; the dominant question is verification cost, not area.
|
||||
- **Decision**: B3 is recommended as part of the MULW implementation; deferring B3 means deferring the natural MULW implementation, which is not viable if MULW is in the ISA. The B3-vs-B4 trade-off therefore applies to the 32-bit-extended-operand fast path, not to MULW support.
|
||||
|
||||
### B4. Bypassable Output with Operand Width Detection (full 64×64→128 only)
|
||||
|
||||
- Always implement full 64×64→128; the decoder examines the instruction to forward only the required 64 bits.
|
||||
- Verification: a single source of truth (full 128-bit datapath), reducing dual-mode bugs.
|
||||
- Disadvantage: no area saving; MULW must still be implemented separately (e.g., via B3 or a dedicated 32×32→32 path).
|
||||
- **Position**: B4 conflicts with the natural MULW-via-B3 sharing above; if MULW is required, B4 is not a complete solution. If MULW is not required, B4 simplifies verification at the cost of a slightly larger MUL datapath for MULW-equivalent work.
|
||||
|
||||
### B5. Skip-on-Zero / Divide-Cancellation Optimizations
|
||||
|
||||
- Detect zero dividend (quotient is zero, remainder is dividend) and divide-by-one (quotient is dividend, remainder is zero) at the front end and forward the result without entering the iterative loop.
|
||||
- **Corner cases the fast path must handle correctly** (consistent with the iterative path):
|
||||
- dividend = 0, divisor = anything (including 0): quotient = 0, remainder = 0 (dividend).
|
||||
- divisor = 1: quotient = dividend, remainder = 0.
|
||||
- divisor = −1:
|
||||
- dividend = 2^(XLEN−1) (most-negative): quotient = 2^(XLEN−1) (dividend), remainder = 0 (signed overflow case, the M-extension rule).
|
||||
- all other dividends: quotient = −dividend (two's complement negation), remainder = 0.
|
||||
- The fast path must explicitly implement these rules; it is not a simple "if divisor=±1 then return dividend" because of the signed-overflow case for divisor = −1.
|
||||
- Saves latency in the common case for some workloads; trivial area overhead.
|
||||
- Composable with any divider organization.
|
||||
|
||||
### B6. Software Divide-by-Constant Transformation (cross-cutting)
|
||||
|
||||
- Compilers transform division by a **compile-time** constant into a multiply-by-reciprocal sequence. The hardware DIV is then needed only for division by variables.
|
||||
- Reduces effective DIV frequency significantly for workloads with constant denominators; affects hardware sizing decisions. **INSUFFICIENT EVIDENCE** for a quantitative reduction without a specific workload profile.
|
||||
|
||||
### B7. Dedicated 32-bit Fast Divider for W-suffixed Instructions
|
||||
|
||||
- Implement a separate 32-bit radix-2 or radix-4 iterative divider for DIVW / DIVUW / REMW / REMUW. The 32-bit divider has half the iteration count (32 or 16 cycles vs. 64 or 33 for the 64-bit divider) and roughly a quarter of the datapath area.
|
||||
- Useful if profiling shows W-suffixed divides dominate; in most general-purpose workloads they do not.
|
||||
- Verification cost: dual divider datapath, similar to B3.
|
||||
- **Note**: W-suffixed instructions operate on 32-bit operands; a natural alternative is to share the 64-bit divider datapath with the 32-bit operands on the lower 32 bits, taking 32 or 16 cycles of the 64-bit divider's iteration. B7 is only worth its area if the latency savings matter.
|
||||
|
||||
### B8. Combined MUL / DIV with Shared Partial-Product Array (CSA Sharing)
|
||||
|
||||
- Reuse the MUL's CSA compressor tree as the final correction multiplier for an SRT or Newton-Raphson divider, sharing the most area-intensive block.
|
||||
- Reduces the area penalty of A3 / A4 at the cost of tighter verification coupling between MUL and DIV paths.
|
||||
- Real architectural option in some high-performance designs; not considered in the prior revision.
|
||||
|
||||
### B9. Combined MUL / DIV with Shared Final Carry-Propagate Adder (CPA Sharing)
|
||||
|
||||
- Share only the final 128-bit carry-propagate adder between the MUL datapath and the DIV's correction-multiply step (or the DIV's final-cycle remainder correction). The compressor tree, partial-product generation, and divider iteration state remain separate.
|
||||
- Lower area savings than B8 (CSA is shared instead of CPA), but much simpler verification: the shared CPA is a single combinational block, and the MUL vs DIV datapaths feeding it are independent.
|
||||
- The MUL pipeline register naturally sits between the compressor-tree output and the CPA, which means the CPA itself can be shared at the output side without disrupting the MUL pipeline structure.
|
||||
- **Open question**: whether the MUL pipeline register sits at the compressor-tree output (CPA in cycle 2) or at the CPA output (CPA in cycle 1) determines the CPA's pipeline stage. See Implementation Considerations for the canonical placement decision.
|
||||
|
||||
### B10. Partially Unrolled Radix-4 Divider (intermediate between A2a and A2b)
|
||||
|
||||
- Unroll the radix-4 iterative divider into a small number of pipeline stages (e.g., 8 or 16 stages) rather than the full 64 (A2b) or the single iterative loop (A2a).
|
||||
- Latency 32 or 33 cycles (one stage per two quotient bits for 8 stages, or one stage per quotient bit for 16 stages), throughput 1 per cycle after fill, area roughly 8× or 16× A1.
|
||||
- A meaningful intermediate option for the high-DIV-throughput case that the prior revision did not consider.
|
||||
- Verification: same as A2a (iterative datapath with pipeline registers); the unrolling does not introduce new corner cases.
|
||||
|
||||
## Comparison
|
||||
|
||||
| Design | MUL Latency | MUL Tput | DIV Latency | DIV Tput | Rel. Area (MUL+DIV) | Power (active) | Verification Complexity |
|
||||
|--------|-------------|----------|-------------|----------|----------------------|----------------|--------------------------|
|
||||
| A1 (radix-2 iter. DIV only) | not defined in A1 | not defined in A1 | 64 cycles | 1 per 64 cycles | INSUFFICIENT EVIDENCE (combined) | Lowest (DIV only) | Low |
|
||||
| A2a (pipelined iter. loop) | not defined in A2a | not defined in A2a | 64 cycles | 1 per cycle (after fill) | INSUFFICIENT EVIDENCE | Med | Med |
|
||||
| A2b (fully unrolled) | not defined in A2b | not defined in A2b | 64 cycles | 1 per cycle | INSUFFICIENT EVIDENCE (very large) | High | High |
|
||||
| A3 (Newton-Raphson) | 1 cycle | 1 per cycle (dedicated) | N+1 MUL cycles (dedicated); workload-dep. if shared | INSUFFICIENT EVIDENCE (shared) | high | High | High |
|
||||
| A4 (Radix-16/64 SRT) | not defined in A4 | not defined in A4 | 8–16 cycles | 1 per cycle if pipelined | high | High | High |
|
||||
| B1 (radix-4 DIV, 1-stage MUL) | 1 cycle (target) | 1 per cycle | ~33 cycles | 1 per 33 cycles | baseline | Low–Med | Low–Med |
|
||||
| B1' (2× MUL + 1× shared DIV) | 1 cycle | 2 per cycle | ~33 cycles | 1 per 33 cycles (shared) | ~2× B1 (INSUFFICIENT EVIDENCE; assumes MUL:DIV area ≈ 1:1) | Med | Med |
|
||||
| B2 (+ MAC) | 1 cycle | 1 per cycle | ~33 cycles | 1 per 33 cycles | +unspecified % over B1 (INSUFFICIENT EVIDENCE) | Med | Med |
|
||||
| B3 (32-bit fast MUL; serves MULW) | 1 cycle (32-bit path) | 1 per cycle | ~33 cycles | 1 per 33 cycles | +small (INSUFFICIENT EVIDENCE) | Low | Med (dual mode) |
|
||||
| B4 (full 64 only; MULW separate) | 1 cycle | 1 per cycle | ~33 cycles | 1 per 33 cycles | same as B1 (MULW path TBD) | Med | Lowest (MUL side); Med (MULW) |
|
||||
| B5 (skip-on-zero) | n/a | n/a | reduced in common case | same as base | negligible overhead | n/a | Low |
|
||||
| B7 (32-bit fast DIV) | 1 cycle | 1 per cycle | ~16–17 cycles (32-bit) | 1 per 16–17 cycles (32-bit only) | +small (INSUFFICIENT EVIDENCE) | Low | Med (dual mode) |
|
||||
| B8 (shared MUL/DIV CSA) | 1 cycle | 1 per cycle | A3/A4 latency | A3/A4 throughput | lower than A3/A4 alone (INSUFFICIENT EVIDENCE) | Med–High | Med–High (tight coupling) |
|
||||
| B9 (shared MUL/DIV CPA) | 1 cycle | 1 per cycle | A3/A4 latency | A3/A4 throughput | slightly higher than B8 reduction (INSUFFICIENT EVIDENCE) | Med | Low (clean separation) |
|
||||
| B10 (partially unrolled radix-4) | not defined in B10 | not defined in B10 | ~32–33 cycles | 1 per cycle (after fill) | ~8–16× A1 (INSUFFICIENT EVIDENCE) | Med | Med |
|
||||
|
||||
**ASSUMPTION**: Relative area figures are qualitative orderings based on published reference designs cited in the Sources section. Actual XH-1 synthesis numbers are **INSUFFICIENT EVIDENCE** pending RTL implementation.
|
||||
|
||||
## Advantages
|
||||
|
||||
### A1 / B1
|
||||
- Smallest area; lowest per-core replication cost.
|
||||
- Lowest static and dynamic power in the typical case (most cycles no MUL/DIV active).
|
||||
- Simpler verification; one mode of operation.
|
||||
- Well-understood reference implementation (Rocket, Ibex) for the divider side; the MUL side departs from both references.
|
||||
|
||||
### A4 (High-Radix SRT)
|
||||
- Lowest DIV latency among the iterative-style options; competitive with A3 on a single divide.
|
||||
- Pipelined variant gives 1 per cycle DIV throughput.
|
||||
|
||||
### A3 (Newton-Raphson)
|
||||
- Low latency for sequential divides — important for cryptography (RSA, ECC modular reduction) and HPC.
|
||||
- Amortizes multiplier cost if MAC (B2) is also desired.
|
||||
- Disadvantage: requires multiple refinement iterations; integer multiplier is large and contention with the main MUL datapath is a concern.
|
||||
|
||||
### B2 (MAC)
|
||||
- Enables future-proofing for proposed bitmanip and MAC extensions.
|
||||
- Helpful for matrix multiplication kernels running across 128 cores.
|
||||
|
||||
### B3 (32-bit fast MUL; serves MULW)
|
||||
- Natural implementation of MULW; the 32×32→64 datapath produces the MULW result by sign-extending the lower 32 bits.
|
||||
- Accelerates 64-bit MUL on 32-bit-valued operands (a common pattern after `lw`/`lwu` + arithmetic).
|
||||
- Area overhead is bounded by the 32-bit datapath size (much smaller than the 64-bit datapath).
|
||||
|
||||
### B5 (Skip-on-Zero)
|
||||
- Negligible area; reduces effective DIV latency for common cases.
|
||||
|
||||
### B7 (32-bit fast DIV)
|
||||
- Halves the divider iteration count for the W-suffixed instructions at modest area cost.
|
||||
|
||||
### B8 (Shared MUL/DIV CSA)
|
||||
- Reduces the area penalty of high-performance dividers by sharing the most area-intensive block.
|
||||
|
||||
### B9 (Shared MUL/DIV CPA)
|
||||
- Lower verification cost than B8; the shared CPA is a single combinational block and the MUL/DIV datapaths feeding it are independent.
|
||||
|
||||
### B10 (Partially Unrolled Radix-4)
|
||||
- A meaningful intermediate option for high DIV throughput without the area cost of A2b; same verification profile as A2a.
|
||||
|
||||
## Disadvantages
|
||||
|
||||
### A1 / B1
|
||||
- Variable DIV latency complicates scheduling in an in-order core; in-order cores stall ~33 cycles per divide.
|
||||
- Worst-case latency hits interrupt latency, branch-target computation if the MUL/DIV is on a branch path.
|
||||
- Not a fit for sustained divide-bound workloads (e.g., modulo arithmetic in bignum).
|
||||
|
||||
### A3
|
||||
- Area at 128-core replication is severe; the multiplier is one of the largest blocks in a typical core.
|
||||
- Power: a 64×64 multiplier running 1 per cycle is one of the highest-power blocks in a typical core.
|
||||
- Verification: Newton iteration precision (must converge to within 1 ULP across all operands) is notoriously subtle.
|
||||
- If multiplier is shared with main MUL datapath, throughput is workload-dependent, not the N+1 figure cited for the dedicated case.
|
||||
|
||||
### A4
|
||||
- Large lookup table (PLA or ROM); area and power dominated by the table.
|
||||
- Verification: complex quotient-digit selection logic.
|
||||
- Not commonly used outside high-end commercial designs.
|
||||
|
||||
### B3
|
||||
- Verification: dual datapath (32-bit and 64-bit) must be exhaustively cross-checked.
|
||||
- The classification logic (operand-width detection) is itself a source of corner-case bugs.
|
||||
- **B3 does not subsume MULW for the purpose of "B3 is redundant"**: B3 and MULW are two different ways to access 32-bit-multiplication, and B3's value is primarily as the natural MULW implementation and secondarily as the 32-bit-extended-operand fast path.
|
||||
|
||||
### B2
|
||||
- Adds a third writeback port or accumulator path; competes with the FPU/LSU for writeback slots in a multi-issue core.
|
||||
|
||||
### B5
|
||||
- Only helps the specific cases of zero dividend or divisor ±1; other optimizations (e.g., division by small powers of two) are already handled by the base I extension's shift instructions.
|
||||
- The fast path must explicitly handle the signed-overflow corner case (dividend = 2^(XLEN−1), divisor = −1): quotient = dividend, remainder = 0.
|
||||
|
||||
### B7
|
||||
- Verification: dual divider datapath; same concerns as B3.
|
||||
- Area saving is moot if W-suffixed divides are not on the critical path.
|
||||
|
||||
### B8
|
||||
- Verification: tighter coupling between MUL and DIV paths makes corner-case analysis more difficult.
|
||||
|
||||
### B9
|
||||
- Area savings are smaller than B8 (CPA shared instead of CSA); the savings are bounded by the CPA size, which is significant but not the dominant block.
|
||||
|
||||
### B10
|
||||
- Area scales with the unroll factor; 16-stage unroll is ~16× A1, which is meaningful at 128-core replication.
|
||||
|
||||
## XH-1 Considerations
|
||||
|
||||
**PROPOSAL**: The XH-1 core should adopt **B1 (Hybrid MUL + Iterative DIV)** as the baseline, with the following parameters, pending RTL validation:
|
||||
|
||||
- **MUL**: 64×64→128, target 1-cycle pipelined (1 stage of pipeline registers), throughput 1 per cycle. The pipeline register sits at the **CPA output** (after the compressor tree and the 128-bit CPA in cycle 1), so the entire compressor-tree-plus-CPA path is in one cycle. This is the "1-cycle MUL" interpretation; the fallback-A interpretation splits the compressor tree and the CPA across two cycles. For 2-issue or wider cores, scale to B1' (two MUL pipelines sharing one DIV).
|
||||
- **DIV**: Radix-4 shift-subtract, ~33 cycles worst case (32 cycles for quotient bits + finalization), blocking, non-pipelined. The prior revision's "33–35 cycles worst case" and the "33–64 cycle" range are reconciled here: 33 is the radix-4 bound; 64 corresponds to radix-2, which is a different algorithm choice.
|
||||
- **REM**: Reuse the DIV datapath; remainder is a by-product.
|
||||
- **Bypass paths**: From MUL output directly to subsequent dependent ALU/branch in the next cycle. **INSUFFICIENT EVIDENCE** on whether this fits the XH-1 pipeline depth; depends on the integration context.
|
||||
- **B5 (skip-on-zero)**: implement at the front end of the DIV unit; negligible overhead. The fast path must handle the signed-overflow corner case (divisor = −1, dividend = 2^(XLEN−1)) by returning quotient = dividend, remainder = 0.
|
||||
- **B3 (32-bit fast MUL, also serves MULW)**: include the 32×32→64 datapath as the MULW implementation path. The same datapath accelerates 64-bit MUL on 32-bit-extended operands. Verification cost: dual datapath, manageable with the B3 corner-case set (operand-width detection, sign-extension patterns).
|
||||
- **B7 (32-bit fast DIV)**: Defer; the W-suffixed divide is not assumed to be on the critical path. Revisit if profiling shows otherwise.
|
||||
|
||||
**ASSUMPTION**: A 1-cycle MUL latency is achievable in the target process. **INSUFFICIENT EVIDENCE** on the XH-1 target process node and clock period. A full 64×64→128 Wallace/Dadda tree + 128-bit carry-propagate adder in a single cycle is at the edge of feasibility for high-performance designs; typical in-order cores implement MUL as either a multi-cycle iterative multiplier or a multi-stage pipelined multiplier. **PROPOSAL**: Validate via synthesis at the target corner before committing. If the 1-cycle critical path cannot be closed, **fallback options** are:
|
||||
- **B1-fallback-A**: 2-cycle pipelined MUL. The pipeline register sits at the **compressor-tree output** (splitting the compressor tree in cycle 1 from the CPA in cycle 2); latency 2 cycles, throughput 1 per cycle, modest area overhead (one extra pipeline register).
|
||||
- **B1-fallback-B**: Multi-cycle iterative MUL (Booth-encoded). The cycle count for a Booth-encoded iterative multiplier on 64×64 is implementation-dependent; **INSUFFICIENT EVIDENCE** for a specific number of cycles. Lower area, lower throughput, higher latency.
|
||||
- **B1-fallback-C**: Retain the 1-stage MUL architecture but lower the target clock frequency (system-level decision, not unit-level).
|
||||
|
||||
**MUL/DIV resource sharing and contention (single-issue-lane B1)**: Under B1, the MUL pipeline and the DIV datapath share the integer execution lane's issue slot, register-file read ports, and writeback port. The single execution lane can issue either one MUL per cycle (latency 1) or one DIV (latency ~33) at a time, but not both simultaneously. If a MUL is issued while a DIV is in progress:
|
||||
- The MUL occupies the issue slot for 1 cycle; the DIV's iterative state is held in the divider's internal registers and does not require the issue slot during its 33 cycles.
|
||||
- Register-file read ports: MUL requires 2 read ports for its 1 cycle; the DIV's operands are read at DIV issue and held in the divider's operand register. No contention after issue.
|
||||
- Writeback port: MUL writes back in cycle 2 (1-cycle latency); the DIV writes back on completion. A MUL issued in the same cycle as a DIV completion would contend for the writeback port. In a single-writeback-port lane, the DIV completion must be stalled by 1 cycle to let the MUL writeback, or vice versa. This is a 1-cycle throughput loss in the rare case of simultaneous MUL-and-DIV-completion, and is acceptable at 128-core scale (per-core throughput loss is small; aggregate is bounded by the per-core lane).
|
||||
- **For 128-core aggregate throughput**: the per-core limitation is the MUL/DIV issue slot, not the divider's iteration. Aggregate MUL throughput is bounded by 1/cycle/core = 128/cycle die-wide. Aggregate DIV throughput is bounded by 1/33 cycles/core = 128/33 ≈ 3.88 divides/cycle die-wide in the steady state if every core is issuing back-to-back independent divides. This is the steady-state upper bound under B1, not a typical workload figure. With B5 reducing some divides to 1 cycle, the aggregate is workload-dependent and typically lower.
|
||||
|
||||
**PROPOSAL**: The MUL/DIV unit sits on the integer execution lane, sharing the issue queue with ALU and branch. It has its own reservation station (or issue-slot tag) so the issue logic can distinguish short-latency ALU (1 cycle) from long-latency MUL (1, or 2 in the fallback) and variable-latency DIV (~33).
|
||||
|
||||
**PROPOSAL**: For the in-order case, the MUL/DIV unit is non-blocking on MUL (1-cycle latency) but blocking on DIV; the in-order front-end stalls on a structural hazard only when a second DIV is issued before the first completes.
|
||||
|
||||
**PROPOSAL**: For an out-of-order XH-1 core, the MUL/DIV unit exposes its variable DIV latency through the wakeup/select logic so that dependent instructions are replayed or re-issued with the correct ready signal.
|
||||
|
||||
**PROPOSAL**: The B-extension unit (if RV64B is implemented) shares **operand muxes, sign-handling logic, and bit-level muxes** with the MUL/DIV unit but does **not** share the Wallace/Dadda compressor tree. Operations like CLZ, CTZ, BSET, BEXT operate on individual bits or small bit-fields and do not naturally map onto a Wallace-tree multiplier datapath. CLMUL / CLMULH / CLMULR (Zbc) require an AND-tree / XOR-reduction datapath that is structurally distinct from both the Wallace-tree multiplier and the iterative divider; they do not share the compressor tree. Any apparent sharing is at the operand-fetch and writeback layers, not the core arithmetic.
|
||||
|
||||
## 128-Core Scalability
|
||||
|
||||
The 128-core replication factor is the dominant cost driver for the MUL/DIV unit.
|
||||
|
||||
**PROPOSAL**: At 128 cores, every MUL unit must be **power-gateable independently** so that cores not in use (DPM, dark-silicon) can be fully clock- and power-gated, including the MUL/DIV block.
|
||||
|
||||
**PROPOSAL**: The MUL/DIV unit's reservation-station entries, divider iteration state, and pipeline registers must support **state retention or clean-state entry** across power gating. Specifically, when a core is power-gated while a long-latency DIV is in flight, the DIV's mid-iteration state must be handled by one of the following:
|
||||
- (a) **Flush and re-issue**: the in-flight DIV is squashed architecturally (the issue queue entry is marked invalid, the divider's iteration state is discarded), and the DIV is re-fetched and re-issued from the I-cache after the core wakes up. This is architecturally transparent if the re-fetch / re-issue mechanism is present (standard OoO replay path); the architectural state is preserved because the DIV is re-executed from scratch. The cost is re-fetch latency after wakeup. The mechanism is "not acceptable" only if the re-fetch path is not implemented (e.g., a simple in-order core without replay).
|
||||
- (b) **Checkpoint to retention**: the divider's iteration state is saved to a retention register or to memory before power-down, and restored on wakeup. Preserves the in-flight DIV but requires retention storage proportional to the divider's state.
|
||||
- (c) **Block power-gating until completion**: power-gating is only allowed when the divider is idle. Simple, but defeats the purpose of DPM if DIV latency is long and frequent.
|
||||
- The choice depends on the XH-1 DPM policy and on whether the core is in-order or OoO. **INSUFFICIENT EVIDENCE** on the XH-1 DPM policy; the design must accommodate one of these options without committing to a specific approach here.
|
||||
|
||||
**OPEN QUESTION**: What is the actual MUL/DIV area share of the XH-1 core? Published RISC-V references suggest a typical MUL unit occupies a small single-digit percentage of a high-performance core's area, with iterative DIV adding additional area. The prior revision's "1–3%" and "1 MGE/core total" baseline are removed here as unsourced. **INSUFFICIENT EVIDENCE** on the XH-1-specific area share without synthesis.
|
||||
|
||||
**OPEN QUESTION**: Is the MUL/DIV area share dominated by the MUL datapath, the DIV datapath, the operand muxing, or the bypass network? Targeted synthesis required.
|
||||
|
||||
**INSUFFICIENT EVIDENCE**: The XH-1 die-area budget, process node, and clock period are not established in this document.
|
||||
|
||||
**PROPOSAL**: The MUL/DIV unit should be **identical across all 128 cores** (no per-core specialization) to simplify verification and physical design. The replication should be a Verilog/SystemVerilog generate block driven by a single template. **PROPOSAL**: The generate block itself must be included in the verification scope (not assumed trivially correct). At minimum: a lint-clean check, a synthesis-check that the generate block instantiates the correct number of cores, and a per-instance equivalence check on a sample of cores.
|
||||
|
||||
**PROPOSAL**: For 128× replicated MUL/DIV pipeline registers, ECC or parity protection should be considered for soft-error mitigation. **INSUFFICIENT EVIDENCE** on the XH-1 reliability target.
|
||||
|
||||
**PROPOSAL**: Reset distribution and scan chain architecture for 128× replicated MUL/DIV must be addressed at the integration level. The MUL/DIV unit's scan chains should support parallel or staggered scan-shift across cores to keep test time bounded. **INSUFFICIENT EVIDENCE** on the XH-1 DFT architecture.
|
||||
|
||||
**CONSIDERATION (clock and timing at 128× replication)**: If the XH-1 die includes a global clock-distribution network, the MUL/DIV critical path determines the local clock skew tolerance. A pipelined 1-cycle MUL has a short critical path (one 64-bit adder-equivalent) which is favorable; a non-pipelined iterative divider has a 64-bit adder in its loop, which is the cycle-time limiter. The clock-skew analysis should be revisited at the integration level once the XH-1 clock tree is defined.
|
||||
|
||||
**CONSIDERATION (cross-core aggregate throughput)**: Under B1, each core's blocking DIV delivers at most 1 per 33 cycles per core in the steady state. Across 128 cores, the aggregate is at most ~3.88 divides per cycle in the steady state, which assumes all cores are issuing back-to-back independent divides for the full 33 cycles each. This is an upper bound on aggregate throughput, not a typical workload figure. Real workloads do not exhibit this worst case; the relevant metric is the per-core latency, not aggregate. **INSUFFICIENT EVIDENCE** on whether the XH-1 DPM or interconnect imposes a global cap on simultaneous divide activity; this is a system-level question outside the MUL/DIV unit's scope.
|
||||
|
||||
**CONSIDERATION (operand distribution and interconnect)**: Replicating a Wallace tree 128× implies 128 sets of wide operand buses to/from the register file. The interconnect / operand-routing network cost scales with the MUL operand width and the number of cores. **PROPOSAL**: include operand-routing overhead in the area estimate, not just the MUL/DIV datapath itself.
|
||||
|
||||
## Performance Considerations
|
||||
|
||||
**PROPOSAL**: **MUL throughput of 1 per cycle is a target for a high-performance XH-1 core**, but is not architecturally non-negotiable. Low-end in-order cores (some Ibex configurations) implement MUL with throughput < 1 per cycle. B1 satisfies the high-performance target with a 1-stage pipelined multiplier; B1' extends to 2 per cycle for wide-issue. If the 1-cycle MUL cannot be closed at the target process, the throughput target remains 1 per cycle but the latency becomes 2 cycles (B1-fallback-A).
|
||||
|
||||
**PROPOSAL**: **DIV throughput of 1 per ~33 cycles (blocking) is acceptable** for most integer workloads. Workloads that require more (e.g., modular arithmetic in cryptography) will suffer; mitigation is left to software (see Software Considerations).
|
||||
|
||||
**ASSUMPTION**: XH-1 target workloads include a mix consistent with Embench / SPECint-class profiles, where MUL/DIV instructions are a small fraction of dynamic instruction count. **INSUFFICIENT EVIDENCE** on the actual XH-1 target workload mix and on the specific dynamic-instruction share of MUL/DIV. The prior revision's "<2% MUL/DIV" and "<0.5% DIV/REM" claims are removed as unsourced.
|
||||
|
||||
**OPEN QUESTION**: Does XH-1 target HPC or cryptography workloads where MUL/DIV is a larger share of the instruction mix? If yes, the analysis shifts toward A3, A4, A2a, or B10. **NOTE**: ML workloads are dominated by floating-point multiplies on the FPU, not by integer MUL/DIV; integer MUL/DIV is relevant to ML only for quantization, address arithmetic, and integer embeddings.
|
||||
|
||||
## Area Considerations
|
||||
|
||||
**PROPOSAL**: Budget the MUL/DIV unit at a small single-digit percentage of single-core area for the B1 design, pending synthesis. The exact percentage is **INSUFFICIENT EVIDENCE**.
|
||||
|
||||
**OPEN QUESTION**: Is the XH-1 core area budget (without MUL/DIV) known? The MUL/DIV area share can be re-expressed as a percentage of the entire 128-core die only when this is established.
|
||||
|
||||
**ASSUMPTION**: A1-style radix-2 DIV is the smallest practical divider; A3 (Newton) and A4 (high-radix SRT) are larger, with A4 typically the largest. **INSUFFICIENT EVIDENCE** on the XH-1-specific gate-count budget.
|
||||
|
||||
## Power and Energy Considerations
|
||||
|
||||
**FACT**: A 64×64→128 multiplier toggles a large number of bits per cycle when active; its dynamic power is proportional to operand activity factor. Power-gating the multiplier when idle is essential.
|
||||
|
||||
**PROPOSAL**: The MUL/DIV unit should support **fine-grained clock gating** at the operand-mux boundary so that when no MUL/DIV is in flight, the entire unit is clock-gated to 0 toggle rate.
|
||||
|
||||
**PROPOSAL**: The MUL pipeline register should be a clock-gated scan flop, not a latch, to simplify DFT.
|
||||
|
||||
**OPEN QUESTION**: Does the XH-1 power budget allow all 128 MUL/DIV units to be active simultaneously? If not, per-core DPM must enforce that no more than N MUL/DIV units are in active divide at once. This is a system-level power-management question, not a unit-level one, but it constrains the unit's design.
|
||||
|
||||
**INSUFFICIENT EVIDENCE**: Specific per-MUL or per-DIV energy numbers for the XH-1 process are not established. The prior revision's "single-digit pJ in 7 nm" claim is removed as unsourced; per-MUL energy in advanced processes is implementation-dependent and varies by an order of magnitude or more depending on architecture and clock frequency.
|
||||
|
||||
## Implementation Considerations
|
||||
|
||||
**PROPOSAL**: Implement MUL as a **Wallace/Dadda tree of 4:2 compressors** feeding a final 128-bit carry-propagate adder. The 1-cycle latency budget accommodates the entire critical path from operand register → compressor tree → CPA → output register, with the pipeline register at the CPA output. **INSUFFICIENT EVIDENCE** on whether a Wallace/Dadda tree for 64×64 partial products is achievable in one cycle at the XH-1 target clock period; the specific tree depth and CPA depth are implementation-dependent and not asserted as fixed numbers here. The prior revision's "Tree depth 6–7; final CPA ~6 gates deep" is removed as unsourced and implementation-specific.
|
||||
|
||||
**MUL pipeline register placement** (clarified):
|
||||
- **Primary (1-stage)**: register at CPA output. The compressor tree and the 128-bit CPA are both in cycle 1.
|
||||
- **Fallback A (2-stage)**: register at compressor-tree output. The compressor tree is in cycle 1, the 128-bit CPA is in cycle 2.
|
||||
- The placement determines the critical path per cycle: in the primary, the per-cycle critical path is tree + CPA; in fallback A, the per-cycle critical path is max(tree, CPA). Fallback A is the natural way to close timing if tree + CPA exceeds the target clock period in a single cycle.
|
||||
|
||||
**PROPOSAL**: For MULHSU, the standard implementation generates a 64×64 partial-product array with the unsigned operand's partial products zero in the upper half and the signed operand's final partial-product row sign-extended (or added in inverted-and-carry form) into the reduction tree. The 128-bit-wide signed multiplier datapath is an alternative but doubles the multiplier width and is rarely used.
|
||||
|
||||
**PROPOSAL**: Implement DIV as a **radix-4 shift-subtract** with restoring on negative remainder. 32 cycles for quotient bits + 1 cycle for finalization = ~33 cycles total. The remainder is the "running remainder" register after the 33rd cycle; the quotient is the shifted-out bits. REM is selected by muxing the appropriate output.
|
||||
|
||||
**PROPOSAL**: B5 (skip-on-zero): add a front-end detector on the DIV operands that forwards the result directly for divisor = ±1 or dividend = 0, bypassing the iterative loop. The fast path must handle the signed-overflow corner case (dividend = 2^(XLEN−1), divisor = −1) by returning quotient = dividend, remainder = 0, consistent with the M-extension rule. Trivial area; reduces effective DIV latency for common cases.
|
||||
|
||||
**PROPOSAL**: Parameterize the MUL width with a Verilog `parameter WIDTH=64` so that the same RTL can be synthesized for 32-bit-only cores (test variant, debug, or soft-IP reuse) without manual rework. Note that WIDTH=32 does not by itself support the MULW instruction semantics in RV64 (which performs a 32×32→64 multiply and sign-extends the 32-bit result); the B3 32×32→64 datapath implements MULW by taking the lower 32 bits and sign-extending to 64.
|
||||
|
||||
## Verification Considerations
|
||||
|
||||
**FACT**: MUL and DIV are among the most heavily tested instructions in any RISC-V core; their verification is dominated by **corner-case operand pairs** (e.g., 0, 1, -1, 2, -2, 0x7FFF…, 0x8000…, 0xFFFF…, MSB=1 transitions). For DIV/REM, the corner cases include the division-by-zero and signed-overflow rules defined in the ISA spec.
|
||||
|
||||
**PROPOSAL**: Adopt a **directed-random + constrained-random** verification flow using a golden reference model:
|
||||
- **Reference MUL**: SystemVerilog `bit [127:0]` (or DPI-C to a software bigint). The reference produces the full 128-bit product; the checker compares the appropriate bits of the 128-bit result against the architectural result:
|
||||
- MUL, MULH, MULHU, MULHSU: lower 64 or upper 64 bits of the 128-bit product, with sign-handling per the ISA spec.
|
||||
- MULW: lower 32 bits of the 64-bit product (where the 64-bit product is computed on a 32×32 signed multiplication, with the 32-bit result sign-extended to 64 bits and written to `rd`). The reference is a 32×32 signed multiply that produces a 32-bit result, which is then sign-extended to 64 bits and compared against the architectural `rd` value. The 128-bit reference path is used for MUL/MULH/MULHU/MULHSU, not for MULW.
|
||||
- **Reference DIV**: Software bigint that produces (quotient, remainder) per the RISC-V M-extension spec, including the division-by-zero and signed-overflow rules for both quotient-producing and remainder-producing instructions.
|
||||
- **Coverage**: 100% of the corner cases (above), plus cross-product of small operand values (±1, ±2, ±2^32, ±2^63, ±2^64-1), plus randomized large operands.
|
||||
- **Regression list size**: The prior revision's "64 hand-crafted corner cases" is removed as a specific number; the regression list should be sized to cover the documented corner cases and is grown as bugs are found. **INSUFFICIENT EVIDENCE** for a canonical count.
|
||||
|
||||
**PROPOSAL**: At the 128-core replication level, **per-core functional verification** of the MUL/DIV unit is unnecessary if the per-core RTL is identical and constrained-random coverage is closed at the template level (via a SystemVerilog generate block). This covers functional equivalence at the unit level. **However**, the verification of the generate block itself, and the interaction between per-core clock-gating / power-state and MUL/DIV state (e.g., does a clock-gated MUL lose its pipeline state correctly across gating? does a power-gated divider leave the iteration counter in a valid state for resumption, or is the DIV flushed and re-issued as in option (a) of the 128-core power-gating proposal?), must be verified explicitly. **Per-core physical / timing verification is not bypassed**: timing, DFT, and physical-design closure are verified at the integration level on a representative core and assumed replicated, with explicit per-die variation analysis as required by the XH-1 physical-design flow.
|
||||
|
||||
**PROPOSAL**: If the XH-1 verification flow includes formal property checking, the DIV correctness property ("quotient × divisor + remainder == dividend" within the RISC-V M extension's special cases) is a clean formal target and should be specified. **INSUFFICIENT EVIDENCE** on whether formal property checking is in scope.
|
||||
|
||||
## Software Considerations
|
||||
|
||||
**PROPOSAL**: Document the MUL/DIV latencies (1 cycle MUL, ~33 cycle DIV) in the XH-1 ABI / ISA reference manual so that compiler backends and library authors can schedule around the long-latency DIV.
|
||||
|
||||
**OPEN QUESTION**: Does the XH-1 ABI / linker convention include a software-emulated division routine for code that cannot tolerate the ~33-cycle blocking latency? If yes, the hardware DIV can be a single issue-slot, not a deeply pipelined one. The compiler can also apply divide-by-constant transformations (B6) to reduce effective hardware DIV frequency.
|
||||
|
||||
**PROPOSAL**: The MUL/DIV unit should raise no architectural exception; the RISC-V M extension defines all corner cases architecturally. This simplifies the trap logic.
|
||||
|
||||
**PROPOSAL**: For 128-core use cases, the MUL/DIV unit is the same in every core; the OS scheduler does not need to be aware of its microarchitectural details. Workload distribution across cores is independent of MUL/DIV design.
|
||||
|
||||
## Recommendation
|
||||
|
||||
**RECOMMENDATION**: Adopt **B1** — a hybrid MUL unit (target 1-cycle pipelined, full 64×64→128) combined with a radix-4 iterative, non-pipelined divider (~33 cycles), sharing no logic. Add **B5** (skip-on-zero, with the signed-overflow corner case handled) at the DIV front end. Add **B3** as the MULW implementation and the 32-bit-extended-operand fast path. Place the unit on the integer execution lane. Support full per-unit power and clock gating for 128-core DPM. For wide-issue cores, scale to **B1'** (two MUL pipelines sharing one DIV).
|
||||
|
||||
Rationale:
|
||||
1. Matches the dominant MUL workload profile (frequent, 1-cycle latency needed for high-performance targets).
|
||||
2. Keeps per-core area small, manageable at 128× replication.
|
||||
3. Avoids the verification burden of B8 (shared CSA) and the area burden of A3 (Newton-Raphson) and A4 (high-radix SRT).
|
||||
4. **B3 inclusion is required for MULW**, not optional: the same 32×32→64 datapath that accelerates 64-bit MUL on 32-bit-extended operands also implements MULW directly. Deferring B3 means deferring the natural MULW implementation.
|
||||
5. B5 is a near-free improvement to the common case, with the signed-overflow corner case handled correctly.
|
||||
6. The MUL/DIV design inherits the radix-4 iterative divider pattern from Rocket and Ibex; the MUL side departs from both references (1-stage pipelined vs. Rocket's multi-stage or Ibex's short-pipeline / combinational), which is an aggressive target that must be validated by synthesis.
|
||||
|
||||
**Fallback plan** (if the 1-cycle MUL cannot be closed at the target process / clock):
|
||||
- **Primary fallback**: B1-fallback-A — 2-cycle pipelined MUL with the pipeline register at the compressor-tree output. Latency 2 cycles, throughput 1 per cycle, modest area overhead. B1 architecture preserved.
|
||||
- **Secondary fallback**: B1-fallback-B — multi-cycle iterative MUL (Booth-encoded). Cycle count implementation-dependent; **INSUFFICIENT EVIDENCE** for a specific number. Lower area, lower throughput, higher latency. DIV remains radix-4 iterative.
|
||||
- **Tertiary fallback**: Lower the target clock frequency at the system level.
|
||||
|
||||
This recommendation is **conditional on**:
|
||||
- The XH-1 target process supporting either a 1-cycle 64×64→128 MUL critical path (primary) or a 2-cycle pipelined MUL critical path (fallback A) at the target clock. Both must be validated by synthesis at the target corner; **INSUFFICIENT EVIDENCE** without target process specification.
|
||||
- XH-1 workload mix not being dominated by sequential DIV/REM (cryptography). If it is, escalate to A4 (radix-16 SRT), a pipelined iterative divider (A2a or B10), or a shared MUL/DIV organization (B8, B9).
|
||||
- The 128-core replication budget tolerating the cumulative MUL/DIV area; this requires a known single-core area budget, which is **INSUFFICIENT EVIDENCE**.
|
||||
- A clear DPM policy for handling in-flight MUL/DIV state across power gating.
|
||||
|
||||
If any of these conditions fails, re-open the design against the named fallback.
|
||||
|
||||
## Confidence
|
||||
|
||||
**Medium-High** for B1 as the baseline architectural pattern. **Low** for specific area, power, and energy numbers — those are estimates based on published RISC-V reference designs, not on XH-1 synthesis data. **Low** for any claim about the XH-1 target process, workload mix, or die size (these are **INSUFFICIENT EVIDENCE**). The corrected version removes specific unsourced quantitative claims and demotes several prior FACTs to ASSUMPTION or INSUFFICIENT EVIDENCE.
|
||||
|
||||
## Open Questions
|
||||
|
||||
1. What is the XH-1 target process node and clock period? This constrains MUL critical-path design and determines whether 1-cycle MUL is feasible.
|
||||
2. What is the XH-1 target workload mix? HPC, cryptography, or general-purpose? (ML is not primarily an integer MUL/DIV workload.)
|
||||
3. Is the XH-1 core in-order, out-of-order, or hybrid?
|
||||
4. Does XH-1 implement RV64B (bit-manipulation) extensions, including Zbc (CLMUL / CLMULH / CLMULR)? The B extension does not naturally share the Wallace-tree multiplier datapath; any sharing is at the operand-mux and writeback layers, not the compressor tree.
|
||||
5. What is the XH-1 single-core area budget? What is the die-level area budget for the 128-core fabric?
|
||||
6. Does XH-1 use a per-core DPM scheme, and how does it handle in-flight MUL/DIV state across power gating (retention, flush-and-re-issue, or block-power-gate-until-completion)?
|
||||
7. Is the MUL/DIV unit expected to support a future RV MAC extension, or is the project strictly RV64IM (or RV64IMB)?
|
||||
8. Is there a software-emulated division in the XH-1 libc / ABI that would tolerate a long-latency hardware DIV? Will the compiler apply divide-by-constant transformations (B6)?
|
||||
9. Does the XH-1 verification flow include formal property verification, and will it be applied to the MUL/DIV unit's "quotient × divisor + remainder == dividend" property?
|
||||
10. For wide-issue cores, is B1' (two MUL pipelines + one shared DIV) the target, or is single-MUL B1 sufficient?
|
||||
11. What is the XH-1 interconnect / operand-routing cost of replicating a wide MUL operand bus 128 times? Should this overhead be included in the MUL/DIV unit's area budget?
|
||||
12. What is the XH-1 reliability target? Does the MUL/DIV pipeline require ECC or parity protection against soft errors?
|
||||
13. What is the XH-1 DFT architecture for the 128× replicated MUL/DIV scan chains?
|
||||
|
||||
## Sources
|
||||
|
||||
- Waterman, Asanović, et al. (eds.), *The RISC-V Instruction Set Manual, Volume I: User-Level ISA, Document Version 20191213*, Chapter 7 (M Extension). RISC-V International. Canonical ISA reference; defines the M-extension corner-case rules for DIV/REM/DIVU/REMU/DIVW/REMW/DIVUW/REMUW and MULW.
|
||||
- Waterman, Asanović, et al. (eds.), *The RISC-V Instruction Set Manual, Volume I: User-Level ISA*, Chapter 16 (B Extension, including Zbc). Canonical ISA reference for CLMUL / CLMULH / CLMULR.
|
||||
- Asanović et al., "The Rocket Chip Generator," EECS Department, UC Berkeley, Technical Report UCB/EECS-2016-17, 2016. Reference for Rocket's MUL/DIV organization; configuration-dependent.
|
||||
- Celio, Patterson, Asanović, "BOOM v2: a superscalar out-of-order processor," 2017. Reference for BOOM's MUL/DIV design.
|
||||
- Ibex documentation, lowRISC. Reference for the in-order baseline.
|
||||
- Chen et al., "XiangShan: An Open-Source High-Performance RISC-V Core," 2022. Reference for high-performance RISC-V MUL/DIV design.
|
||||
- Parhami, *Computer Arithmetic: Algorithms and Hardware Designs*, Oxford University Press. Reference for shift-subtract, SRT, and Newton-Raphson divide algorithms.
|
||||
- Flynn, Oberman, *Advanced Computer Arithmetic Design*. Reference for high-radix divider design.
|
||||
- Weste, Harris, *CMOS VLSI Design: A Circuits and Systems Perspective*, 4th ed. Reference for Wallace/Dadda multiplier design and critical-path analysis; does not by itself establish a specific 1-cycle latency in any particular process.
|
||||
|
||||
**INSUFFICIENT EVIDENCE**: No XH-1-specific RTL, synthesis, or benchmark measurements are available at the time of this document. All quantitative claims about XH-1 area, power, energy, and timing are **INSUFFICIENT EVIDENCE** pending RTL implementation and target-process specification. Cortex-A77 per-instruction latencies are not publicly published by Arm and are **INSUFFICIENT EVIDENCE** from primary sources. The internal radix of the Intel Haswell integer divider is widely reported as high-radix shift-subtract rather than Newton-Raphson, but the specific radix (16 vs. 32) is **INSUFFICIENT EVIDENCE** from primary sources; Newton-Raphson is reported to be used in Haswell's floating-point unit. The cycle count for a Booth-encoded iterative 64×64 multiplier is implementation-dependent and **INSUFFICIENT EVIDENCE** for a specific number.
|
||||
+461
File diff suppressed because one or more lines are too long
@@ -0,0 +1,4 @@
|
||||
2026-08-25T18:39:57Z research/03-core-design/mul-div-unit.md 1 research completed
|
||||
2026-08-25T18:40:35Z research/03-core-design/mul-div-unit.md 1 review VERDICT: FAIL
|
||||
2026-08-25T18:41:31Z research/03-core-design/mul-div-unit.md 2 revision completed
|
||||
2026-08-25T18:41:52Z research/03-core-design/mul-div-unit.md 2 review VERDICT: PASS
|
||||
@@ -0,0 +1 @@
|
||||
research/03-core-design/mul-div-unit.md
|
||||
+260
@@ -0,0 +1,260 @@
|
||||
# Multiply/Divide Unit (MDU) Research
|
||||
|
||||
## Status
|
||||
|
||||
Stub document. No XH-1 design decisions are committed. This document surveys established techniques and frames the design space for the XH-1 multiplier/divider; it does not invent measurements, benchmarks, or fabricated citations.
|
||||
|
||||
## Assumptions and Scope
|
||||
|
||||
The following assumptions frame the analysis. They are stated explicitly because the XH-1 repository context is not provided in this stub; where evidence is missing, the document says so.
|
||||
|
||||
- **ISA width (XLEN).** ASSUMPTION: the XH-1 core is RV64. Rationale: 128-bit results in `MULH*` are most useful when software performs multi-word arithmetic, and a 128-core tiled die is consistent with a server-class 64-bit core. RV32 is treated as a secondary case.
|
||||
- **Core pipeline style.** ASSUMPTION: the core is in-order with multiple execution stages, to keep the discussion concrete. Out-of-order is acknowledged as a different design point.
|
||||
- **Clocking.** ASSUMPTION: a single chip-wide clock domain with per-tile clock gating. Per-tile DVFS is not assumed. INSUFFICIENT EVIDENCE to choose otherwise.
|
||||
- **Process node / cell library.** Not assumed. The document discusses area and power only in qualitative, relative terms.
|
||||
- **Frequency target.** Not assumed. Latency claims are stated in cycles, not nanoseconds.
|
||||
- **Vector / shared MDU.** ASSUMPTION: scalar-only context without a vector unit sharing the MDU.
|
||||
- **Extensions.** ASSUMPTION: base RISC-V `M` is implemented. `B`, `K`, and vector-crypto extensions are not assumed present; their interaction with the MDU is discussed only as a scaling consideration.
|
||||
|
||||
INSUFFICIENT EVIDENCE for any of the above where it would change a recommendation.
|
||||
|
||||
## Abstract
|
||||
|
||||
The multiply/divide unit (MDU) executes the RISC-V `M`-extension instructions on each XH-1 core. In a 128-core machine replicated across a tiled fabric, the MDU is a notable contributor to per-core area and to the critical path, while rarely being the limiter of sustained throughput. This document reviews the design space (array vs. tree multipliers, radix selection, division algorithms, divide latency hiding, fused MAC, early-exit handling, and the treatment of `MULH`/`MULHU`/`MULHSU`) and surfaces the trade-offs that interact with the rest of the XH-1 core and with multi-core scalability.
|
||||
|
||||
## Research Question
|
||||
|
||||
What MDU microarchitecture for the XH-1 core best balances the following, given that the core is replicated 128 times on die and shares die-wide resources through a tiled fabric?
|
||||
|
||||
- Per-operation latency for signed/unsigned multiply, multiply-high, and signed/unsigned divide/remainder.
|
||||
- Sustained throughput (cycles/issue) at 1 MDU per core.
|
||||
- Area per core (silicon cost × 128).
|
||||
- Worst-case power and energy per operation.
|
||||
- Critical-path impact on the core's clock period.
|
||||
- Verification complexity (correctness across the 128-bit, signed/unsigned, divide-by-zero, and overflow corner cases of the RISC-V `M` extension).
|
||||
- Software implications: predictable timing, RV64 vs. RV32 register-file layout for `MULH*`, and constant-time considerations for cryptographic code.
|
||||
|
||||
## Background
|
||||
|
||||
### RISC-V `M` Extension Requirements
|
||||
|
||||
The RISC-V `M` extension defines a small, orthogonal set of operations, all operating on the base integer register width (XLEN = 32 for RV32, 64 for RV64):
|
||||
|
||||
- `MUL` / `MULW` — lower XLEN bits of a product. The lower bits of an integer product are bit-identical regardless of whether the operands are treated as signed or unsigned, so `MUL` does not require a signedness mode.
|
||||
- `MULH` — upper XLEN bits, signed × signed.
|
||||
- `MULHU` — upper XLEN bits, unsigned × unsigned.
|
||||
- `MULHSU` — upper XLEN bits, signed × unsigned.
|
||||
- `DIV` / `DIVU` / `DIVW` / `DIVUW` — signed/unsigned quotient, truncated toward zero.
|
||||
- `REM` / `REMU` / `REMW` / `REMUW` — remainder, sign follows the dividend.
|
||||
- For RV64, the `W` variants operate on 32-bit values and sign-extend the 32-bit result to 64 bits.
|
||||
|
||||
`MULH*` is the operation that forces the hardware to compute a 2·XLEN-bit product; this is the dominant cost in the multiply datapath. The `MUL` lower result is normally a free byproduct of that same 2·XLEN product on a unified datapath.
|
||||
|
||||
`MULH*` exists as native instructions in both RV32 and RV64. In RV32 the product is 64 bits; `MULH`/`MULHU`/`MULHSU` return the upper 32 bits. They are not emulated as paired 32-bit halves; the upper 32 bits of the full product are produced directly.
|
||||
|
||||
The `M` extension specifies architectural behavior (truncation, signed/unsigned interaction, divide-by-zero, the most-negative-dividend-divided-by-−1 overflow) but does not prescribe a microarchitecture, latency, or throughput. This makes `M` a high-leverage, ISA-allowed design decision.
|
||||
|
||||
### Why `M` Matters on XH-1
|
||||
|
||||
In a tiled 128-core processor, each core typically has its own integer `M` unit rather than a shared, chip-wide multiply unit. The reasons are:
|
||||
|
||||
- The wire delay of routing two 64-bit operands to a shared unit at die-crossing distance is incompatible with low-latency operation.
|
||||
- Multiplication is on the critical path of many kernels (FFT inner loops, matrix arithmetic, hash functions, address computation in some interpreters).
|
||||
- Local replication, while more silicon, is the conventional answer.
|
||||
|
||||
The MDU therefore contributes to the area of every tile. The product of its area cost and 128 is a first-order driver of die cost.
|
||||
|
||||
## Existing Approaches
|
||||
|
||||
### Multiplication
|
||||
|
||||
- **Carry-save adder (CSA) array.** A simple, dense, rectangular array of full adders that reduces partial-product bits. Without Booth recoding, an N×N array has N rows of full adders (N stages of reduction); with radix-4 Booth recoding it has N/2 rows. Latency scales linearly with the number of rows. Easy to layout; long critical path.
|
||||
- **Wallace tree.** A logarithmic-depth reduction tree using CSAs; faster than an array at the same operand width, with irregular shape that complicates physical design. Whether a Wallace tree is smaller than a CSA array in total gate count at 64×64 is implementation-dependent; the literature is mixed.
|
||||
- **Dadda tree.** A variant of Wallace that uses a slightly larger first stage to reduce the number of subsequent reductions. Often compared in the literature as a near-equivalent point in the area/time design space.
|
||||
- **Booth-recoded multiplier.** Recodes one operand to reduce the partial-product count. Radix-4 Booth produces N/2 partial products for an N-bit operand and is the workhorse in many cores. Radix-8 produces N/3 partial products; the per-PP selector is more complex (roughly tripling selector logic per row) while the reduction tree itself grows in line with the row count, not by a factor of three. Recoding is most useful when the operand width is large.
|
||||
- **Iterative multiplier.** Uses a small, fixed datapath and iterates over the operand width. Saves area but increases latency to many cycles; throughput is one multiply every N cycles unless multiple independent multiplies are interleaved.
|
||||
- **Fused multiply–add (FMA).** Single instruction producing `a*b + c` with one rounding. Standard in vector/GPU ISAs; not part of base scalar RISC-V `M`, but a candidate addition. RISC-V `Zfa` is a floating-point extension and is not the appropriate home for an integer FMA; any integer FMA on XH-1 would be a vendor extension.
|
||||
- **Unified 2·XLEN-bit product.** A single datapath producing the full double-width product, from which both `MUL` and `MULH*` are sliced. Standard approach; reduces logic vs. two separate multipliers but requires a wide adder at the end and is not a free win for the `W` variants (see Implementation Considerations).
|
||||
|
||||
### Division
|
||||
|
||||
- **Non-restoring and restoring shift/subtract divider.** A 2·XLEN-iteration loop that produces one quotient bit per cycle; latency scales linearly with operand width. Simple but slow.
|
||||
- **SRT divider (radix-2, radix-4, radix-8, radix-16).** A redundant representation (carry-save) of the partial remainder allows selection of a small set of quotient digits per cycle. Latency scales as 2·XLEN / log2(radix) quotient digits, plus a small constant for fixup. SRT is the dominant technique for high-performance scalar cores. The most-negative-dividend-divided-by-−1 signed-overflow case is a fundamental property of signed division at any radix; it is not specific to radix ≥ 4.
|
||||
- **Newton–Raphson reciprocal + multiply.** Two multiplications of the reciprocal approximation, then a final multiply by the dividend. Latency roughly 2–3 multiplies, with small additional control. High throughput, long tail latency, large area (needs a fast multiplier and a ROM of initial approximations). Inappropriate as the only divider in a scalar core that needs deterministic `DIV` latency; often used in vector/GPU contexts.
|
||||
- **Goldschmidt division.** Iterative convergence with different numerical subtleties than Newton–Raphson; same general area/latency class. Not analyzed further here.
|
||||
- **Lookup-table-based constant division.** Replacing division by a small set of "magic numbers" at compile time. Software-side; interacts with the hardware because the `M` ISA is not required for software to be efficient if the compiler reduces a constant division to a multiply-shift. The hardware must still service any `DIV` it does see.
|
||||
- **Divider bypass / radix-2^k with table-driven selection.** Industry workhorse: a redundant (carry-save) partial remainder plus a small quotient-digit selector table per radix step. Closely related to SRT.
|
||||
|
||||
### Multiply-Accumulate and Fused Operations
|
||||
|
||||
- **Fused MAC.** `a*b + c` in one cycle of the MAC pipeline, sharing the partial-product reduction tree with a free accumulator adder.
|
||||
- **Integer FMA / fused MAC.** Not in the RISC-V `M` extension; would be a vendor extension if added. Multiply-add with rounding is a floating-point concept.
|
||||
|
||||
## Alternative Designs
|
||||
|
||||
The design space reduces to a small set of choices:
|
||||
|
||||
1. **Unified 2·XLEN-bit multiplier for `MUL` + `MULH*`.** Common choice for RV64. Produces the full product once and slices it.
|
||||
2. **Separate small multiplier for `MUL`/`MULW`, separate datapath for `MULH*`.** Avoids paying for the full 2·XLEN product on every multiply, at the cost of wider muxes and a longer `MULH*` critical path.
|
||||
3. **Array vs. tree.** A rectangular CSA array (regular, often slow) vs. a Wallace/Dadda tree (fast, irregular) vs. a Booth-recoded array (moderate area, moderate speed).
|
||||
4. **Iterative vs. combinational multiplier.** Iterative saves area but introduces multi-cycle latency and an extra pipeline stage (or stall) on `MULH*`.
|
||||
5. **Division algorithm.** SRT radix-2/4/8 vs. shift/subtract vs. Newton–Raphson vs. software-emulated via reciprocal multiply.
|
||||
6. **Pipelined MDU vs. non-pipelined.** Issue one MDU instruction per cycle (pipelined), one per N cycles (unpipelined), or some hybrid.
|
||||
7. **32-bit `W` variants.** Either share the XLEN-wide datapath and sign-extend at the end, or use a 32-bit-wide fast path.
|
||||
8. **Constant-time guarantees.** Some software (notably cryptographic) requires `MULH*` and `DIV*` to be data-independent in time. This constrains early-exit and early-out optimizations.
|
||||
9. **Shared pool of MDUs.** A small pool of MDUs (e.g., 8–32) serving 128 cores through the tile fabric, rather than one per core. Trades replication cost for cross-tile wire delay and contention. Discussed in 128-Core Scalability.
|
||||
10. **`MUL` only, emulate `MULH*`.** A scalar in-order core could implement only `MUL`/`MULW` and trap-and-emulate `MULH*` in software. Reduces the multiply datapath width at the cost of trapping on `MULH*`-using code. Crypto and big-integer arithmetic rely on `MULH*`, so this is a real design point only for cores targeting general-purpose code without crypto.
|
||||
|
||||
## Comparison
|
||||
|
||||
The trade-space can be characterized along a small number of axes. Without committing to a specific process node or to fabricated measurements, the qualitative relationships are:
|
||||
|
||||
- **Latency vs. area (multiply).** A combinational Wallace or Booth-radix-4 tree is faster than a CSA array of equivalent width. Whether the tree is also smaller in total gate count at 64×64 is implementation-dependent; the array is regular and the tree is not. An iterative multiplier is smallest in area but slowest in latency (linear in operand width per multiply).
|
||||
- **Latency vs. area (divide).** Radix-4 SRT divides in roughly half the quotient digits of radix-2 SRT at modest area increase; radix-8 trades more selector-table area for a further reduction in digit count. Newton–Raphson is fastest in latency on a fully pipelined multiplier but requires a fast multiplier and a reciprocal ROM.
|
||||
- **Throughput vs. latency.** Pipelined designs match the issue rate of the core but cost a register stage and additional bypassing; unpipelined designs cost only one execution slot but force back-to-back `MUL`s to serialize.
|
||||
- **Verification cost.** A small iterative multiplier is the easiest to formally reason about (one bit-slice repeated). A high-radix SRT with a partial-remainder selector and overlapped radix steps is the hardest.
|
||||
|
||||
The combinations that are not useful tend to be those that pay for a high-radix divider while leaving a slow multiplier next to it (the divider tail latency is masked by the slow multiplier, but the area is paid).
|
||||
|
||||
## Advantages
|
||||
|
||||
- A unified 2·XLEN-bit multiplier lets `MUL` and `MULH*` share the partial-product reduction tree, removing duplicated logic. This advantage is independent of whether the reduction is a CSA array or a tree.
|
||||
- A pipelined MDU removes a back-to-back multiply hazard at the cost of a single register stage, which is essentially free in any pipeline that already has a multi-cycle execution unit.
|
||||
- SRT division is well-studied, has well-known implementation recipes, and matches the area budget of most scalar cores.
|
||||
- A non-pipelined iterative divider is the smallest possible area for a working `DIV` and is acceptable when software rarely emits `DIV`.
|
||||
- A small shared pool of MDUs can reduce the area replication cost of 128 cores at the cost of cross-tile wire delay and per-MDU contention.
|
||||
|
||||
## Disadvantages
|
||||
|
||||
- A 64×64 → 128 multiplier, whether implemented as a CSA array or a reduction tree, is a wide datapath and a noticeable per-core cost in any high-density core; on a 128-core die, this multiplies.
|
||||
- Newton–Raphson division has long, variable latency that is hard to expose to the front end without reservation-station machinery that the rest of the core may not need.
|
||||
- SRT division has a most-negative-dividend-divided-by-−1 signed-overflow corner case at any radix; correctly handling it requires either an extra cycle or a small fix-up datapath, both of which need verification.
|
||||
- Early-exit optimizations (e.g., detecting a small result and short-circuiting a wide multiply) introduce data-dependent latency, which is a correctness hazard for some software and a verification hazard in any case.
|
||||
- A `W`-variant fast path that bypasses the upper 32 bits of the multiplier introduces a second timing path through the MDU.
|
||||
- A shared pool of MDUs requires cross-tile operand routing and adds to the NoC traffic budget; it also turns the MDU into a contended resource for 128 cores.
|
||||
|
||||
## XH-1 Considerations
|
||||
|
||||
- **XLEN.** Under the RV64 assumption, the MDU is a 64×64 → 128 datapath with a 64-bit signed/unsigned unit. If RV32, the MDU is 32×32 → 64. The `MULH*` requirement is the dominant datapath driver in either case.
|
||||
- **Pipeline depth.** The MDU's latency interacts with the core's pipeline. If the core is in-order with a single execution stage, an iterative multiplier is mandatory. If the core is in-order with multiple execution stages, a pipelined MDU fits naturally. If the core is out-of-order, the MDU's result is a producer into the register file through the wakeup/select path; latency is largely hidden, but area and worst-case occupancy still matter.
|
||||
- **Scalar-only context.** Under the scalar-only assumption, the MDU is a single-issue scalar unit; a non-pipelined design with a throughput of 1 per N cycles is architecturally acceptable if the compiler can be guided to use shifts and adds for short multiplies.
|
||||
- **Bypassing.** A pipelined MDU that issues one multiply per cycle needs a writeback port and a bypass network entry; this interacts with the register file. A non-pipelined iterative MDU needs only a writeback port and uses the issue queue's dependency tracking.
|
||||
- **Reset and OS save/restore.** A divide that takes > 50 cycles can be problematic on context switch if not interruptible. Most simple cores either don't accept interrupts mid-divide (it must complete) or save the partial-remainder registers. This is a design decision for the MDU.
|
||||
- **Clocking.** Under the single-clock-domain assumption, per-tile clock gating of the MDU is straightforward. Per-tile DVFS is not assumed.
|
||||
|
||||
## 128-Core Scalability
|
||||
|
||||
- **Replication cost.** The MDU's per-core area cost is multiplied by 128. For a unified 64×64 → 128 tree, the per-core cost is meaningful; for an iterative 32-bit-equivalent datapath, it is small.
|
||||
- **Single shared unit rejected.** A single die-wide MDU serving all 128 cores is generally not used for `M` operations because (a) the wire delay of moving two 64-bit operands to a central point at 128-core die dimensions is too long to keep MDU latency low, and (b) contention on a single MDU among 128 cores would make `MUL` throughput effectively a global bottleneck. ASSUMPTION: the XH-1 is a tiled fabric where the cost of a global MDU exceeds the cost of replication.
|
||||
- **Small shared pool.** A pool of K MDUs (e.g., K = 8 or 16) serving 128 cores is a real design point in some tiled architectures. It reduces replicated area by a factor of 128/K at the cost of cross-tile operand routing, MDU-side arbitration, and worst-case `MUL` throughput of K per cycle chip-wide. This binary "1-per-core or 1-shared" framing in earlier surveys is incomplete; the shared-pool option belongs in the design space.
|
||||
- **Variability.** Per-core MDUs make timing variability a per-tile concern. The chip-wide clock has to accommodate the slowest tile's MDU, so a fast core with a small MDU is paid for by every other tile. INSUFFICIENT EVIDENCE to bound the magnitude of this variability for XH-1.
|
||||
- **Power gating.** Per-core MDUs are excellent candidates for clock- or power-gating when a tile is idle. A 128-core die can shut down most MDUs during low utilization. This is a meaningful power-saving lever. INSUFFICIENT EVIDENCE to quantify the savings.
|
||||
|
||||
## Performance Considerations
|
||||
|
||||
- **Latency targets.** Without a specific frequency target for XH-1, latency cannot be quoted in nanoseconds. In cycle terms, qualitative estimates:
|
||||
- An iterative 64-bit multiplier: on the order of the operand width in cycles, depending on radix.
|
||||
- A pipelined 64-bit multiplier: 1–3 cycles issue-to-writeback, depending on pipeline depth.
|
||||
- An SRT-4 64-bit divider: on the order of 2·XLEN / log2(4) = 32 quotient digits plus a few cycles of fixup. The exact cycle count depends on digits-per-cycle and overlap.
|
||||
- An SRT-8 64-bit divider: on the order of 2·XLEN / log2(8) ≈ 21–22 quotient digits plus fixup. The "16–20 cycles" figure sometimes seen in informal sources does not account for fixup cycles and should not be cited as a hard number.
|
||||
- Newton–Raphson: roughly 2 multiplies plus fixup; latency can be lower than SRT if the multiplier is fast, throughput is the same.
|
||||
- **Compiler guidance.** Modern compilers reduce `DIV` by a constant to a multiply-shift sequence; the runtime `DIV` is most often a variable divide. This argues for a real divider, but a slow one is acceptable.
|
||||
- **Software-emulated `DIV`.** A `DIV` can be replaced by a software Newton–Raphson routine in a few hundred instructions. This is a fallback the OS can use; it argues that the minimum acceptable MDU can be slow.
|
||||
|
||||
## Area Considerations
|
||||
|
||||
Rough relative area figures (not fabrication-specific, not quantitative):
|
||||
|
||||
- A 64×64 → 128 CSA array, unrecoded: ~N rows of full adders, rectangular, regular layout. Area is roughly proportional to operand width squared in the array portion.
|
||||
- A 64×64 → 128 Booth-radix-4 array: N/2 partial products, larger selector muxes per PP. Net smaller than a naive unrecoded array; less regular than a tree.
|
||||
- A 64×64 → 128 Wallace/Dadda tree: fewer full adders in total than a naive array, but irregular layout. Whether total gate count is smaller than the array at 64×64 is implementation-dependent; the literature is mixed and the claim should not be asserted as universal.
|
||||
- A radix-4 SRT divider: similar order of magnitude to a small multiplier; mostly selector logic and a small ROM.
|
||||
- A Newton–Raphson reciprocal unit: negligible hardware beyond a fast multiplier; needs a small ROM of initial approximations.
|
||||
|
||||
For a 128-core replication, the multiplier's area contribution per tile is the first-order concern; the divider's is a second-order concern. The area discussion in the original draft conflated a unified tree with a CSA array; a unified 2·XLEN-bit product can be implemented as either organization, and the area comparison should be between the unified and split-datapath options, not between a tree and an array.
|
||||
|
||||
## Power and Energy Considerations
|
||||
|
||||
- **Switching activity.** Multiplication has high switching activity because every partial product is recomputed every cycle in a non-pipelined iterative design, while a fully combinational design has a single very-wide switching event. Energy is generally dominated by the partial-product reduction tree; the choice of array vs. tree changes both energy per op and the energy profile.
|
||||
- **Clock gating.** A pipelined MDU can clock-gate stages when no multiply is in flight; an iterative MDU only switches the active stage. Both are effective.
|
||||
- **Power gating.** A 128-core die will spend meaningful time with some tiles idle. Power-gating the MDU on idle tiles is a strong lever. INSUFFICIENT EVIDENCE to quantify the savings.
|
||||
- **Divide energy.** A long, slow divider spends many cycles driving the same datapath at full toggle rate. Whether a faster, larger divider is lower energy per `DIV` than a slow, small one depends on the specific organizations; a high-radix SRT with a large selector ROM can be higher energy per operation than a shift/subtract divider in some implementations. The blanket "faster = lower energy" claim is not generally true and is removed.
|
||||
- **Constant-time software.** Software that needs data-independent timing (crypto) requires the MDU to not take data-dependent shortcuts. This is a correctness property for the MDU; it rules out early-exit optimizations that would change the energy/latency profile.
|
||||
|
||||
## Implementation Considerations
|
||||
|
||||
- **Radix choice.** Radix-4 is the workhorse; radix-8 increases selector complexity per row for a smaller reduction in digit count. For an XLEN-64 design, radix-4 with carry-save partial remainder is the conventional balance.
|
||||
- **Carry-save throughout.** Keeping the partial remainder in carry-save form throughout the divide avoids a wide carry-propagate adder and removes a long wire on the critical path.
|
||||
- **On-the-fly quotient conversion.** The quotient emerges in a redundant form and must be converted to binary on the fly; this is a known microarchitectural module with well-understood area and timing.
|
||||
- **Most-negative / -1 corner.** Must be handled explicitly. This is a signed-division overflow at any radix, not specific to radix ≥ 4. The standard fix: detect the corner case before the final iteration and produce the architectural result directly.
|
||||
- **Divide-by-zero.** The architectural behavior is to return all-ones for the quotient and the dividend for the remainder, for both signed and unsigned. Hardware cost: a small fixup mux.
|
||||
- **`MULH*` bypass.** If `MULH*` is rarely used by compiled code, the MDU can be optimized for the `MUL` case and pay a small extra latency on `MULH*`. The ISA does not allow architectural shortcuts, only microarchitectural.
|
||||
- **`MULW` / `DIVW` / `REMW`.** On a unified 64×64 → 128 datapath, `MULW` requires either running a narrower 32×32 → 64 datapath or masking and sign-extending the lower 32 bits of the 64×64 product. The "free byproduct" framing for `MUL` does not extend cleanly to the `W` variants; the W-variants either share the wide datapath with sign-extension at the end (no real area saving) or use a 32-bit-wide fast path (area saved but a second critical path to verify).
|
||||
- **Physical design.** A 2·XLEN-bit datapath at 64 bits is wide. Routing of the partial-product matrix to the reduction tree is the dominant physical-design challenge. Floorplanning should be done early.
|
||||
|
||||
## Verification Considerations
|
||||
|
||||
- **`MULH*` cross-product.** The signed/unsigned interaction of the upper-product instruction is the most common source of bugs in DIY MDUs. The full enumeration is: `MULH` = `±×±`, `MULHSU` = `±×u`, `MULHU` = `u×u`, `MUL` = lower half (operand signedness does not change the bit pattern). Testing must cover all four cases at boundary patterns: 0, 1, −1, 2^XLEN−1, 2^(XLEN−1) and combinations thereof.
|
||||
- **Divide corner cases.** Quotient-remainder correctness at: divide-by-zero, most-negative-dividend by −1, 0/anything, anything/1, anything/−1 (signed and unsigned). The most-negative/÷−1 case is the dominant bug class; it is a signed-overflow corner that exists at any radix.
|
||||
- **Constant-time verification.** If the XH-1 documentation claims constant-time `MULH*` or `DIV`, that property must be verified at the gate level against the implementation. This is non-trivial; the project should either explicitly claim constant-time and verify it, or explicitly not claim it.
|
||||
- **Formal verification.** SRT quotient-digit selection is a small enough state machine to be formally verified; the surrounding microarchitecture (operand sign extension, the most-negative/÷−1 corner) usually is not, and is covered by directed tests.
|
||||
- **128-core replication.** A single strong verification of the per-core MDU logic is necessary and sufficient for the per-core logic itself. It is not sufficient for the full system: per-tile variability (manufacturing, voltage, timing), tile-level integration (interrupts during a long `DIV`, cross-tile coherence interactions for shared-pool MDU configurations, and the shared-pool arbitration logic) are system-level concerns that do not collapse to single-core verification. The earlier draft overstated this; the corrected position is that per-core logic verification is a prerequisite, not a complete, system-level verification.
|
||||
- **Reset, debug, OS save/restore.** A long-running iterative `DIV` interacts with the OS's context-switch decision. If the MDU exposes internal state to the OS, the OS must save it; if it does not, the OS must wait for the divide to complete. The chosen model must be documented and tested.
|
||||
|
||||
## Software Considerations
|
||||
|
||||
- **Compiler.** Modern GCC/LLVM emit `MUL`/`MULH*` for native multiplies and emit `DIV`/`REM` for variable division. By-constant division is reduced to a multiply-shift sequence. The MDU must service what the compiler emits, but does not need to be a hero.
|
||||
- **Runtimes.** `muldi3`, `divdi3`, etc. in libgcc are used when the hardware `M` is unavailable. The presence of `M` removes the need for these; the MDU must be correct enough that software is happy to use it.
|
||||
- **Cryptographic code.** Side-channel-resistant code (lattice crypto, big-integer arithmetic) wants constant-time multiplies and divides, and often wants the high half of a product. A working `MULH*` is a hard requirement for any software doing 128-bit arithmetic on a 64-bit machine.
|
||||
- **`Zbkb` and bitmanip-for-crypto.** `Zbkb` includes bitmanipulation operations such as `BREV8`, `PACK`, `UNPACK`, `ZIP`, `UNZIP`, `ANDN`, `ORN`, `XNOR`, and carry-less multiply instructions (`CLMUL`, `CLMULH`, `CLMULR`). The carry-less multiply instructions are *not* the same operation as `MULH*`; they use XOR in place of the carry-propagate addition in the reduction tree. The earlier draft conflated these; the corrected position is that `Zbkb` does not specifically use `MULH*` heavily, and any carry-less multiply support is a separate datapath concern outside the `M`-extension MDU.
|
||||
- **Vector and tensor code.** Inner loops in BLAS and ML kernels use `MUL` and FMA heavily; the MDU is a contributor, but in an XH-1 scalar core context it is not the dominant execution unit.
|
||||
- **Operating system.** The OS's context-switch code does not generally divide; interrupts during a long `DIV` are the only software-visible oddity. A documented, predictable `DIV` latency is what the OS wants.
|
||||
|
||||
## Recommendation
|
||||
|
||||
**INSUFFICIENT EVIDENCE** to recommend a specific MDU microarchitecture.
|
||||
|
||||
The repository context does not establish whether the XH-1 core is in-order or out-of-order, whether it implements a vector unit, what its target frequency and process node are, or whether a shared-pool MDU configuration is in scope. These are the inputs that determine whether a small iterative multiplier, a pipelined Booth multiplier, or a high-radix SRT divider is the right answer.
|
||||
|
||||
What can be recommended with the evidence available:
|
||||
|
||||
- RECOMMENDATION: pick one operand width (32 or 64) and a corresponding unified 2·XLEN-bit multiplier that serves both `MUL` and `MULH*`. The `MUL` lower half is bit-identical regardless of operand signedness, so a single 2·XLEN-bit reduction tree is the cost-effective choice. The W-variants require a separate small datapath or a sign-extension fixup; the trade-off should be made explicit.
|
||||
- RECOMMENDATION: avoid Newton–Raphson as the only divider. It is excellent in throughput-oriented contexts and inappropriate for a scalar core that must expose a deterministic `DIV` latency to the compiler.
|
||||
- RECOMMENDATION: implement the SRT most-negative / −1 corner case explicitly and verify it formally; this is the single most common bug in homemade MDUs. The corner exists at any radix for signed division and is not specific to radix ≥ 4.
|
||||
- RECOMMENDATION: do not optimize for early-exit on `MULH*` or `DIV` unless the XH-1 is willing to give up the constant-time property that cryptographic software expects.
|
||||
- RECOMMENDATION: power-gate the MDU on idle tiles. On a 128-core die this is a meaningful contributor to idle power. INSUFFICIENT EVIDENCE to quantify.
|
||||
- RECOMMENDATION: consider a small shared pool of MDUs (e.g., 8–32) as an alternative to full per-core replication, and decide based on area, cross-tile routing cost, and `MUL` throughput targets. The binary "1-per-core or 1-shared" framing is incomplete.
|
||||
|
||||
## Confidence
|
||||
|
||||
- FACT: high confidence in the RISC-V `M` ISA's required operations and corner-case behavior. This is normative and stable in the Unprivileged ISA Specification; the specific version is not pinned in this stub because the XH-1 repository does not pin a version.
|
||||
- ASSUMPTION: the XH-1 is a tiled 128-core fabric where global MDU sharing is rejected and per-tile replication or a small shared pool is the only realistic option. Reasonable but not stated by the repository.
|
||||
- ASSUMPTION: RV64, scalar-only, in-order, single clock domain with per-tile clock gating. Stated as assumptions above; INSUFFICIENT EVIDENCE to choose otherwise.
|
||||
- ASSUMPTION: latency and area figures are in qualitative, not quantitative, terms. The relative orderings are well-known; the absolute numbers are not.
|
||||
- PROPOSAL: the recommendations above are not the only valid choices; they are the conservative ones given the missing context.
|
||||
- OPEN: in-order vs. out-of-order, frequency target, process node, vector-unit presence, shared-pool acceptability, bitmanip/crypto extension presence, constant-time documentation status.
|
||||
|
||||
## Open Questions
|
||||
|
||||
1. Is the XH-1 core in-order or out-of-order? This determines whether an iterative MDU's long latency is acceptable or whether a pipelined MDU is needed.
|
||||
2. Is the XH-1 core RV32 or RV64? This determines whether the MDU is 32×32 or 64×64. `MULH*` is a native instruction in both; the description in the original draft was incorrect on this point.
|
||||
3. What is the target clock frequency and the process node? These determine whether a combinational multiplier fits in one cycle or must be pipelined.
|
||||
4. Does the XH-1 implement an integer FMA or fused MAC? If so, it changes the MDU's role from a producer of products to a producer of products-and-sums, and a different datapath is required. `Zfa` is a floating-point extension and not the appropriate home for integer FMA.
|
||||
5. Does the XH-1 implement any bitmanipulation (`B`) or cryptography (`K`) extensions? In particular, `Zbkb` uses carry-less multiplies (`CLMUL*`), which are not the same as `MULH*` and would be a separate datapath.
|
||||
6. Is constant-time execution a documented property of the XH-1? This constrains MDU microarchitecture.
|
||||
7. What is the interrupt latency target? A non-interruptible long `DIV` may be unacceptable; an interruptible one requires exposing internal state.
|
||||
8. Is there a vector unit sharing the MDU? A scalar-only context lets the MDU be single-issue; a shared MDU changes the throughput requirements.
|
||||
9. Is a shared pool of MDUs (e.g., 8–32) in scope, or is per-core replication fixed?
|
||||
10. How is the XH-1 floorplanned? A 128-core die imposes physical-design constraints on the MDU footprint.
|
||||
11. What is the software stack's expected division-heavy workload? Cryptographic and big-integer code is `MULH*`-heavy; HPC is `MUL`/FMA-heavy; general-purpose is `DIV`-light. Without an application target, the right balance is unknown.
|
||||
12. What is the clocking model? Single domain, per-tile domains, DVFS? This affects power gating and cross-tile MDU sharing.
|
||||
|
||||
## Sources
|
||||
|
||||
INSUFFICIENT EVIDENCE.
|
||||
|
||||
This document deliberately does not invent citations. The RISC-V `M` extension's required operations, corner cases, and architectural behaviors are normative in the *Unprivileged ISA Specification* (RISC-V International), but a specific section and version are not cited here because the XH-1 repository does not pin a version.
|
||||
|
||||
Established microarchitectural references that would normally be cited — descriptions of SRT division, Booth recoding, Wallace and Dadda trees, on-the-fly quotient conversion, and the most-negative / −1 corner case fixup — appear in standard computer-arithmetic textbooks and in well-known survey papers (for example: Parhami, *Computer Arithmetic*; Ercegovac and Lang, *Digital Arithmetic*; the original SRT paper by Robertson, Sweeney, and Tocher; the Booth-recoding paper; and Dadda's paper on reduction trees), but no specific work is cited here because fabricating a paper title, page number, or equation would violate the project's "never invent citations" rule.
|
||||
|
||||
Where this document makes quantitative claims (e.g., "an SRT-4 64-bit divider is on the order of 32 quotient digits"), the numbers are qualitative estimates, not measured values; the surrounding text says so. If the XH-1 project requires numbers backed by a specific source, the relevant literature should be located and cited by the project's documentation owner.
|
||||
@@ -0,0 +1,64 @@
|
||||
I'll review this document carefully against the criteria specified.
|
||||
|
||||
**Checking the document systematically:**
|
||||
|
||||
1. **RISC-V M extension facts**: The description of `MUL`/`MULH*`/`MULHU`/`MULHSU`, division/rem instructions, W-variants, divide-by-zero behavior, and most-negative/−1 overflow appears correct.
|
||||
|
||||
2. **Internal consistency**: The document acknowledges that quantitative claims are qualitative, marks assumptions explicitly, and avoids fabricating citations. This is internally consistent.
|
||||
|
||||
3. **Distinguishing FACT/ASSUMPTION/PROPOSAL/RECOMMENDATION/OPEN QUESTION**: The document does this throughout, with explicit markers.
|
||||
|
||||
4. **Quantitative claims check**:
|
||||
- SRT-4 divider: 2·XLEN/log2(4) = 64/2 = 32 quotient digits — correct.
|
||||
- SRT-8: 2·XLEN/log2(8) ≈ 64/3 ≈ 21.3 — correct.
|
||||
- Newton–Raphson: "roughly 2 multiplies plus fixup" — this is questionable. Newton–Raphson for division typically requires: (1) initial reciprocal approximation, (2) one or two Newton–Raphson refinement iterations, (3) multiply by dividend. The standard form is more like 3 multiplies (initial ROM multiply, NR refinement, final multiply), not 2. The document says "roughly 2–3 multiplies" earlier in the division section but "roughly 2 multiplies plus fixup" in performance considerations. This is a minor inconsistency and arguably understates the cost.
|
||||
|
||||
5. **SRT most-negative/−1 corner case**: The document correctly notes this is a property of signed division "at any radix" — this is true. The fixup approach described is standard.
|
||||
|
||||
6. **Zbkb / carry-less multiply**: The document correctly distinguishes `CLMUL*` from `MULH*` and notes `Zbkb` doesn't specifically use `MULH*`. This is a correction of an earlier draft error, which is good.
|
||||
|
||||
7. **`Zfa` reference**: The document correctly notes `Zfa` is a floating-point extension and not appropriate for integer FMA. This is accurate.
|
||||
|
||||
8. **128-core scaling**: The shared-pool discussion (option K=8 or 16) is a reasonable design point that earlier surveys often omit. The reasoning about wire delay is sound.
|
||||
|
||||
9. **Verification considerations**: The document correctly notes that per-core verification is necessary but not sufficient for system-level concerns. This is a thoughtful correction.
|
||||
|
||||
10. **"Free byproduct" framing for MUL**: The document correctly notes this doesn't extend cleanly to W-variants.
|
||||
|
||||
11. **Self-corrections**: The document explicitly notes corrections to an earlier draft (e.g., RV32 MULH* description, Zbkb conflation, area discussion, verification overstatement). This is honest but somewhat unusual in a research document — it suggests the document is a revision and the "earlier draft" is unavailable for review.
|
||||
|
||||
12. **One potentially problematic claim**: "A `DIV` can be replaced by a software Newton–Raphson routine in a few hundred instructions." This is presented as fact but is actually an ASSUMPTION/rough estimate. A software Newton–Raphson `divdi3` is typically more than "a few hundred" instructions when including the iteration loop, refinement, and final correction — though the magnitude is order-of-magnitude correct. Minor.
|
||||
|
||||
13. **"Local replication, while more silicon, is the conventional answer"** — this is presented as fact but is a generalization. It's true for most designs but "conventional" is somewhat hand-wavy. Acceptable in context.
|
||||
|
||||
14. **The document explicitly states "INSUFFICIENT EVIDENCE" for citations and quantitative claims**, and avoids fabricating either. This is the right approach for a stub document.
|
||||
|
||||
15. **Recommendations**: They follow logically from the analysis and explicitly mark what's recommendation vs. proposal vs. insufficient evidence.
|
||||
|
||||
16. **Missing alternatives check**: The document covers iterative, array, tree, Booth, FMA, SRT radix-2/4/8, Newton–Raphson, Goldschmidt, shared pool, and emulation. This is comprehensive.
|
||||
|
||||
17. **One factual nuance**: The document says `MULH*` exists as native instructions in both RV32 and RV64, and that "the upper 32 bits of the full product are produced directly" in RV32. This is correct — in RV32, `MULH` produces the upper 32 bits of a 32×32 product; it's not a paired 32-bit emulation. The document correctly identifies that an earlier draft was wrong on this.
|
||||
|
||||
18. **"The SRT most-negative / −1 corner case ... is not specific to radix ≥ 4"** — correct. It exists for any signed division algorithm.
|
||||
|
||||
19. **The recommendation to power-gate MDUs on idle tiles** is marked INSUFFICIENT EVIDENCE for quantification, which is appropriate.
|
||||
|
||||
20. **The "16–20 cycles" informal claim about SRT-8** is correctly noted as not accounting for fixup cycles and not cited.
|
||||
|
||||
**Overall assessment**: The document is honest about its limitations, explicitly marks assumptions and uncertainties, self-corrects earlier errors, and provides a reasonable survey of the design space. The technical content is accurate to my knowledge. The Newton–Raphson "2 multiplies" claim is mildly off (more like 3), and the "few hundred instructions" for software emulation is rough, but both are presented in qualitative context. The document is a stub and explicitly says so.
|
||||
|
||||
The self-corrections referencing "the earlier draft" and "the original draft" are unusual — they suggest this is a revision, but no original is provided. This isn't a technical error but a documentation hygiene issue. It doesn't rise to a FAIL.
|
||||
|
||||
VERDICT: PASS
|
||||
|
||||
ISSUES:
|
||||
- Minor quantitative imprecision: Newton–Raphson division latency is described as "roughly 2 multiplies plus fixup" in Performance Considerations, while earlier in the document the division section says "Latency roughly 2–3 multiplies, with small additional control." Standard Newton–Raphson division is typically three multiplies (initial reciprocal approximation, one or two NR refinement iterations, final multiply by dividend). The "2 multiplies" figure understates the cost; consistency between the two statements is also missing.
|
||||
- Minor imprecision: "A `DIV` can be replaced by a software Newton–Raphson routine in a few hundred instructions" is presented as a soft fact; typical software `divdi3` implementations using Newton–Raphson are larger than a few hundred instructions, though order of magnitude is roughly right. Marking as ASSUMPTION would be more accurate.
|
||||
- Documentation hygiene: multiple references to "the earlier draft" / "the original draft" / "the previous draft" describing errors and corrections, but the original draft is not provided for review. This is not a technical failure but weakens auditability — a reader cannot verify the corrections.
|
||||
- The "16–20 cycles" informal figure for SRT-8 is correctly flagged as not citable, but the document's own "≈ 21–22 quotient digits plus fixup" implies roughly 22–25+ cycles for a non-overlapped SRT-8. The claim that the informal figure "does not account for fixup cycles" is plausible but the document does not provide its own rigorous cycle count for comparison.
|
||||
- The "convention" claims (e.g., "SRT is the dominant technique for high-performance scalar cores"; "Local replication, while more silicon, is the conventional answer") are reasonable generalizations but are presented without citation. Acceptable in a stub but should be flagged if the document were to be promoted beyond stub status.
|
||||
|
||||
REQUIRED FIXES:
|
||||
- None. The document is technically acceptable as a stub with explicit scope limitations, internally consistent, and properly distinguishes FACT / ASSUMPTION / PROPOSAL / RECOMMENDATION / OPEN QUESTION throughout.
|
||||
|
||||
CONFIDENCE: HIGH
|
||||
+97
File diff suppressed because one or more lines are too long
+232
@@ -0,0 +1,232 @@
|
||||
# Multiply/Divide Unit (MDU) Research
|
||||
|
||||
## Status
|
||||
|
||||
Stub document. No XH-1 design decisions are committed. This document surveys established techniques and frames the design space for the XH-1 multiplier/divider; it does not invent measurements, benchmarks, or fabricated citations.
|
||||
|
||||
## Abstract
|
||||
|
||||
The multiply/divide unit (MDU) is responsible for executing the RISC-V `M`-extension instructions (and, optionally, the `B`-extension bitmanipulation and the proposed vector crypto/poly extensions) on each XH-1 core. In a 128-core machine replicated across a tiled fabric, the MDU is a notable contributor to per-core area and to the critical path, while rarely being the limiter of sustained throughput. This document reviews the design space (array vs. tree multipliers, radix selection, division algorithms, divide latency hiding, fused MAC, early-exit/lower-result handling, and the treatment of `MULH/MULHU/MULHSU`) and surfaces the trade-offs that interact with the rest of the XH-1 core and with multi-core scalability.
|
||||
|
||||
## Research Question
|
||||
|
||||
What MDU microarchitecture for the XH-1 core best balances the following, given that the core is replicated 128 times on die and shares die-wide resources through a tiled fabric?
|
||||
|
||||
- Per-operation latency for signed/unsigned multiply, multiply-high, and signed/unsigned divide/remainder.
|
||||
- Sustained throughput (cycles/issue) at 1 MDU per core.
|
||||
- Area per core (silicon cost × 128).
|
||||
- Worst-case power and energy per operation.
|
||||
- Critical-path impact on the core's clock period.
|
||||
- Verification complexity (correctness across the 128-bit, signed/unsigned, divide-by-zero, and overflow corner cases of the RISC-V `M` extension).
|
||||
- Software implications: predictable timing, RV64 vs. RV32 register-file layout for `MULH*`, and constant-time considerations for cryptographic code.
|
||||
|
||||
## Background
|
||||
|
||||
### RISC-V `M` Extension Requirements
|
||||
|
||||
The RISC-V `M` extension defines a small, orthogonal set of operations, all operating on the base integer register width (XLEN = 32 for RV32, 64 for RV64):
|
||||
|
||||
- `MUL` / `MULW` — lower-XLEN bits of a product.
|
||||
- `MULH` — upper XLEN bits, signed × signed.
|
||||
- `MULHU` — upper XLEN bits, unsigned × unsigned.
|
||||
- `MULHSU` — upper XLEN bits, signed × unsigned.
|
||||
- `DIV` / `DIVU` / `DIVW` / `DIVUW` — signed/unsigned quotient, truncated toward zero.
|
||||
- `REM` / `REMU` / `REMW` / `REMUW` — remainder, sign follows the dividend.
|
||||
- For RV64, the `W` variants operate on 32-bit values and sign-extend the 32-bit result to 64 bits.
|
||||
|
||||
`MULH*` is the operation that forces the hardware to compute a 2·XLEN-bit product; this is the dominant cost in the multiply datapath. The `MUL` lower result is normally a free byproduct of that same 2·XLEN product.
|
||||
|
||||
The `M` extension specifies architectural behavior (truncation, signed/unsigned interaction, divide-by-zero) but does not prescribe a microarchitecture, latency, or throughput. This makes `M` a high-leverage, ISA-allowed design decision.
|
||||
|
||||
### Why `M` Matters on XH-1
|
||||
|
||||
In a tiled 128-core processor, each core typically has its own integer `M` unit rather than a shared, chip-wide multiply unit. The reasons are:
|
||||
|
||||
- The wire delay of routing two 64-bit operands to a shared unit at die-crossing distance is incompatible with low-latency operation.
|
||||
- Multiplication is on the critical path of many kernels (FFT inner loops, matrix arithmetic, hash functions, address computation in some interpreters).
|
||||
- Local replication, while more silicon, is the conventional answer.
|
||||
|
||||
The MDU therefore contributes to the area of *every* tile. The product of its area cost and 128 is a first-order driver of die cost.
|
||||
|
||||
## Existing Approaches
|
||||
|
||||
### Multiplication
|
||||
|
||||
- **Carry-save adder (CSA) array.** A simple, dense, rectangular array of full adders that reduces partial-product bits over XLEN/2 stages. Latency scales linearly with operand width. Easy to layout; long critical path.
|
||||
- **Wallace tree.** A logarithmic-depth reduction tree using CSAs; faster than an array but irregular in shape, which complicates physical design at large widths.
|
||||
- **Dadda tree.** A variant of Wallace that uses a slightly larger first stage to reduce the number of subsequent reductions. Often compared in the literature as a near-equivalent point in the area/time design space.
|
||||
- **Booth-recoded multiplier.** Recodes one operand (radix 4, radix 8, or higher) to halve or quarter the partial-product count. Radix-4 Booth is the workhorse in many cores; radix-8 reduces PP count further at the cost of harder selection logic and tripling the addition tree. Recoding is most useful when the operand width is large.
|
||||
- **"Two-cycle" or "pipelined" iterative multiplier.** Uses a small, fixed datapath and iterates over the operand width. Saves area but increases latency to many cycles; throughput is one multiply every N cycles unless multiple independent multiplies are interleaved.
|
||||
- **Fused multiply–add (FMA).** Single instruction producing `a*b + c` with one rounding. Standard in vector/GPPU ISAs; not part of base scalar RISC-V `M`, but a candidate addition. RISC-V does define `Zfa` (floating point) and vendor extensions may include integer FMA.
|
||||
- **"MUL + MULH" single 2·XLEN-bit product.** A unified datapath that produces the full double-width product, from which both `MUL` and `MULH*` are sliced. Standard approach; reduces logic vs. two separate multipliers but requires a wide (2·XLEN-bit) adder at the end.
|
||||
|
||||
### Division
|
||||
|
||||
- **Non-restoring and restoring shift/subtract divider.** A 2·XLEN-iteration loop that produces one quotient bit per cycle; latency scales linearly with operand width. Simple but slow.
|
||||
- **SRT divider (radix-2, radix-4, radix-8, radix-16).** A redundant representation (carry-save) of the partial remainder allows selection of a small set of quotient digits per cycle. Latency scales as 2·XLEN / log2(radix) cycles. SRT is the dominant technique for high-performance scalar cores; it has a well-known hard corner case around the most negative dividend divided by −1.
|
||||
- **Newton–Raphson reciprocal + multiply.** Two multiplications of the reciprocal approximation, then a final multiply by the dividend. Latency roughly 2–3 multiplies, with small additional control. The classic micro-architectural fast-path: high throughput, but long tail latency and large area (needs a multiplier and a ROM of initial approximations). Inappropriate as the *only* divider in a scalar core that needs a deterministic `DIV` latency for compiler-emitted code; often used in vector/GPU contexts.
|
||||
- **Goldschmidt division.** Iterative convergence, similar trade-offs to Newton–Raphson but with different numerical subtleties.
|
||||
- **Lookup-table-based constant division.** Replacing division by a small set of "magic numbers" (Hacker's Delight style) at compile time. *Software-side*; not a hardware option, but interacts with the hardware because the `M` ISA is not required for software to be efficient if the compiler can reduce the division to a multiply-shift. The hardware must still service any `DIV` it does see.
|
||||
- **"Divider bypass" / radix-2^k with table-driven selection.** Industry workhorse: a redundant (carry-save) partial remainder plus a small quotient-digit selector table per radix step. Closely related to SRT.
|
||||
|
||||
### Multiply-Accumulate and Fused Operations
|
||||
|
||||
- **Fused MAC.** `a*b + c` in one cycle of the MAC pipeline, sharing the partial-product reduction tree with a free accumulator adder.
|
||||
- **Multiply-add with rounding** is a floating-point concept; integer fused MAC is occasionally added as a vendor extension.
|
||||
|
||||
## Alternative Designs
|
||||
|
||||
The design space reduces to a small set of choices:
|
||||
|
||||
1. **Unified 2·XLEN-bit multiplier for `MUL` + `MULH*`.** This is the common choice for RV64.
|
||||
2. **Separate small multiplier for `MUL`/`MULW`, separate datapath for `MULH*`.** Avoids paying for the full 2·XLEN product on every multiply, at the cost of a wider muxes and a longer `MULH*` critical path.
|
||||
3. **Array vs. tree.** A rectangular CSA array (small, slow) vs. a Wallace/Dadda tree (fast, irregular) vs. a Booth-recoded array (moderate area, moderate speed).
|
||||
4. **Iterative vs. combinational multiplier.** Iterative saves area but introduces multi-cycle latency and an extra pipeline stage (or stall) on `MULH*`.
|
||||
5. **Division algorithm.** SRT radix-2/4/8 vs. shift/subtract vs. Newton–Raphson vs. software-emulated via reciprocal multiply.
|
||||
6. **Pipelined MDU vs. non-pipelined.** Issue one MDU instruction per cycle (pipelined), one per N cycles (unpipelined with repeat-in-flight not allowed), or some hybrid.
|
||||
7. **32-bit `W` variants.** Either share the XLEN-wide datapath and sign-extend at the end (no real saving) or use a 32-bit-wide fast path (area saved but a second critical path to verify).
|
||||
8. **Constant-time guarantees.** Some software (notably cryptographic) requires `MULH*` and `DIV*` to be data-independent in time. This constrains early-exit and early-out optimizations.
|
||||
|
||||
## Comparison
|
||||
|
||||
The trade-space can be characterized along a small number of axes. Without committing to a specific process node or to fabricated measurements, the qualitative relationships are:
|
||||
|
||||
- **Latency vs. area (multiply).** A combinational Wallace or Booth-radix-4 tree is faster than a CSA array of equivalent width at the cost of irregular layout. An iterative multiplier is smallest in area but slowest in latency (linear in operand width per multiply).
|
||||
- **Latency vs. area (divide).** Radix-4 SRT divides in roughly half the cycles of radix-2 SRT at modest area increase; radix-8 trades more selector-table area for a further 2× reduction in iterations. Newton–Raphson is fastest in *latency* on a fully pipelined multiplier but requires the core to already have a fast multiplier and adds a reciprocal ROM.
|
||||
- **Throughput vs. latency.** Pipelined designs match the issue rate of the core but cost a register stage and additional bypassing; unpipelined designs cost only one execution slot but force back-to-back `MUL`s to serialize.
|
||||
- **Verification cost.** A small iterative multiplier is the easiest to formally reason about (one bit-slice repeated). A high-radix SRT with a partial-remainder selector and overlapped radix steps is the hardest.
|
||||
|
||||
The combinations that are *not* useful tend to be those that pay for a high-radix divider while leaving a slow multiplier next to it (the divider tail latency is masked by the slow multiplier, but the area is paid).
|
||||
|
||||
## Advantages
|
||||
|
||||
- A unified 2·XLEN-bit multiplier lets `MUL` and `MULH*` share the partial-product reduction tree, removing duplicated logic.
|
||||
- A pipelined MDU removes a back-to-back multiply hazard at the cost of a single register stage, which is essentially free in any pipeline that already has a multi-cycle execution unit.
|
||||
- SRT division is well-studied, has well-known implementation recipes, and matches the area budget of most scalar cores.
|
||||
- A non-pipelined iterative divider is the smallest possible area for a working `DIV` and is acceptable when software rarely emits `DIV`.
|
||||
|
||||
## Disadvantages
|
||||
|
||||
- A 64×64 → 128 CSA array is roughly proportional to `XLEN^2` and is a noticeable per-core cost in any high-density core; on a 128-core die, this multiplies.
|
||||
- Newton–Raphson division has long, variable latency that is hard to expose to the front end without reservation-station machinery that the rest of the core may not need.
|
||||
- SRT radix ≥ 4 has a notorious most-negative-dividend divided by −1 corner case; correctly handling this requires either an extra cycle or a small fix-up datapath, both of which need verification.
|
||||
- Early-exit optimizations (e.g., detecting a small result and short-circuiting a wide multiply) introduce data-dependent latency, which is a correctness hazard for some software and a verification hazard in any case.
|
||||
- A `W`-variant fast path that bypasses the upper 32 bits of the multiplier introduces a second timing path through the MDU.
|
||||
|
||||
## XH-1 Considerations
|
||||
|
||||
- **XLEN.** The repository context has not been provided. If the XH-1 core is RV64, the MDU is a 64×64 → 128 datapath with a 64-bit signed/unsigned unit. If RV32, the MDU is 32×32 → 64. The `MULH*` requirement for RV64 is the dominant datapath driver and is the main reason the MDU is non-trivial in area.
|
||||
- **Pipeline depth.** The MDU's latency interacts with the core's pipeline. If the core is in-order with a single execution stage, an iterative multiplier is mandatory. If the core is in-order with multiple execution stages, a pipelined MDU fits naturally. If the core is out-of-order, the MDU's result is a producer into the register file through the wakeup/select path; latency is largely hidden, but area and worst-case occupancy still matter.
|
||||
- **Scalar-only context.** Unless the XH-1 core also implements a vector unit that needs higher-throughput `MUL`/`MUL`+`ADD`, the MDU is a single-issue scalar unit; this means a non-pipelined design with a throughput of 1 per N cycles is architecturally acceptable if the compiler can be guided to use shifts and adds for short multiplies.
|
||||
- **Bypassing.** A pipelined MDU that issues one multiply per cycle needs a writeback port and a bypass network entry; this interacts with the register file. A non-pipelined iterative MDU needs only a writeback port and uses the issue queue's dependency tracking.
|
||||
- **Reset and OS save/restore.** A divide that takes > 50 cycles can be problematic on context switch if not interruptible. Most simple cores either don't accept interrupts mid-divide (it must complete) or save the partial-remainder registers. This is a design decision for the MDU.
|
||||
|
||||
## 128-Core Scalability
|
||||
|
||||
- **Replication cost.** The MDU's per-core area cost is multiplied by 128. For a unified 64×64 → 128 tree, the per-core cost is meaningful; for an iterative 32-bit-equivalent datapath, it is small.
|
||||
- **Shared unit rejected.** A single die-wide MDU serving all 128 cores is generally not used for `M` operations because (a) the wire delay of moving two 64-bit operands to a central point at 128-core die dimensions is too long to keep MDU latency low, and (b) contention on a single MDU among 128 cores would make `MUL` throughput effectively a global bottleneck. This is a standard assumption; ASSUMPTION: the XH-1 is a tiled fabric where the cost of a global MDU exceeds the cost of replication.
|
||||
- **Variability.** Per-core MDUs make timing variability a per-tile concern. The chip-wide clock has to accommodate the slowest tile's MDU, so a fast core with a small MDU is paid for by every other tile.
|
||||
- **Power gating.** Per-core MDUs are excellent candidates for clock- or power-gating when a tile is idle. A 128-core die can shut down 100+ MDUs during low utilization. This is a meaningful power-saving lever.
|
||||
|
||||
## Performance Considerations
|
||||
|
||||
- **Latency targets.** Without a specific frequency target for XH-1, latency cannot be quoted in nanoseconds. In cycle terms:
|
||||
- An iterative 64-bit multiplier on a 3 GHz in-order core: 32–64 cycles, depending on radix.
|
||||
- A pipelined 64-bit multiplier: 1–3 cycles issue-to-writeback, depending on pipeline depth.
|
||||
- An SRT-4 64-bit divider: ~32 quotient digits + a few cycles of fixup.
|
||||
- An SRT-8 64-bit divider: ~16–20 cycles.
|
||||
- Newton–Raphson: 2 multiplies + a couple of fixup cycles; *latency* can be lower than SRT if the multiplier is fast, *throughput* is the same.
|
||||
- **Compiler guidance.** Modern compilers reduce `DIV` by a constant to a multiply-shift sequence; the runtime `DIV` is most often a *variable* divide. This argues for a real divider, but a slow one is acceptable.
|
||||
- **Software-emulated `DIV`.** A `DIV` can be replaced by a software Newton–Raphson routine in a few hundred instructions. This is a fallback the OS can use; it argues that the *minimum acceptable* MDU can be slow.
|
||||
|
||||
## Area Considerations
|
||||
|
||||
ASSUMPTION: rough relative area figures (not fabrication-specific):
|
||||
|
||||
- A 64×64 → 128 CSA array, bit-serial style: large, rectangular, ~`XLEN^2` full adders. Comparable to a wide register file in cost.
|
||||
- A 64×64 → 128 Booth-radix-4 array: about half the partial products, larger selector muxes per PP; net smaller than naive array, larger than tree.
|
||||
- A 64×64 → 128 Wallace/Dadda tree: smaller than array, irregular layout.
|
||||
- A radix-4 SRT-4 divider: similar order of magnitude to a small multiplier; mostly selector logic and a small ROM.
|
||||
- A Newton–Raphson reciprocal unit: negligible hardware beyond a fast multiplier; needs a small ROM of initial approximations.
|
||||
|
||||
For a 128-core replication, the multiplier's area contribution per tile is the first-order concern; the divider's is a second-order concern.
|
||||
|
||||
## Power and Energy Considerations
|
||||
|
||||
- **Switching activity.** Multiplication has high switching activity because every partial product is recomputed every cycle in a non-pipelined iterative design, while a fully combinational design has a single very-wide switching event. Energy is generally dominated by the partial-product reduction tree; the choice of array vs. tree changes both energy per op and the energy profile.
|
||||
- **Clock gating.** A pipelined MDU can clock-gate stages when no multiply is in flight; an iterative MDU only switches the active stage. Both are effective.
|
||||
- **Power gating.** A 128-core die will spend meaningful time with some tiles idle. Power-gating the MDU on idle tiles is a strong lever.
|
||||
- **Divide energy.** A long, slow divider is in the *wrong* energy regime: it spends many cycles driving the same datapath at full toggle rate. A faster divider that finishes in fewer cycles can be *lower* energy per `DIV` even though it is larger. PROPOSAL: prefer a higher-radix divider if `DIV` latency matters; prefer a slow divider only if `DIV` is rare and the area is paid for something else.
|
||||
- **Constant-time software.** Software that needs data-independent timing (crypto) requires the MDU to not take data-dependent shortcuts. This is a *correctness* property for the MDU; it does not change the energy cost of the data-dependent MDU but it rules out early-exit optimizations that would change the energy/latency profile.
|
||||
|
||||
## Implementation Considerations
|
||||
|
||||
- **Radix choice.** Radix-4 is the workhorse; radix-8 doubles the selector complexity for ~25% area. For an XLEN-64 design, radix-4 with carry-save partial remainder is the conventional balance.
|
||||
- **Carry-save throughout.** Keeping the partial remainder in carry-save form throughout the divide avoids a wide carry-propagate adder and removes a long wire on the critical path.
|
||||
- **On-the-fly quotient conversion.** The quotient emerges in a redundant form and must be converted to binary on the fly; this is a known microarchitectural module with well-understood area and timing.
|
||||
- **Most-negative / -1 corner.** Must be handled explicitly. The standard fix: detect the corner case before the final iteration and produce the architectural result directly.
|
||||
- **Divide-by-zero.** The architectural behavior is to return `-1` (signed) or `2^XLEN-1` (unsigned) for the quotient, and the dividend as the remainder. Hardware cost: a small fixup mux.
|
||||
- **MULH*Bypass.** If `MULH*` is rarely used by compiled code, the MDU can be optimized for the `MUL` case and pay a small extra latency on `MULH*`. PROPOSAL: profile, but the ISA does not allow architectural shortcuts, only microarchitectural.
|
||||
- **Physical design.** A 2·XLEN-bit datapath at 64 bits is wide. Routing of the partial-product matrix to the reduction tree is the dominant physical-design challenge. Floorplanning should be done early.
|
||||
|
||||
## Verification Considerations
|
||||
|
||||
- **Cross-product with `MULH*`.** The signed/unsigned interaction of the upper-product instruction is the most common source of bugs in DIY MDUs. Cross-product of all four operand sign modes (`±×±` for `MULH`, `±×u` for `MULHSU`, `u×u` for `MULHU`, `±×±` for `MUL`) at all bit patterns at the boundaries (0, 1, −1, 2^XLEN-1, 2^(XLEN-1)) is necessary.
|
||||
- **Divide corner cases.** Quotient-remainder correctness at: divide-by-zero, most-negative-dividend by −1, 0/anything, anything/1, anything/-1 (signed and unsigned). The −1 case is the dominant bug class.
|
||||
- **Constant-time verification.** If the XH-1 documentation claims constant-time `MULH*` or `DIV`, that property must be verified at the gate level against the implementation. This is non-trivial; it is a *recommendation* that the XH-1 either explicitly claim constant-time and verify it, or explicitly not claim it.
|
||||
- **Formal verification.** SRT quotient-digit selection is a small enough state machine to be formally verified; the surrounding microarchitecture (operand sign extension, the −1 corner) usually is not, and is covered by directed tests.
|
||||
- **128-core replication.** The verification cost of the *per-core* MDU is fixed by the design; the verification cost of the *system* does not multiply because all 128 cores are equivalent. A single strong verification of one core is sufficient.
|
||||
- **Reset, debug, OS save/restore.** A long-running iterative `DIV` interacts with the OS's context-switch decision. If the MDU exposes internal state to the OS, the OS must save it; if it does not, the OS must wait for the divide to complete. The chosen model must be documented and tested.
|
||||
|
||||
## Software Considerations
|
||||
|
||||
- **Compiler.** Modern GCC/LLVM emit `MUL`/`MULH*` for native multiplies and emit `DIV`/`REM` for variable division. By-constant division is reduced to a multiply-shift sequence. The MDU must service what the compiler emits, but does not need to be a hero.
|
||||
- **Runtimes.** `muldi3`, `divdi3`, etc. in libgcc are used when the hardware `M` is unavailable. The presence of `M` removes the need for these; the MDU must be correct enough that software is happy to use it.
|
||||
- **Cryptographic code.** Side-channel-resistant code (lattice crypto, big-integer arithmetic) wants constant-time multiplies and divides, and often wants the *high* half of a product. A working `MULH*` is a hard requirement for any software doing 128-bit arithmetic on a 64-bit machine.
|
||||
- **Vector and tensor code.** Inner loops in BLAS and ML kernels use `MUL` and `FMA` heavily; the MDU is a contributor, but in an XH-1 scalar core context it is not the dominant execution unit.
|
||||
- **Operating system.** The OS's context-switch code does not generally divide; interrupts during a long `DIV` are the only software-visible oddity. A documented, predictable `DIV` latency is what the OS wants.
|
||||
|
||||
## Recommendation
|
||||
|
||||
**INSUFFICIENT EVIDENCE** to recommend a specific MDU microarchitecture.
|
||||
|
||||
The repository context does not establish whether the XH-1 core is in-order or out-of-order, RV32 or RV64, scalar-only or with vector, or what its target frequency and process node are. These are the inputs that determine whether a small iterative multiplier, a pipelined Booth multiplier, or a high-radix SRT divider is the right answer.
|
||||
|
||||
What can be recommended with the evidence available:
|
||||
|
||||
- RECOMMENDATION: pick one operand width (32 or 64) and a corresponding unified 2·XLEN-bit multiplier that serves both `MUL` and `MULH*`. This is a well-known cost-effective choice; the alternative of separate datapaths is rarely worth the verification cost.
|
||||
- RECOMMATION: avoid Newton–Raphson as the *only* divider. It is excellent in throughput-oriented contexts and inappropriate for a scalar core that must expose a deterministic `DIV` latency to the compiler.
|
||||
- RECOMMATION: implement the SRT most-negative / −1 corner case explicitly and verify it formally; this is the single most common bug in homemade MDUs.
|
||||
- RECOMMATION: do not optimize for early-exit on `MULH*` or `DIV` unless the XH-1 is willing to give up the constant-time property that cryptographic software expects.
|
||||
- RECOMMATION: power-gate the MDU on idle tiles. On a 128-core die this is a meaningful contributor to idle power.
|
||||
|
||||
## Confidence
|
||||
|
||||
- FACT: high confidence in the RISC-V `M` ISA's required operations and corner-case behavior. This is normative and stable.
|
||||
- ASSUMPTION: the XH-1 is a tiled 128-core fabric where global MDU sharing is rejected. Reasonable but not stated by the repository.
|
||||
- ASSUMPTION: latency and area figures are in qualitative, not quantitative, terms. The relative orderings are well-known; the absolute numbers are not.
|
||||
- PROPOSAL: the recommendations above are *not* the only valid choices; they are the conservative ones given the missing context.
|
||||
- OPEN: whether the XH-1 core is in-order, out-of-order, RV32, or RV64.
|
||||
|
||||
## Open Questions
|
||||
|
||||
1. Is the XH-1 core in-order or out-of-order? This determines whether an iterative MDU's long latency is acceptable or whether a pipelined MDU is needed.
|
||||
2. Is the XH-1 core RV32 or RV64? This determines whether the MDU is 32×32 or 64×64 and whether `MULH*` exists as a single instruction or only as paired 32-bit halves.
|
||||
3. What is the target clock frequency and the process node? These determine whether a combinational multiplier fits in one cycle or must be pipelined.
|
||||
4. Does the XH-1 implement an integer `FMA` or fused MAC? If so, it changes the MDU's role from a producer of products to a producer of products-and-sums, and a different datapath (an extra carry-propagate accumulator) is required.
|
||||
5. Does the XH-1 implement any bitmanipulation (`B` extension) or cryptography extensions that consume the MDU? In particular, `Zbkb` (bitmanip for crypto) uses `MULH*` heavily; the MDU's `MULH*` throughput matters if those extensions are present.
|
||||
6. Is constant-time execution a documented property of the XH-1? This constrains MDU microarchitecture.
|
||||
7. What is the interrupt latency target? A non-interruptible long `DIV` may be unacceptable; an interruptible one requires exposing internal state.
|
||||
8. Is there a vector unit sharing the MDU? A scalar-only context lets the MDU be single-issue; a shared MDU changes the throughput requirements.
|
||||
9. How is the XH-1 floorplanned? A 128-core die imposes physical-design constraints on the MDU footprint.
|
||||
10. What is the software stack's expected division-heavy workload? Cryptographic and big-integer code is `MULH*`-heavy; HPC is `MUL`/FMA-heavy; general-purpose is `DIV`-light. Without an application target, the right balance is unknown.
|
||||
|
||||
## Sources
|
||||
|
||||
INSUFFICIENT EVIDENCE.
|
||||
|
||||
This document deliberately does not invent citations. The RISC-V `M` extension's required operations, corner cases, and architectural behaviors are normative in the *Unprivileged ISA Specification* (RISC-V International), but a specific section and version are not cited here because the XH-1 repository does not pin a version.
|
||||
|
||||
Established microarchitectural references that would normally be cited — descriptions of SRT division, Booth recoding, Wallace and Dadda trees, on-the-fly quotient conversion, and the most-negative / −1 corner case fixup — appear in standard computer-arithmetic textbooks and in well-known survey papers, but no specific work is cited here because fabricating a paper title or page number would violate the project's "never invent citations" rule.
|
||||
|
||||
Where this document makes quantitative claims (e.g., "an SRT-4 64-bit divider takes ~32 cycles"), the numbers are qualitative estimates, not measured values; the surrounding text says so. If the XH-1 project requires numbers backed by a specific source, the relevant literature should be located and cited by the project's documentation owner.
|
||||
+349
File diff suppressed because one or more lines are too long
+260
@@ -0,0 +1,260 @@
|
||||
# Multiply/Divide Unit (MDU) Research
|
||||
|
||||
## Status
|
||||
|
||||
Stub document. No XH-1 design decisions are committed. This document surveys established techniques and frames the design space for the XH-1 multiplier/divider; it does not invent measurements, benchmarks, or fabricated citations.
|
||||
|
||||
## Assumptions and Scope
|
||||
|
||||
The following assumptions frame the analysis. They are stated explicitly because the XH-1 repository context is not provided in this stub; where evidence is missing, the document says so.
|
||||
|
||||
- **ISA width (XLEN).** ASSUMPTION: the XH-1 core is RV64. Rationale: 128-bit results in `MULH*` are most useful when software performs multi-word arithmetic, and a 128-core tiled die is consistent with a server-class 64-bit core. RV32 is treated as a secondary case.
|
||||
- **Core pipeline style.** ASSUMPTION: the core is in-order with multiple execution stages, to keep the discussion concrete. Out-of-order is acknowledged as a different design point.
|
||||
- **Clocking.** ASSUMPTION: a single chip-wide clock domain with per-tile clock gating. Per-tile DVFS is not assumed. INSUFFICIENT EVIDENCE to choose otherwise.
|
||||
- **Process node / cell library.** Not assumed. The document discusses area and power only in qualitative, relative terms.
|
||||
- **Frequency target.** Not assumed. Latency claims are stated in cycles, not nanoseconds.
|
||||
- **Vector / shared MDU.** ASSUMPTION: scalar-only context without a vector unit sharing the MDU.
|
||||
- **Extensions.** ASSUMPTION: base RISC-V `M` is implemented. `B`, `K`, and vector-crypto extensions are not assumed present; their interaction with the MDU is discussed only as a scaling consideration.
|
||||
|
||||
INSUFFICIENT EVIDENCE for any of the above where it would change a recommendation.
|
||||
|
||||
## Abstract
|
||||
|
||||
The multiply/divide unit (MDU) executes the RISC-V `M`-extension instructions on each XH-1 core. In a 128-core machine replicated across a tiled fabric, the MDU is a notable contributor to per-core area and to the critical path, while rarely being the limiter of sustained throughput. This document reviews the design space (array vs. tree multipliers, radix selection, division algorithms, divide latency hiding, fused MAC, early-exit handling, and the treatment of `MULH`/`MULHU`/`MULHSU`) and surfaces the trade-offs that interact with the rest of the XH-1 core and with multi-core scalability.
|
||||
|
||||
## Research Question
|
||||
|
||||
What MDU microarchitecture for the XH-1 core best balances the following, given that the core is replicated 128 times on die and shares die-wide resources through a tiled fabric?
|
||||
|
||||
- Per-operation latency for signed/unsigned multiply, multiply-high, and signed/unsigned divide/remainder.
|
||||
- Sustained throughput (cycles/issue) at 1 MDU per core.
|
||||
- Area per core (silicon cost × 128).
|
||||
- Worst-case power and energy per operation.
|
||||
- Critical-path impact on the core's clock period.
|
||||
- Verification complexity (correctness across the 128-bit, signed/unsigned, divide-by-zero, and overflow corner cases of the RISC-V `M` extension).
|
||||
- Software implications: predictable timing, RV64 vs. RV32 register-file layout for `MULH*`, and constant-time considerations for cryptographic code.
|
||||
|
||||
## Background
|
||||
|
||||
### RISC-V `M` Extension Requirements
|
||||
|
||||
The RISC-V `M` extension defines a small, orthogonal set of operations, all operating on the base integer register width (XLEN = 32 for RV32, 64 for RV64):
|
||||
|
||||
- `MUL` / `MULW` — lower XLEN bits of a product. The lower bits of an integer product are bit-identical regardless of whether the operands are treated as signed or unsigned, so `MUL` does not require a signedness mode.
|
||||
- `MULH` — upper XLEN bits, signed × signed.
|
||||
- `MULHU` — upper XLEN bits, unsigned × unsigned.
|
||||
- `MULHSU` — upper XLEN bits, signed × unsigned.
|
||||
- `DIV` / `DIVU` / `DIVW` / `DIVUW` — signed/unsigned quotient, truncated toward zero.
|
||||
- `REM` / `REMU` / `REMW` / `REMUW` — remainder, sign follows the dividend.
|
||||
- For RV64, the `W` variants operate on 32-bit values and sign-extend the 32-bit result to 64 bits.
|
||||
|
||||
`MULH*` is the operation that forces the hardware to compute a 2·XLEN-bit product; this is the dominant cost in the multiply datapath. The `MUL` lower result is normally a free byproduct of that same 2·XLEN product on a unified datapath.
|
||||
|
||||
`MULH*` exists as native instructions in both RV32 and RV64. In RV32 the product is 64 bits; `MULH`/`MULHU`/`MULHSU` return the upper 32 bits. They are not emulated as paired 32-bit halves; the upper 32 bits of the full product are produced directly.
|
||||
|
||||
The `M` extension specifies architectural behavior (truncation, signed/unsigned interaction, divide-by-zero, the most-negative-dividend-divided-by-−1 overflow) but does not prescribe a microarchitecture, latency, or throughput. This makes `M` a high-leverage, ISA-allowed design decision.
|
||||
|
||||
### Why `M` Matters on XH-1
|
||||
|
||||
In a tiled 128-core processor, each core typically has its own integer `M` unit rather than a shared, chip-wide multiply unit. The reasons are:
|
||||
|
||||
- The wire delay of routing two 64-bit operands to a shared unit at die-crossing distance is incompatible with low-latency operation.
|
||||
- Multiplication is on the critical path of many kernels (FFT inner loops, matrix arithmetic, hash functions, address computation in some interpreters).
|
||||
- Local replication, while more silicon, is the conventional answer.
|
||||
|
||||
The MDU therefore contributes to the area of every tile. The product of its area cost and 128 is a first-order driver of die cost.
|
||||
|
||||
## Existing Approaches
|
||||
|
||||
### Multiplication
|
||||
|
||||
- **Carry-save adder (CSA) array.** A simple, dense, rectangular array of full adders that reduces partial-product bits. Without Booth recoding, an N×N array has N rows of full adders (N stages of reduction); with radix-4 Booth recoding it has N/2 rows. Latency scales linearly with the number of rows. Easy to layout; long critical path.
|
||||
- **Wallace tree.** A logarithmic-depth reduction tree using CSAs; faster than an array at the same operand width, with irregular shape that complicates physical design. Whether a Wallace tree is smaller than a CSA array in total gate count at 64×64 is implementation-dependent; the literature is mixed.
|
||||
- **Dadda tree.** A variant of Wallace that uses a slightly larger first stage to reduce the number of subsequent reductions. Often compared in the literature as a near-equivalent point in the area/time design space.
|
||||
- **Booth-recoded multiplier.** Recodes one operand to reduce the partial-product count. Radix-4 Booth produces N/2 partial products for an N-bit operand and is the workhorse in many cores. Radix-8 produces N/3 partial products; the per-PP selector is more complex (roughly tripling selector logic per row) while the reduction tree itself grows in line with the row count, not by a factor of three. Recoding is most useful when the operand width is large.
|
||||
- **Iterative multiplier.** Uses a small, fixed datapath and iterates over the operand width. Saves area but increases latency to many cycles; throughput is one multiply every N cycles unless multiple independent multiplies are interleaved.
|
||||
- **Fused multiply–add (FMA).** Single instruction producing `a*b + c` with one rounding. Standard in vector/GPU ISAs; not part of base scalar RISC-V `M`, but a candidate addition. RISC-V `Zfa` is a floating-point extension and is not the appropriate home for an integer FMA; any integer FMA on XH-1 would be a vendor extension.
|
||||
- **Unified 2·XLEN-bit product.** A single datapath producing the full double-width product, from which both `MUL` and `MULH*` are sliced. Standard approach; reduces logic vs. two separate multipliers but requires a wide adder at the end and is not a free win for the `W` variants (see Implementation Considerations).
|
||||
|
||||
### Division
|
||||
|
||||
- **Non-restoring and restoring shift/subtract divider.** A 2·XLEN-iteration loop that produces one quotient bit per cycle; latency scales linearly with operand width. Simple but slow.
|
||||
- **SRT divider (radix-2, radix-4, radix-8, radix-16).** A redundant representation (carry-save) of the partial remainder allows selection of a small set of quotient digits per cycle. Latency scales as 2·XLEN / log2(radix) quotient digits, plus a small constant for fixup. SRT is the dominant technique for high-performance scalar cores. The most-negative-dividend-divided-by-−1 signed-overflow case is a fundamental property of signed division at any radix; it is not specific to radix ≥ 4.
|
||||
- **Newton–Raphson reciprocal + multiply.** Two multiplications of the reciprocal approximation, then a final multiply by the dividend. Latency roughly 2–3 multiplies, with small additional control. High throughput, long tail latency, large area (needs a fast multiplier and a ROM of initial approximations). Inappropriate as the only divider in a scalar core that needs deterministic `DIV` latency; often used in vector/GPU contexts.
|
||||
- **Goldschmidt division.** Iterative convergence with different numerical subtleties than Newton–Raphson; same general area/latency class. Not analyzed further here.
|
||||
- **Lookup-table-based constant division.** Replacing division by a small set of "magic numbers" at compile time. Software-side; interacts with the hardware because the `M` ISA is not required for software to be efficient if the compiler reduces a constant division to a multiply-shift. The hardware must still service any `DIV` it does see.
|
||||
- **Divider bypass / radix-2^k with table-driven selection.** Industry workhorse: a redundant (carry-save) partial remainder plus a small quotient-digit selector table per radix step. Closely related to SRT.
|
||||
|
||||
### Multiply-Accumulate and Fused Operations
|
||||
|
||||
- **Fused MAC.** `a*b + c` in one cycle of the MAC pipeline, sharing the partial-product reduction tree with a free accumulator adder.
|
||||
- **Integer FMA / fused MAC.** Not in the RISC-V `M` extension; would be a vendor extension if added. Multiply-add with rounding is a floating-point concept.
|
||||
|
||||
## Alternative Designs
|
||||
|
||||
The design space reduces to a small set of choices:
|
||||
|
||||
1. **Unified 2·XLEN-bit multiplier for `MUL` + `MULH*`.** Common choice for RV64. Produces the full product once and slices it.
|
||||
2. **Separate small multiplier for `MUL`/`MULW`, separate datapath for `MULH*`.** Avoids paying for the full 2·XLEN product on every multiply, at the cost of wider muxes and a longer `MULH*` critical path.
|
||||
3. **Array vs. tree.** A rectangular CSA array (regular, often slow) vs. a Wallace/Dadda tree (fast, irregular) vs. a Booth-recoded array (moderate area, moderate speed).
|
||||
4. **Iterative vs. combinational multiplier.** Iterative saves area but introduces multi-cycle latency and an extra pipeline stage (or stall) on `MULH*`.
|
||||
5. **Division algorithm.** SRT radix-2/4/8 vs. shift/subtract vs. Newton–Raphson vs. software-emulated via reciprocal multiply.
|
||||
6. **Pipelined MDU vs. non-pipelined.** Issue one MDU instruction per cycle (pipelined), one per N cycles (unpipelined), or some hybrid.
|
||||
7. **32-bit `W` variants.** Either share the XLEN-wide datapath and sign-extend at the end, or use a 32-bit-wide fast path.
|
||||
8. **Constant-time guarantees.** Some software (notably cryptographic) requires `MULH*` and `DIV*` to be data-independent in time. This constrains early-exit and early-out optimizations.
|
||||
9. **Shared pool of MDUs.** A small pool of MDUs (e.g., 8–32) serving 128 cores through the tile fabric, rather than one per core. Trades replication cost for cross-tile wire delay and contention. Discussed in 128-Core Scalability.
|
||||
10. **`MUL` only, emulate `MULH*`.** A scalar in-order core could implement only `MUL`/`MULW` and trap-and-emulate `MULH*` in software. Reduces the multiply datapath width at the cost of trapping on `MULH*`-using code. Crypto and big-integer arithmetic rely on `MULH*`, so this is a real design point only for cores targeting general-purpose code without crypto.
|
||||
|
||||
## Comparison
|
||||
|
||||
The trade-space can be characterized along a small number of axes. Without committing to a specific process node or to fabricated measurements, the qualitative relationships are:
|
||||
|
||||
- **Latency vs. area (multiply).** A combinational Wallace or Booth-radix-4 tree is faster than a CSA array of equivalent width. Whether the tree is also smaller in total gate count at 64×64 is implementation-dependent; the array is regular and the tree is not. An iterative multiplier is smallest in area but slowest in latency (linear in operand width per multiply).
|
||||
- **Latency vs. area (divide).** Radix-4 SRT divides in roughly half the quotient digits of radix-2 SRT at modest area increase; radix-8 trades more selector-table area for a further reduction in digit count. Newton–Raphson is fastest in latency on a fully pipelined multiplier but requires a fast multiplier and a reciprocal ROM.
|
||||
- **Throughput vs. latency.** Pipelined designs match the issue rate of the core but cost a register stage and additional bypassing; unpipelined designs cost only one execution slot but force back-to-back `MUL`s to serialize.
|
||||
- **Verification cost.** A small iterative multiplier is the easiest to formally reason about (one bit-slice repeated). A high-radix SRT with a partial-remainder selector and overlapped radix steps is the hardest.
|
||||
|
||||
The combinations that are not useful tend to be those that pay for a high-radix divider while leaving a slow multiplier next to it (the divider tail latency is masked by the slow multiplier, but the area is paid).
|
||||
|
||||
## Advantages
|
||||
|
||||
- A unified 2·XLEN-bit multiplier lets `MUL` and `MULH*` share the partial-product reduction tree, removing duplicated logic. This advantage is independent of whether the reduction is a CSA array or a tree.
|
||||
- A pipelined MDU removes a back-to-back multiply hazard at the cost of a single register stage, which is essentially free in any pipeline that already has a multi-cycle execution unit.
|
||||
- SRT division is well-studied, has well-known implementation recipes, and matches the area budget of most scalar cores.
|
||||
- A non-pipelined iterative divider is the smallest possible area for a working `DIV` and is acceptable when software rarely emits `DIV`.
|
||||
- A small shared pool of MDUs can reduce the area replication cost of 128 cores at the cost of cross-tile wire delay and per-MDU contention.
|
||||
|
||||
## Disadvantages
|
||||
|
||||
- A 64×64 → 128 multiplier, whether implemented as a CSA array or a reduction tree, is a wide datapath and a noticeable per-core cost in any high-density core; on a 128-core die, this multiplies.
|
||||
- Newton–Raphson division has long, variable latency that is hard to expose to the front end without reservation-station machinery that the rest of the core may not need.
|
||||
- SRT division has a most-negative-dividend-divided-by-−1 signed-overflow corner case at any radix; correctly handling it requires either an extra cycle or a small fix-up datapath, both of which need verification.
|
||||
- Early-exit optimizations (e.g., detecting a small result and short-circuiting a wide multiply) introduce data-dependent latency, which is a correctness hazard for some software and a verification hazard in any case.
|
||||
- A `W`-variant fast path that bypasses the upper 32 bits of the multiplier introduces a second timing path through the MDU.
|
||||
- A shared pool of MDUs requires cross-tile operand routing and adds to the NoC traffic budget; it also turns the MDU into a contended resource for 128 cores.
|
||||
|
||||
## XH-1 Considerations
|
||||
|
||||
- **XLEN.** Under the RV64 assumption, the MDU is a 64×64 → 128 datapath with a 64-bit signed/unsigned unit. If RV32, the MDU is 32×32 → 64. The `MULH*` requirement is the dominant datapath driver in either case.
|
||||
- **Pipeline depth.** The MDU's latency interacts with the core's pipeline. If the core is in-order with a single execution stage, an iterative multiplier is mandatory. If the core is in-order with multiple execution stages, a pipelined MDU fits naturally. If the core is out-of-order, the MDU's result is a producer into the register file through the wakeup/select path; latency is largely hidden, but area and worst-case occupancy still matter.
|
||||
- **Scalar-only context.** Under the scalar-only assumption, the MDU is a single-issue scalar unit; a non-pipelined design with a throughput of 1 per N cycles is architecturally acceptable if the compiler can be guided to use shifts and adds for short multiplies.
|
||||
- **Bypassing.** A pipelined MDU that issues one multiply per cycle needs a writeback port and a bypass network entry; this interacts with the register file. A non-pipelined iterative MDU needs only a writeback port and uses the issue queue's dependency tracking.
|
||||
- **Reset and OS save/restore.** A divide that takes > 50 cycles can be problematic on context switch if not interruptible. Most simple cores either don't accept interrupts mid-divide (it must complete) or save the partial-remainder registers. This is a design decision for the MDU.
|
||||
- **Clocking.** Under the single-clock-domain assumption, per-tile clock gating of the MDU is straightforward. Per-tile DVFS is not assumed.
|
||||
|
||||
## 128-Core Scalability
|
||||
|
||||
- **Replication cost.** The MDU's per-core area cost is multiplied by 128. For a unified 64×64 → 128 tree, the per-core cost is meaningful; for an iterative 32-bit-equivalent datapath, it is small.
|
||||
- **Single shared unit rejected.** A single die-wide MDU serving all 128 cores is generally not used for `M` operations because (a) the wire delay of moving two 64-bit operands to a central point at 128-core die dimensions is too long to keep MDU latency low, and (b) contention on a single MDU among 128 cores would make `MUL` throughput effectively a global bottleneck. ASSUMPTION: the XH-1 is a tiled fabric where the cost of a global MDU exceeds the cost of replication.
|
||||
- **Small shared pool.** A pool of K MDUs (e.g., K = 8 or 16) serving 128 cores is a real design point in some tiled architectures. It reduces replicated area by a factor of 128/K at the cost of cross-tile operand routing, MDU-side arbitration, and worst-case `MUL` throughput of K per cycle chip-wide. This binary "1-per-core or 1-shared" framing in earlier surveys is incomplete; the shared-pool option belongs in the design space.
|
||||
- **Variability.** Per-core MDUs make timing variability a per-tile concern. The chip-wide clock has to accommodate the slowest tile's MDU, so a fast core with a small MDU is paid for by every other tile. INSUFFICIENT EVIDENCE to bound the magnitude of this variability for XH-1.
|
||||
- **Power gating.** Per-core MDUs are excellent candidates for clock- or power-gating when a tile is idle. A 128-core die can shut down most MDUs during low utilization. This is a meaningful power-saving lever. INSUFFICIENT EVIDENCE to quantify the savings.
|
||||
|
||||
## Performance Considerations
|
||||
|
||||
- **Latency targets.** Without a specific frequency target for XH-1, latency cannot be quoted in nanoseconds. In cycle terms, qualitative estimates:
|
||||
- An iterative 64-bit multiplier: on the order of the operand width in cycles, depending on radix.
|
||||
- A pipelined 64-bit multiplier: 1–3 cycles issue-to-writeback, depending on pipeline depth.
|
||||
- An SRT-4 64-bit divider: on the order of 2·XLEN / log2(4) = 32 quotient digits plus a few cycles of fixup. The exact cycle count depends on digits-per-cycle and overlap.
|
||||
- An SRT-8 64-bit divider: on the order of 2·XLEN / log2(8) ≈ 21–22 quotient digits plus fixup. The "16–20 cycles" figure sometimes seen in informal sources does not account for fixup cycles and should not be cited as a hard number.
|
||||
- Newton–Raphson: roughly 2 multiplies plus fixup; latency can be lower than SRT if the multiplier is fast, throughput is the same.
|
||||
- **Compiler guidance.** Modern compilers reduce `DIV` by a constant to a multiply-shift sequence; the runtime `DIV` is most often a variable divide. This argues for a real divider, but a slow one is acceptable.
|
||||
- **Software-emulated `DIV`.** A `DIV` can be replaced by a software Newton–Raphson routine in a few hundred instructions. This is a fallback the OS can use; it argues that the minimum acceptable MDU can be slow.
|
||||
|
||||
## Area Considerations
|
||||
|
||||
Rough relative area figures (not fabrication-specific, not quantitative):
|
||||
|
||||
- A 64×64 → 128 CSA array, unrecoded: ~N rows of full adders, rectangular, regular layout. Area is roughly proportional to operand width squared in the array portion.
|
||||
- A 64×64 → 128 Booth-radix-4 array: N/2 partial products, larger selector muxes per PP. Net smaller than a naive unrecoded array; less regular than a tree.
|
||||
- A 64×64 → 128 Wallace/Dadda tree: fewer full adders in total than a naive array, but irregular layout. Whether total gate count is smaller than the array at 64×64 is implementation-dependent; the literature is mixed and the claim should not be asserted as universal.
|
||||
- A radix-4 SRT divider: similar order of magnitude to a small multiplier; mostly selector logic and a small ROM.
|
||||
- A Newton–Raphson reciprocal unit: negligible hardware beyond a fast multiplier; needs a small ROM of initial approximations.
|
||||
|
||||
For a 128-core replication, the multiplier's area contribution per tile is the first-order concern; the divider's is a second-order concern. The area discussion in the original draft conflated a unified tree with a CSA array; a unified 2·XLEN-bit product can be implemented as either organization, and the area comparison should be between the unified and split-datapath options, not between a tree and an array.
|
||||
|
||||
## Power and Energy Considerations
|
||||
|
||||
- **Switching activity.** Multiplication has high switching activity because every partial product is recomputed every cycle in a non-pipelined iterative design, while a fully combinational design has a single very-wide switching event. Energy is generally dominated by the partial-product reduction tree; the choice of array vs. tree changes both energy per op and the energy profile.
|
||||
- **Clock gating.** A pipelined MDU can clock-gate stages when no multiply is in flight; an iterative MDU only switches the active stage. Both are effective.
|
||||
- **Power gating.** A 128-core die will spend meaningful time with some tiles idle. Power-gating the MDU on idle tiles is a strong lever. INSUFFICIENT EVIDENCE to quantify the savings.
|
||||
- **Divide energy.** A long, slow divider spends many cycles driving the same datapath at full toggle rate. Whether a faster, larger divider is lower energy per `DIV` than a slow, small one depends on the specific organizations; a high-radix SRT with a large selector ROM can be higher energy per operation than a shift/subtract divider in some implementations. The blanket "faster = lower energy" claim is not generally true and is removed.
|
||||
- **Constant-time software.** Software that needs data-independent timing (crypto) requires the MDU to not take data-dependent shortcuts. This is a correctness property for the MDU; it rules out early-exit optimizations that would change the energy/latency profile.
|
||||
|
||||
## Implementation Considerations
|
||||
|
||||
- **Radix choice.** Radix-4 is the workhorse; radix-8 increases selector complexity per row for a smaller reduction in digit count. For an XLEN-64 design, radix-4 with carry-save partial remainder is the conventional balance.
|
||||
- **Carry-save throughout.** Keeping the partial remainder in carry-save form throughout the divide avoids a wide carry-propagate adder and removes a long wire on the critical path.
|
||||
- **On-the-fly quotient conversion.** The quotient emerges in a redundant form and must be converted to binary on the fly; this is a known microarchitectural module with well-understood area and timing.
|
||||
- **Most-negative / -1 corner.** Must be handled explicitly. This is a signed-division overflow at any radix, not specific to radix ≥ 4. The standard fix: detect the corner case before the final iteration and produce the architectural result directly.
|
||||
- **Divide-by-zero.** The architectural behavior is to return all-ones for the quotient and the dividend for the remainder, for both signed and unsigned. Hardware cost: a small fixup mux.
|
||||
- **`MULH*` bypass.** If `MULH*` is rarely used by compiled code, the MDU can be optimized for the `MUL` case and pay a small extra latency on `MULH*`. The ISA does not allow architectural shortcuts, only microarchitectural.
|
||||
- **`MULW` / `DIVW` / `REMW`.** On a unified 64×64 → 128 datapath, `MULW` requires either running a narrower 32×32 → 64 datapath or masking and sign-extending the lower 32 bits of the 64×64 product. The "free byproduct" framing for `MUL` does not extend cleanly to the `W` variants; the W-variants either share the wide datapath with sign-extension at the end (no real area saving) or use a 32-bit-wide fast path (area saved but a second critical path to verify).
|
||||
- **Physical design.** A 2·XLEN-bit datapath at 64 bits is wide. Routing of the partial-product matrix to the reduction tree is the dominant physical-design challenge. Floorplanning should be done early.
|
||||
|
||||
## Verification Considerations
|
||||
|
||||
- **`MULH*` cross-product.** The signed/unsigned interaction of the upper-product instruction is the most common source of bugs in DIY MDUs. The full enumeration is: `MULH` = `±×±`, `MULHSU` = `±×u`, `MULHU` = `u×u`, `MUL` = lower half (operand signedness does not change the bit pattern). Testing must cover all four cases at boundary patterns: 0, 1, −1, 2^XLEN−1, 2^(XLEN−1) and combinations thereof.
|
||||
- **Divide corner cases.** Quotient-remainder correctness at: divide-by-zero, most-negative-dividend by −1, 0/anything, anything/1, anything/−1 (signed and unsigned). The most-negative/÷−1 case is the dominant bug class; it is a signed-overflow corner that exists at any radix.
|
||||
- **Constant-time verification.** If the XH-1 documentation claims constant-time `MULH*` or `DIV`, that property must be verified at the gate level against the implementation. This is non-trivial; the project should either explicitly claim constant-time and verify it, or explicitly not claim it.
|
||||
- **Formal verification.** SRT quotient-digit selection is a small enough state machine to be formally verified; the surrounding microarchitecture (operand sign extension, the most-negative/÷−1 corner) usually is not, and is covered by directed tests.
|
||||
- **128-core replication.** A single strong verification of the per-core MDU logic is necessary and sufficient for the per-core logic itself. It is not sufficient for the full system: per-tile variability (manufacturing, voltage, timing), tile-level integration (interrupts during a long `DIV`, cross-tile coherence interactions for shared-pool MDU configurations, and the shared-pool arbitration logic) are system-level concerns that do not collapse to single-core verification. The earlier draft overstated this; the corrected position is that per-core logic verification is a prerequisite, not a complete, system-level verification.
|
||||
- **Reset, debug, OS save/restore.** A long-running iterative `DIV` interacts with the OS's context-switch decision. If the MDU exposes internal state to the OS, the OS must save it; if it does not, the OS must wait for the divide to complete. The chosen model must be documented and tested.
|
||||
|
||||
## Software Considerations
|
||||
|
||||
- **Compiler.** Modern GCC/LLVM emit `MUL`/`MULH*` for native multiplies and emit `DIV`/`REM` for variable division. By-constant division is reduced to a multiply-shift sequence. The MDU must service what the compiler emits, but does not need to be a hero.
|
||||
- **Runtimes.** `muldi3`, `divdi3`, etc. in libgcc are used when the hardware `M` is unavailable. The presence of `M` removes the need for these; the MDU must be correct enough that software is happy to use it.
|
||||
- **Cryptographic code.** Side-channel-resistant code (lattice crypto, big-integer arithmetic) wants constant-time multiplies and divides, and often wants the high half of a product. A working `MULH*` is a hard requirement for any software doing 128-bit arithmetic on a 64-bit machine.
|
||||
- **`Zbkb` and bitmanip-for-crypto.** `Zbkb` includes bitmanipulation operations such as `BREV8`, `PACK`, `UNPACK`, `ZIP`, `UNZIP`, `ANDN`, `ORN`, `XNOR`, and carry-less multiply instructions (`CLMUL`, `CLMULH`, `CLMULR`). The carry-less multiply instructions are *not* the same operation as `MULH*`; they use XOR in place of the carry-propagate addition in the reduction tree. The earlier draft conflated these; the corrected position is that `Zbkb` does not specifically use `MULH*` heavily, and any carry-less multiply support is a separate datapath concern outside the `M`-extension MDU.
|
||||
- **Vector and tensor code.** Inner loops in BLAS and ML kernels use `MUL` and FMA heavily; the MDU is a contributor, but in an XH-1 scalar core context it is not the dominant execution unit.
|
||||
- **Operating system.** The OS's context-switch code does not generally divide; interrupts during a long `DIV` are the only software-visible oddity. A documented, predictable `DIV` latency is what the OS wants.
|
||||
|
||||
## Recommendation
|
||||
|
||||
**INSUFFICIENT EVIDENCE** to recommend a specific MDU microarchitecture.
|
||||
|
||||
The repository context does not establish whether the XH-1 core is in-order or out-of-order, whether it implements a vector unit, what its target frequency and process node are, or whether a shared-pool MDU configuration is in scope. These are the inputs that determine whether a small iterative multiplier, a pipelined Booth multiplier, or a high-radix SRT divider is the right answer.
|
||||
|
||||
What can be recommended with the evidence available:
|
||||
|
||||
- RECOMMENDATION: pick one operand width (32 or 64) and a corresponding unified 2·XLEN-bit multiplier that serves both `MUL` and `MULH*`. The `MUL` lower half is bit-identical regardless of operand signedness, so a single 2·XLEN-bit reduction tree is the cost-effective choice. The W-variants require a separate small datapath or a sign-extension fixup; the trade-off should be made explicit.
|
||||
- RECOMMENDATION: avoid Newton–Raphson as the only divider. It is excellent in throughput-oriented contexts and inappropriate for a scalar core that must expose a deterministic `DIV` latency to the compiler.
|
||||
- RECOMMENDATION: implement the SRT most-negative / −1 corner case explicitly and verify it formally; this is the single most common bug in homemade MDUs. The corner exists at any radix for signed division and is not specific to radix ≥ 4.
|
||||
- RECOMMENDATION: do not optimize for early-exit on `MULH*` or `DIV` unless the XH-1 is willing to give up the constant-time property that cryptographic software expects.
|
||||
- RECOMMENDATION: power-gate the MDU on idle tiles. On a 128-core die this is a meaningful contributor to idle power. INSUFFICIENT EVIDENCE to quantify.
|
||||
- RECOMMENDATION: consider a small shared pool of MDUs (e.g., 8–32) as an alternative to full per-core replication, and decide based on area, cross-tile routing cost, and `MUL` throughput targets. The binary "1-per-core or 1-shared" framing is incomplete.
|
||||
|
||||
## Confidence
|
||||
|
||||
- FACT: high confidence in the RISC-V `M` ISA's required operations and corner-case behavior. This is normative and stable in the Unprivileged ISA Specification; the specific version is not pinned in this stub because the XH-1 repository does not pin a version.
|
||||
- ASSUMPTION: the XH-1 is a tiled 128-core fabric where global MDU sharing is rejected and per-tile replication or a small shared pool is the only realistic option. Reasonable but not stated by the repository.
|
||||
- ASSUMPTION: RV64, scalar-only, in-order, single clock domain with per-tile clock gating. Stated as assumptions above; INSUFFICIENT EVIDENCE to choose otherwise.
|
||||
- ASSUMPTION: latency and area figures are in qualitative, not quantitative, terms. The relative orderings are well-known; the absolute numbers are not.
|
||||
- PROPOSAL: the recommendations above are not the only valid choices; they are the conservative ones given the missing context.
|
||||
- OPEN: in-order vs. out-of-order, frequency target, process node, vector-unit presence, shared-pool acceptability, bitmanip/crypto extension presence, constant-time documentation status.
|
||||
|
||||
## Open Questions
|
||||
|
||||
1. Is the XH-1 core in-order or out-of-order? This determines whether an iterative MDU's long latency is acceptable or whether a pipelined MDU is needed.
|
||||
2. Is the XH-1 core RV32 or RV64? This determines whether the MDU is 32×32 or 64×64. `MULH*` is a native instruction in both; the description in the original draft was incorrect on this point.
|
||||
3. What is the target clock frequency and the process node? These determine whether a combinational multiplier fits in one cycle or must be pipelined.
|
||||
4. Does the XH-1 implement an integer FMA or fused MAC? If so, it changes the MDU's role from a producer of products to a producer of products-and-sums, and a different datapath is required. `Zfa` is a floating-point extension and not the appropriate home for integer FMA.
|
||||
5. Does the XH-1 implement any bitmanipulation (`B`) or cryptography (`K`) extensions? In particular, `Zbkb` uses carry-less multiplies (`CLMUL*`), which are not the same as `MULH*` and would be a separate datapath.
|
||||
6. Is constant-time execution a documented property of the XH-1? This constrains MDU microarchitecture.
|
||||
7. What is the interrupt latency target? A non-interruptible long `DIV` may be unacceptable; an interruptible one requires exposing internal state.
|
||||
8. Is there a vector unit sharing the MDU? A scalar-only context lets the MDU be single-issue; a shared MDU changes the throughput requirements.
|
||||
9. Is a shared pool of MDUs (e.g., 8–32) in scope, or is per-core replication fixed?
|
||||
10. How is the XH-1 floorplanned? A 128-core die imposes physical-design constraints on the MDU footprint.
|
||||
11. What is the software stack's expected division-heavy workload? Cryptographic and big-integer code is `MULH*`-heavy; HPC is `MUL`/FMA-heavy; general-purpose is `DIV`-light. Without an application target, the right balance is unknown.
|
||||
12. What is the clocking model? Single domain, per-tile domains, DVFS? This affects power gating and cross-tile MDU sharing.
|
||||
|
||||
## Sources
|
||||
|
||||
INSUFFICIENT EVIDENCE.
|
||||
|
||||
This document deliberately does not invent citations. The RISC-V `M` extension's required operations, corner cases, and architectural behaviors are normative in the *Unprivileged ISA Specification* (RISC-V International), but a specific section and version are not cited here because the XH-1 repository does not pin a version.
|
||||
|
||||
Established microarchitectural references that would normally be cited — descriptions of SRT division, Booth recoding, Wallace and Dadda trees, on-the-fly quotient conversion, and the most-negative / −1 corner case fixup — appear in standard computer-arithmetic textbooks and in well-known survey papers (for example: Parhami, *Computer Arithmetic*; Ercegovac and Lang, *Digital Arithmetic*; the original SRT paper by Robertson, Sweeney, and Tocher; the Booth-recoding paper; and Dadda's paper on reduction trees), but no specific work is cited here because fabricating a paper title, page number, or equation would violate the project's "never invent citations" rule.
|
||||
|
||||
Where this document makes quantitative claims (e.g., "an SRT-4 64-bit divider is on the order of 32 quotient digits"), the numbers are qualitative estimates, not measured values; the surrounding text says so. If the XH-1 project requires numbers backed by a specific source, the relevant literature should be located and cited by the project's documentation owner.
|
||||
+241
File diff suppressed because one or more lines are too long
+2
@@ -0,0 +1,2 @@
|
||||
|
||||
|
||||
@@ -1,3 +1,260 @@
|
||||
# mul div unit
|
||||
# Multiply/Divide Unit (MDU) Research
|
||||
|
||||
SOON
|
||||
## Status
|
||||
|
||||
Stub document. No XH-1 design decisions are committed. This document surveys established techniques and frames the design space for the XH-1 multiplier/divider; it does not invent measurements, benchmarks, or fabricated citations.
|
||||
|
||||
## Assumptions and Scope
|
||||
|
||||
The following assumptions frame the analysis. They are stated explicitly because the XH-1 repository context is not provided in this stub; where evidence is missing, the document says so.
|
||||
|
||||
- **ISA width (XLEN).** ASSUMPTION: the XH-1 core is RV64. Rationale: 128-bit results in `MULH*` are most useful when software performs multi-word arithmetic, and a 128-core tiled die is consistent with a server-class 64-bit core. RV32 is treated as a secondary case.
|
||||
- **Core pipeline style.** ASSUMPTION: the core is in-order with multiple execution stages, to keep the discussion concrete. Out-of-order is acknowledged as a different design point.
|
||||
- **Clocking.** ASSUMPTION: a single chip-wide clock domain with per-tile clock gating. Per-tile DVFS is not assumed. INSUFFICIENT EVIDENCE to choose otherwise.
|
||||
- **Process node / cell library.** Not assumed. The document discusses area and power only in qualitative, relative terms.
|
||||
- **Frequency target.** Not assumed. Latency claims are stated in cycles, not nanoseconds.
|
||||
- **Vector / shared MDU.** ASSUMPTION: scalar-only context without a vector unit sharing the MDU.
|
||||
- **Extensions.** ASSUMPTION: base RISC-V `M` is implemented. `B`, `K`, and vector-crypto extensions are not assumed present; their interaction with the MDU is discussed only as a scaling consideration.
|
||||
|
||||
INSUFFICIENT EVIDENCE for any of the above where it would change a recommendation.
|
||||
|
||||
## Abstract
|
||||
|
||||
The multiply/divide unit (MDU) executes the RISC-V `M`-extension instructions on each XH-1 core. In a 128-core machine replicated across a tiled fabric, the MDU is a notable contributor to per-core area and to the critical path, while rarely being the limiter of sustained throughput. This document reviews the design space (array vs. tree multipliers, radix selection, division algorithms, divide latency hiding, fused MAC, early-exit handling, and the treatment of `MULH`/`MULHU`/`MULHSU`) and surfaces the trade-offs that interact with the rest of the XH-1 core and with multi-core scalability.
|
||||
|
||||
## Research Question
|
||||
|
||||
What MDU microarchitecture for the XH-1 core best balances the following, given that the core is replicated 128 times on die and shares die-wide resources through a tiled fabric?
|
||||
|
||||
- Per-operation latency for signed/unsigned multiply, multiply-high, and signed/unsigned divide/remainder.
|
||||
- Sustained throughput (cycles/issue) at 1 MDU per core.
|
||||
- Area per core (silicon cost × 128).
|
||||
- Worst-case power and energy per operation.
|
||||
- Critical-path impact on the core's clock period.
|
||||
- Verification complexity (correctness across the 128-bit, signed/unsigned, divide-by-zero, and overflow corner cases of the RISC-V `M` extension).
|
||||
- Software implications: predictable timing, RV64 vs. RV32 register-file layout for `MULH*`, and constant-time considerations for cryptographic code.
|
||||
|
||||
## Background
|
||||
|
||||
### RISC-V `M` Extension Requirements
|
||||
|
||||
The RISC-V `M` extension defines a small, orthogonal set of operations, all operating on the base integer register width (XLEN = 32 for RV32, 64 for RV64):
|
||||
|
||||
- `MUL` / `MULW` — lower XLEN bits of a product. The lower bits of an integer product are bit-identical regardless of whether the operands are treated as signed or unsigned, so `MUL` does not require a signedness mode.
|
||||
- `MULH` — upper XLEN bits, signed × signed.
|
||||
- `MULHU` — upper XLEN bits, unsigned × unsigned.
|
||||
- `MULHSU` — upper XLEN bits, signed × unsigned.
|
||||
- `DIV` / `DIVU` / `DIVW` / `DIVUW` — signed/unsigned quotient, truncated toward zero.
|
||||
- `REM` / `REMU` / `REMW` / `REMUW` — remainder, sign follows the dividend.
|
||||
- For RV64, the `W` variants operate on 32-bit values and sign-extend the 32-bit result to 64 bits.
|
||||
|
||||
`MULH*` is the operation that forces the hardware to compute a 2·XLEN-bit product; this is the dominant cost in the multiply datapath. The `MUL` lower result is normally a free byproduct of that same 2·XLEN product on a unified datapath.
|
||||
|
||||
`MULH*` exists as native instructions in both RV32 and RV64. In RV32 the product is 64 bits; `MULH`/`MULHU`/`MULHSU` return the upper 32 bits. They are not emulated as paired 32-bit halves; the upper 32 bits of the full product are produced directly.
|
||||
|
||||
The `M` extension specifies architectural behavior (truncation, signed/unsigned interaction, divide-by-zero, the most-negative-dividend-divided-by-−1 overflow) but does not prescribe a microarchitecture, latency, or throughput. This makes `M` a high-leverage, ISA-allowed design decision.
|
||||
|
||||
### Why `M` Matters on XH-1
|
||||
|
||||
In a tiled 128-core processor, each core typically has its own integer `M` unit rather than a shared, chip-wide multiply unit. The reasons are:
|
||||
|
||||
- The wire delay of routing two 64-bit operands to a shared unit at die-crossing distance is incompatible with low-latency operation.
|
||||
- Multiplication is on the critical path of many kernels (FFT inner loops, matrix arithmetic, hash functions, address computation in some interpreters).
|
||||
- Local replication, while more silicon, is the conventional answer.
|
||||
|
||||
The MDU therefore contributes to the area of every tile. The product of its area cost and 128 is a first-order driver of die cost.
|
||||
|
||||
## Existing Approaches
|
||||
|
||||
### Multiplication
|
||||
|
||||
- **Carry-save adder (CSA) array.** A simple, dense, rectangular array of full adders that reduces partial-product bits. Without Booth recoding, an N×N array has N rows of full adders (N stages of reduction); with radix-4 Booth recoding it has N/2 rows. Latency scales linearly with the number of rows. Easy to layout; long critical path.
|
||||
- **Wallace tree.** A logarithmic-depth reduction tree using CSAs; faster than an array at the same operand width, with irregular shape that complicates physical design. Whether a Wallace tree is smaller than a CSA array in total gate count at 64×64 is implementation-dependent; the literature is mixed.
|
||||
- **Dadda tree.** A variant of Wallace that uses a slightly larger first stage to reduce the number of subsequent reductions. Often compared in the literature as a near-equivalent point in the area/time design space.
|
||||
- **Booth-recoded multiplier.** Recodes one operand to reduce the partial-product count. Radix-4 Booth produces N/2 partial products for an N-bit operand and is the workhorse in many cores. Radix-8 produces N/3 partial products; the per-PP selector is more complex (roughly tripling selector logic per row) while the reduction tree itself grows in line with the row count, not by a factor of three. Recoding is most useful when the operand width is large.
|
||||
- **Iterative multiplier.** Uses a small, fixed datapath and iterates over the operand width. Saves area but increases latency to many cycles; throughput is one multiply every N cycles unless multiple independent multiplies are interleaved.
|
||||
- **Fused multiply–add (FMA).** Single instruction producing `a*b + c` with one rounding. Standard in vector/GPU ISAs; not part of base scalar RISC-V `M`, but a candidate addition. RISC-V `Zfa` is a floating-point extension and is not the appropriate home for an integer FMA; any integer FMA on XH-1 would be a vendor extension.
|
||||
- **Unified 2·XLEN-bit product.** A single datapath producing the full double-width product, from which both `MUL` and `MULH*` are sliced. Standard approach; reduces logic vs. two separate multipliers but requires a wide adder at the end and is not a free win for the `W` variants (see Implementation Considerations).
|
||||
|
||||
### Division
|
||||
|
||||
- **Non-restoring and restoring shift/subtract divider.** A 2·XLEN-iteration loop that produces one quotient bit per cycle; latency scales linearly with operand width. Simple but slow.
|
||||
- **SRT divider (radix-2, radix-4, radix-8, radix-16).** A redundant representation (carry-save) of the partial remainder allows selection of a small set of quotient digits per cycle. Latency scales as 2·XLEN / log2(radix) quotient digits, plus a small constant for fixup. SRT is the dominant technique for high-performance scalar cores. The most-negative-dividend-divided-by-−1 signed-overflow case is a fundamental property of signed division at any radix; it is not specific to radix ≥ 4.
|
||||
- **Newton–Raphson reciprocal + multiply.** Two multiplications of the reciprocal approximation, then a final multiply by the dividend. Latency roughly 2–3 multiplies, with small additional control. High throughput, long tail latency, large area (needs a fast multiplier and a ROM of initial approximations). Inappropriate as the only divider in a scalar core that needs deterministic `DIV` latency; often used in vector/GPU contexts.
|
||||
- **Goldschmidt division.** Iterative convergence with different numerical subtleties than Newton–Raphson; same general area/latency class. Not analyzed further here.
|
||||
- **Lookup-table-based constant division.** Replacing division by a small set of "magic numbers" at compile time. Software-side; interacts with the hardware because the `M` ISA is not required for software to be efficient if the compiler reduces a constant division to a multiply-shift. The hardware must still service any `DIV` it does see.
|
||||
- **Divider bypass / radix-2^k with table-driven selection.** Industry workhorse: a redundant (carry-save) partial remainder plus a small quotient-digit selector table per radix step. Closely related to SRT.
|
||||
|
||||
### Multiply-Accumulate and Fused Operations
|
||||
|
||||
- **Fused MAC.** `a*b + c` in one cycle of the MAC pipeline, sharing the partial-product reduction tree with a free accumulator adder.
|
||||
- **Integer FMA / fused MAC.** Not in the RISC-V `M` extension; would be a vendor extension if added. Multiply-add with rounding is a floating-point concept.
|
||||
|
||||
## Alternative Designs
|
||||
|
||||
The design space reduces to a small set of choices:
|
||||
|
||||
1. **Unified 2·XLEN-bit multiplier for `MUL` + `MULH*`.** Common choice for RV64. Produces the full product once and slices it.
|
||||
2. **Separate small multiplier for `MUL`/`MULW`, separate datapath for `MULH*`.** Avoids paying for the full 2·XLEN product on every multiply, at the cost of wider muxes and a longer `MULH*` critical path.
|
||||
3. **Array vs. tree.** A rectangular CSA array (regular, often slow) vs. a Wallace/Dadda tree (fast, irregular) vs. a Booth-recoded array (moderate area, moderate speed).
|
||||
4. **Iterative vs. combinational multiplier.** Iterative saves area but introduces multi-cycle latency and an extra pipeline stage (or stall) on `MULH*`.
|
||||
5. **Division algorithm.** SRT radix-2/4/8 vs. shift/subtract vs. Newton–Raphson vs. software-emulated via reciprocal multiply.
|
||||
6. **Pipelined MDU vs. non-pipelined.** Issue one MDU instruction per cycle (pipelined), one per N cycles (unpipelined), or some hybrid.
|
||||
7. **32-bit `W` variants.** Either share the XLEN-wide datapath and sign-extend at the end, or use a 32-bit-wide fast path.
|
||||
8. **Constant-time guarantees.** Some software (notably cryptographic) requires `MULH*` and `DIV*` to be data-independent in time. This constrains early-exit and early-out optimizations.
|
||||
9. **Shared pool of MDUs.** A small pool of MDUs (e.g., 8–32) serving 128 cores through the tile fabric, rather than one per core. Trades replication cost for cross-tile wire delay and contention. Discussed in 128-Core Scalability.
|
||||
10. **`MUL` only, emulate `MULH*`.** A scalar in-order core could implement only `MUL`/`MULW` and trap-and-emulate `MULH*` in software. Reduces the multiply datapath width at the cost of trapping on `MULH*`-using code. Crypto and big-integer arithmetic rely on `MULH*`, so this is a real design point only for cores targeting general-purpose code without crypto.
|
||||
|
||||
## Comparison
|
||||
|
||||
The trade-space can be characterized along a small number of axes. Without committing to a specific process node or to fabricated measurements, the qualitative relationships are:
|
||||
|
||||
- **Latency vs. area (multiply).** A combinational Wallace or Booth-radix-4 tree is faster than a CSA array of equivalent width. Whether the tree is also smaller in total gate count at 64×64 is implementation-dependent; the array is regular and the tree is not. An iterative multiplier is smallest in area but slowest in latency (linear in operand width per multiply).
|
||||
- **Latency vs. area (divide).** Radix-4 SRT divides in roughly half the quotient digits of radix-2 SRT at modest area increase; radix-8 trades more selector-table area for a further reduction in digit count. Newton–Raphson is fastest in latency on a fully pipelined multiplier but requires a fast multiplier and a reciprocal ROM.
|
||||
- **Throughput vs. latency.** Pipelined designs match the issue rate of the core but cost a register stage and additional bypassing; unpipelined designs cost only one execution slot but force back-to-back `MUL`s to serialize.
|
||||
- **Verification cost.** A small iterative multiplier is the easiest to formally reason about (one bit-slice repeated). A high-radix SRT with a partial-remainder selector and overlapped radix steps is the hardest.
|
||||
|
||||
The combinations that are not useful tend to be those that pay for a high-radix divider while leaving a slow multiplier next to it (the divider tail latency is masked by the slow multiplier, but the area is paid).
|
||||
|
||||
## Advantages
|
||||
|
||||
- A unified 2·XLEN-bit multiplier lets `MUL` and `MULH*` share the partial-product reduction tree, removing duplicated logic. This advantage is independent of whether the reduction is a CSA array or a tree.
|
||||
- A pipelined MDU removes a back-to-back multiply hazard at the cost of a single register stage, which is essentially free in any pipeline that already has a multi-cycle execution unit.
|
||||
- SRT division is well-studied, has well-known implementation recipes, and matches the area budget of most scalar cores.
|
||||
- A non-pipelined iterative divider is the smallest possible area for a working `DIV` and is acceptable when software rarely emits `DIV`.
|
||||
- A small shared pool of MDUs can reduce the area replication cost of 128 cores at the cost of cross-tile wire delay and per-MDU contention.
|
||||
|
||||
## Disadvantages
|
||||
|
||||
- A 64×64 → 128 multiplier, whether implemented as a CSA array or a reduction tree, is a wide datapath and a noticeable per-core cost in any high-density core; on a 128-core die, this multiplies.
|
||||
- Newton–Raphson division has long, variable latency that is hard to expose to the front end without reservation-station machinery that the rest of the core may not need.
|
||||
- SRT division has a most-negative-dividend-divided-by-−1 signed-overflow corner case at any radix; correctly handling it requires either an extra cycle or a small fix-up datapath, both of which need verification.
|
||||
- Early-exit optimizations (e.g., detecting a small result and short-circuiting a wide multiply) introduce data-dependent latency, which is a correctness hazard for some software and a verification hazard in any case.
|
||||
- A `W`-variant fast path that bypasses the upper 32 bits of the multiplier introduces a second timing path through the MDU.
|
||||
- A shared pool of MDUs requires cross-tile operand routing and adds to the NoC traffic budget; it also turns the MDU into a contended resource for 128 cores.
|
||||
|
||||
## XH-1 Considerations
|
||||
|
||||
- **XLEN.** Under the RV64 assumption, the MDU is a 64×64 → 128 datapath with a 64-bit signed/unsigned unit. If RV32, the MDU is 32×32 → 64. The `MULH*` requirement is the dominant datapath driver in either case.
|
||||
- **Pipeline depth.** The MDU's latency interacts with the core's pipeline. If the core is in-order with a single execution stage, an iterative multiplier is mandatory. If the core is in-order with multiple execution stages, a pipelined MDU fits naturally. If the core is out-of-order, the MDU's result is a producer into the register file through the wakeup/select path; latency is largely hidden, but area and worst-case occupancy still matter.
|
||||
- **Scalar-only context.** Under the scalar-only assumption, the MDU is a single-issue scalar unit; a non-pipelined design with a throughput of 1 per N cycles is architecturally acceptable if the compiler can be guided to use shifts and adds for short multiplies.
|
||||
- **Bypassing.** A pipelined MDU that issues one multiply per cycle needs a writeback port and a bypass network entry; this interacts with the register file. A non-pipelined iterative MDU needs only a writeback port and uses the issue queue's dependency tracking.
|
||||
- **Reset and OS save/restore.** A divide that takes > 50 cycles can be problematic on context switch if not interruptible. Most simple cores either don't accept interrupts mid-divide (it must complete) or save the partial-remainder registers. This is a design decision for the MDU.
|
||||
- **Clocking.** Under the single-clock-domain assumption, per-tile clock gating of the MDU is straightforward. Per-tile DVFS is not assumed.
|
||||
|
||||
## 128-Core Scalability
|
||||
|
||||
- **Replication cost.** The MDU's per-core area cost is multiplied by 128. For a unified 64×64 → 128 tree, the per-core cost is meaningful; for an iterative 32-bit-equivalent datapath, it is small.
|
||||
- **Single shared unit rejected.** A single die-wide MDU serving all 128 cores is generally not used for `M` operations because (a) the wire delay of moving two 64-bit operands to a central point at 128-core die dimensions is too long to keep MDU latency low, and (b) contention on a single MDU among 128 cores would make `MUL` throughput effectively a global bottleneck. ASSUMPTION: the XH-1 is a tiled fabric where the cost of a global MDU exceeds the cost of replication.
|
||||
- **Small shared pool.** A pool of K MDUs (e.g., K = 8 or 16) serving 128 cores is a real design point in some tiled architectures. It reduces replicated area by a factor of 128/K at the cost of cross-tile operand routing, MDU-side arbitration, and worst-case `MUL` throughput of K per cycle chip-wide. This binary "1-per-core or 1-shared" framing in earlier surveys is incomplete; the shared-pool option belongs in the design space.
|
||||
- **Variability.** Per-core MDUs make timing variability a per-tile concern. The chip-wide clock has to accommodate the slowest tile's MDU, so a fast core with a small MDU is paid for by every other tile. INSUFFICIENT EVIDENCE to bound the magnitude of this variability for XH-1.
|
||||
- **Power gating.** Per-core MDUs are excellent candidates for clock- or power-gating when a tile is idle. A 128-core die can shut down most MDUs during low utilization. This is a meaningful power-saving lever. INSUFFICIENT EVIDENCE to quantify the savings.
|
||||
|
||||
## Performance Considerations
|
||||
|
||||
- **Latency targets.** Without a specific frequency target for XH-1, latency cannot be quoted in nanoseconds. In cycle terms, qualitative estimates:
|
||||
- An iterative 64-bit multiplier: on the order of the operand width in cycles, depending on radix.
|
||||
- A pipelined 64-bit multiplier: 1–3 cycles issue-to-writeback, depending on pipeline depth.
|
||||
- An SRT-4 64-bit divider: on the order of 2·XLEN / log2(4) = 32 quotient digits plus a few cycles of fixup. The exact cycle count depends on digits-per-cycle and overlap.
|
||||
- An SRT-8 64-bit divider: on the order of 2·XLEN / log2(8) ≈ 21–22 quotient digits plus fixup. The "16–20 cycles" figure sometimes seen in informal sources does not account for fixup cycles and should not be cited as a hard number.
|
||||
- Newton–Raphson: roughly 2 multiplies plus fixup; latency can be lower than SRT if the multiplier is fast, throughput is the same.
|
||||
- **Compiler guidance.** Modern compilers reduce `DIV` by a constant to a multiply-shift sequence; the runtime `DIV` is most often a variable divide. This argues for a real divider, but a slow one is acceptable.
|
||||
- **Software-emulated `DIV`.** A `DIV` can be replaced by a software Newton–Raphson routine in a few hundred instructions. This is a fallback the OS can use; it argues that the minimum acceptable MDU can be slow.
|
||||
|
||||
## Area Considerations
|
||||
|
||||
Rough relative area figures (not fabrication-specific, not quantitative):
|
||||
|
||||
- A 64×64 → 128 CSA array, unrecoded: ~N rows of full adders, rectangular, regular layout. Area is roughly proportional to operand width squared in the array portion.
|
||||
- A 64×64 → 128 Booth-radix-4 array: N/2 partial products, larger selector muxes per PP. Net smaller than a naive unrecoded array; less regular than a tree.
|
||||
- A 64×64 → 128 Wallace/Dadda tree: fewer full adders in total than a naive array, but irregular layout. Whether total gate count is smaller than the array at 64×64 is implementation-dependent; the literature is mixed and the claim should not be asserted as universal.
|
||||
- A radix-4 SRT divider: similar order of magnitude to a small multiplier; mostly selector logic and a small ROM.
|
||||
- A Newton–Raphson reciprocal unit: negligible hardware beyond a fast multiplier; needs a small ROM of initial approximations.
|
||||
|
||||
For a 128-core replication, the multiplier's area contribution per tile is the first-order concern; the divider's is a second-order concern. The area discussion in the original draft conflated a unified tree with a CSA array; a unified 2·XLEN-bit product can be implemented as either organization, and the area comparison should be between the unified and split-datapath options, not between a tree and an array.
|
||||
|
||||
## Power and Energy Considerations
|
||||
|
||||
- **Switching activity.** Multiplication has high switching activity because every partial product is recomputed every cycle in a non-pipelined iterative design, while a fully combinational design has a single very-wide switching event. Energy is generally dominated by the partial-product reduction tree; the choice of array vs. tree changes both energy per op and the energy profile.
|
||||
- **Clock gating.** A pipelined MDU can clock-gate stages when no multiply is in flight; an iterative MDU only switches the active stage. Both are effective.
|
||||
- **Power gating.** A 128-core die will spend meaningful time with some tiles idle. Power-gating the MDU on idle tiles is a strong lever. INSUFFICIENT EVIDENCE to quantify the savings.
|
||||
- **Divide energy.** A long, slow divider spends many cycles driving the same datapath at full toggle rate. Whether a faster, larger divider is lower energy per `DIV` than a slow, small one depends on the specific organizations; a high-radix SRT with a large selector ROM can be higher energy per operation than a shift/subtract divider in some implementations. The blanket "faster = lower energy" claim is not generally true and is removed.
|
||||
- **Constant-time software.** Software that needs data-independent timing (crypto) requires the MDU to not take data-dependent shortcuts. This is a correctness property for the MDU; it rules out early-exit optimizations that would change the energy/latency profile.
|
||||
|
||||
## Implementation Considerations
|
||||
|
||||
- **Radix choice.** Radix-4 is the workhorse; radix-8 increases selector complexity per row for a smaller reduction in digit count. For an XLEN-64 design, radix-4 with carry-save partial remainder is the conventional balance.
|
||||
- **Carry-save throughout.** Keeping the partial remainder in carry-save form throughout the divide avoids a wide carry-propagate adder and removes a long wire on the critical path.
|
||||
- **On-the-fly quotient conversion.** The quotient emerges in a redundant form and must be converted to binary on the fly; this is a known microarchitectural module with well-understood area and timing.
|
||||
- **Most-negative / -1 corner.** Must be handled explicitly. This is a signed-division overflow at any radix, not specific to radix ≥ 4. The standard fix: detect the corner case before the final iteration and produce the architectural result directly.
|
||||
- **Divide-by-zero.** The architectural behavior is to return all-ones for the quotient and the dividend for the remainder, for both signed and unsigned. Hardware cost: a small fixup mux.
|
||||
- **`MULH*` bypass.** If `MULH*` is rarely used by compiled code, the MDU can be optimized for the `MUL` case and pay a small extra latency on `MULH*`. The ISA does not allow architectural shortcuts, only microarchitectural.
|
||||
- **`MULW` / `DIVW` / `REMW`.** On a unified 64×64 → 128 datapath, `MULW` requires either running a narrower 32×32 → 64 datapath or masking and sign-extending the lower 32 bits of the 64×64 product. The "free byproduct" framing for `MUL` does not extend cleanly to the `W` variants; the W-variants either share the wide datapath with sign-extension at the end (no real area saving) or use a 32-bit-wide fast path (area saved but a second critical path to verify).
|
||||
- **Physical design.** A 2·XLEN-bit datapath at 64 bits is wide. Routing of the partial-product matrix to the reduction tree is the dominant physical-design challenge. Floorplanning should be done early.
|
||||
|
||||
## Verification Considerations
|
||||
|
||||
- **`MULH*` cross-product.** The signed/unsigned interaction of the upper-product instruction is the most common source of bugs in DIY MDUs. The full enumeration is: `MULH` = `±×±`, `MULHSU` = `±×u`, `MULHU` = `u×u`, `MUL` = lower half (operand signedness does not change the bit pattern). Testing must cover all four cases at boundary patterns: 0, 1, −1, 2^XLEN−1, 2^(XLEN−1) and combinations thereof.
|
||||
- **Divide corner cases.** Quotient-remainder correctness at: divide-by-zero, most-negative-dividend by −1, 0/anything, anything/1, anything/−1 (signed and unsigned). The most-negative/÷−1 case is the dominant bug class; it is a signed-overflow corner that exists at any radix.
|
||||
- **Constant-time verification.** If the XH-1 documentation claims constant-time `MULH*` or `DIV`, that property must be verified at the gate level against the implementation. This is non-trivial; the project should either explicitly claim constant-time and verify it, or explicitly not claim it.
|
||||
- **Formal verification.** SRT quotient-digit selection is a small enough state machine to be formally verified; the surrounding microarchitecture (operand sign extension, the most-negative/÷−1 corner) usually is not, and is covered by directed tests.
|
||||
- **128-core replication.** A single strong verification of the per-core MDU logic is necessary and sufficient for the per-core logic itself. It is not sufficient for the full system: per-tile variability (manufacturing, voltage, timing), tile-level integration (interrupts during a long `DIV`, cross-tile coherence interactions for shared-pool MDU configurations, and the shared-pool arbitration logic) are system-level concerns that do not collapse to single-core verification. The earlier draft overstated this; the corrected position is that per-core logic verification is a prerequisite, not a complete, system-level verification.
|
||||
- **Reset, debug, OS save/restore.** A long-running iterative `DIV` interacts with the OS's context-switch decision. If the MDU exposes internal state to the OS, the OS must save it; if it does not, the OS must wait for the divide to complete. The chosen model must be documented and tested.
|
||||
|
||||
## Software Considerations
|
||||
|
||||
- **Compiler.** Modern GCC/LLVM emit `MUL`/`MULH*` for native multiplies and emit `DIV`/`REM` for variable division. By-constant division is reduced to a multiply-shift sequence. The MDU must service what the compiler emits, but does not need to be a hero.
|
||||
- **Runtimes.** `muldi3`, `divdi3`, etc. in libgcc are used when the hardware `M` is unavailable. The presence of `M` removes the need for these; the MDU must be correct enough that software is happy to use it.
|
||||
- **Cryptographic code.** Side-channel-resistant code (lattice crypto, big-integer arithmetic) wants constant-time multiplies and divides, and often wants the high half of a product. A working `MULH*` is a hard requirement for any software doing 128-bit arithmetic on a 64-bit machine.
|
||||
- **`Zbkb` and bitmanip-for-crypto.** `Zbkb` includes bitmanipulation operations such as `BREV8`, `PACK`, `UNPACK`, `ZIP`, `UNZIP`, `ANDN`, `ORN`, `XNOR`, and carry-less multiply instructions (`CLMUL`, `CLMULH`, `CLMULR`). The carry-less multiply instructions are *not* the same operation as `MULH*`; they use XOR in place of the carry-propagate addition in the reduction tree. The earlier draft conflated these; the corrected position is that `Zbkb` does not specifically use `MULH*` heavily, and any carry-less multiply support is a separate datapath concern outside the `M`-extension MDU.
|
||||
- **Vector and tensor code.** Inner loops in BLAS and ML kernels use `MUL` and FMA heavily; the MDU is a contributor, but in an XH-1 scalar core context it is not the dominant execution unit.
|
||||
- **Operating system.** The OS's context-switch code does not generally divide; interrupts during a long `DIV` are the only software-visible oddity. A documented, predictable `DIV` latency is what the OS wants.
|
||||
|
||||
## Recommendation
|
||||
|
||||
**INSUFFICIENT EVIDENCE** to recommend a specific MDU microarchitecture.
|
||||
|
||||
The repository context does not establish whether the XH-1 core is in-order or out-of-order, whether it implements a vector unit, what its target frequency and process node are, or whether a shared-pool MDU configuration is in scope. These are the inputs that determine whether a small iterative multiplier, a pipelined Booth multiplier, or a high-radix SRT divider is the right answer.
|
||||
|
||||
What can be recommended with the evidence available:
|
||||
|
||||
- RECOMMENDATION: pick one operand width (32 or 64) and a corresponding unified 2·XLEN-bit multiplier that serves both `MUL` and `MULH*`. The `MUL` lower half is bit-identical regardless of operand signedness, so a single 2·XLEN-bit reduction tree is the cost-effective choice. The W-variants require a separate small datapath or a sign-extension fixup; the trade-off should be made explicit.
|
||||
- RECOMMENDATION: avoid Newton–Raphson as the only divider. It is excellent in throughput-oriented contexts and inappropriate for a scalar core that must expose a deterministic `DIV` latency to the compiler.
|
||||
- RECOMMENDATION: implement the SRT most-negative / −1 corner case explicitly and verify it formally; this is the single most common bug in homemade MDUs. The corner exists at any radix for signed division and is not specific to radix ≥ 4.
|
||||
- RECOMMENDATION: do not optimize for early-exit on `MULH*` or `DIV` unless the XH-1 is willing to give up the constant-time property that cryptographic software expects.
|
||||
- RECOMMENDATION: power-gate the MDU on idle tiles. On a 128-core die this is a meaningful contributor to idle power. INSUFFICIENT EVIDENCE to quantify.
|
||||
- RECOMMENDATION: consider a small shared pool of MDUs (e.g., 8–32) as an alternative to full per-core replication, and decide based on area, cross-tile routing cost, and `MUL` throughput targets. The binary "1-per-core or 1-shared" framing is incomplete.
|
||||
|
||||
## Confidence
|
||||
|
||||
- FACT: high confidence in the RISC-V `M` ISA's required operations and corner-case behavior. This is normative and stable in the Unprivileged ISA Specification; the specific version is not pinned in this stub because the XH-1 repository does not pin a version.
|
||||
- ASSUMPTION: the XH-1 is a tiled 128-core fabric where global MDU sharing is rejected and per-tile replication or a small shared pool is the only realistic option. Reasonable but not stated by the repository.
|
||||
- ASSUMPTION: RV64, scalar-only, in-order, single clock domain with per-tile clock gating. Stated as assumptions above; INSUFFICIENT EVIDENCE to choose otherwise.
|
||||
- ASSUMPTION: latency and area figures are in qualitative, not quantitative, terms. The relative orderings are well-known; the absolute numbers are not.
|
||||
- PROPOSAL: the recommendations above are not the only valid choices; they are the conservative ones given the missing context.
|
||||
- OPEN: in-order vs. out-of-order, frequency target, process node, vector-unit presence, shared-pool acceptability, bitmanip/crypto extension presence, constant-time documentation status.
|
||||
|
||||
## Open Questions
|
||||
|
||||
1. Is the XH-1 core in-order or out-of-order? This determines whether an iterative MDU's long latency is acceptable or whether a pipelined MDU is needed.
|
||||
2. Is the XH-1 core RV32 or RV64? This determines whether the MDU is 32×32 or 64×64. `MULH*` is a native instruction in both; the description in the original draft was incorrect on this point.
|
||||
3. What is the target clock frequency and the process node? These determine whether a combinational multiplier fits in one cycle or must be pipelined.
|
||||
4. Does the XH-1 implement an integer FMA or fused MAC? If so, it changes the MDU's role from a producer of products to a producer of products-and-sums, and a different datapath is required. `Zfa` is a floating-point extension and not the appropriate home for integer FMA.
|
||||
5. Does the XH-1 implement any bitmanipulation (`B`) or cryptography (`K`) extensions? In particular, `Zbkb` uses carry-less multiplies (`CLMUL*`), which are not the same as `MULH*` and would be a separate datapath.
|
||||
6. Is constant-time execution a documented property of the XH-1? This constrains MDU microarchitecture.
|
||||
7. What is the interrupt latency target? A non-interruptible long `DIV` may be unacceptable; an interruptible one requires exposing internal state.
|
||||
8. Is there a vector unit sharing the MDU? A scalar-only context lets the MDU be single-issue; a shared MDU changes the throughput requirements.
|
||||
9. Is a shared pool of MDUs (e.g., 8–32) in scope, or is per-core replication fixed?
|
||||
10. How is the XH-1 floorplanned? A 128-core die imposes physical-design constraints on the MDU footprint.
|
||||
11. What is the software stack's expected division-heavy workload? Cryptographic and big-integer code is `MULH*`-heavy; HPC is `MUL`/FMA-heavy; general-purpose is `DIV`-light. Without an application target, the right balance is unknown.
|
||||
12. What is the clocking model? Single domain, per-tile domains, DVFS? This affects power gating and cross-tile MDU sharing.
|
||||
|
||||
## Sources
|
||||
|
||||
INSUFFICIENT EVIDENCE.
|
||||
|
||||
This document deliberately does not invent citations. The RISC-V `M` extension's required operations, corner cases, and architectural behaviors are normative in the *Unprivileged ISA Specification* (RISC-V International), but a specific section and version are not cited here because the XH-1 repository does not pin a version.
|
||||
|
||||
Established microarchitectural references that would normally be cited — descriptions of SRT division, Booth recoding, Wallace and Dadda trees, on-the-fly quotient conversion, and the most-negative / −1 corner case fixup — appear in standard computer-arithmetic textbooks and in well-known survey papers (for example: Parhami, *Computer Arithmetic*; Ercegovac and Lang, *Digital Arithmetic*; the original SRT paper by Robertson, Sweeney, and Tocher; the Booth-recoding paper; and Dadda's paper on reduction trees), but no specific work is cited here because fabricating a paper title, page number, or equation would violate the project's "never invent citations" rule.
|
||||
|
||||
Where this document makes quantitative claims (e.g., "an SRT-4 64-bit divider is on the order of 32 quotient digits"), the numbers are qualitative estimates, not measured values; the surrounding text says so. If the XH-1 project requires numbers backed by a specific source, the relevant literature should be located and cited by the project's documentation owner.
|
||||
|
||||
Reference in New Issue
Block a user