Files
main-website/docs/cmt.md
T
2026-06-15 19:20:24 +02:00

3.3 KiB

Chip Multi-Threading (CMT)

ReviveSparc implements the UltraSPARC T1/T2 chip multi-threading architecture, allowing multiple threads to execute concurrently on each core.

Architecture

Core 0          Core 1          Core 2          Core 3
+--------+     +--------+     +--------+     +--------+
| T0 T1  |     | T0 T1  |     | T0 T1  |     | T0 T1  |
| T2 T3  |     | T2 T3  |     | T2 T3  |     | T2 T3  |
+--------+     +--------+     +--------+     +--------+
     |              |              |              |
     +--------------+--------------+--------------+
                        |
                  +-----------+
                  | L2 Cache  |
                  +-----------+

Each core supports up to 4 hardware threads with fine-grained interleaving. The pipeline switches between threads every cycle, hiding memory latency and maximizing throughput.

Thread Scheduling

Threads are scheduled using a round-robin policy with two modes:

Fine-Grained (Default)

The pipeline switches to a different thread each cycle. If the selected thread is stalled (cache miss, long latency operation), the pipeline switches to the next ready thread.

// Pseudocode for thread selection
while (true) {
    for (int i = 0; i < NUM_THREADS; i++) {
        thread = (current_thread + i) % NUM_THREADS;
        if (thread_is_ready(thread)) {
            execute(thread);
            break;
        }
    }
    current_thread = (current_thread + 1) % NUM_THREADS;
}

Coarse-Grained

Each thread runs for a fixed quantum (configurable, default 64 cycles) before yielding the pipeline. Useful for workloads with good instruction-level parallelism.

Thread State

Each thread has its own fully replicated architectural state:

  • Register window (16 registers, with window pointer)
  • PC and NPC (program counters)
  • PSTATE (processor state register)
  • TL (trap level, up to 4)
  • Interrupt state (pending interrupts, PIL)

Memory Model

The T1/T2 memory model is Total Store Order (TSO) with per-thread write buffers. ReviveSparc implements:

  • TSO by default (SPARC V9 default)
  • Partial Store Order (PSO) as an alternative
  • Relaxed Memory Order (RMO) for maximum performance

Memory Barrier Instructions

Instruction Description
MEMBAR #Sync Full memory barrier
MEMBAR #StoreLoad Store before Load ordering
MEMBAR #StoreStore Store before Store ordering
MEMBAR #LoadLoad Load before Load ordering
MEMBAR #LoadStore Load before Store ordering
MEMBAR #Lookaside Invalidate lookaside buffers

Configuration

CMT parameters are configured at core initialization:

sparc_core_config_t config = {
    .num_threads   = 4,
    .sched_mode    = SCHED_FINE_GRAINED,
    .pipeline_mode = TSO,
    .quantum       = 64   // coarse-grained quantum
};
sparc_core_t *core = sparc_core_new(&config);

Or via command line:

./sparc-emu -cores 4 -threads 4 -sched fine kernel.bin

Performance Considerations

  • Fine-grained threading works best with high cache-miss rates (database, web serving)
  • Coarse-grained threading works best with compute-intensive workloads
  • The optimal thread count depends on the L2 cache size and memory latency
  • Monitor thread utilization with the -stats flag:
    ./sparc-emu -stats kernel.bin