3.3 KiB
Chip Multi-Threading (CMT)
ReviveSparc implements the UltraSPARC T1/T2 chip multi-threading architecture, allowing multiple threads to execute concurrently on each core.
Architecture
Core 0 Core 1 Core 2 Core 3
+--------+ +--------+ +--------+ +--------+
| T0 T1 | | T0 T1 | | T0 T1 | | T0 T1 |
| T2 T3 | | T2 T3 | | T2 T3 | | T2 T3 |
+--------+ +--------+ +--------+ +--------+
| | | |
+--------------+--------------+--------------+
|
+-----------+
| L2 Cache |
+-----------+
Each core supports up to 4 hardware threads with fine-grained interleaving. The pipeline switches between threads every cycle, hiding memory latency and maximizing throughput.
Thread Scheduling
Threads are scheduled using a round-robin policy with two modes:
Fine-Grained (Default)
The pipeline switches to a different thread each cycle. If the selected thread is stalled (cache miss, long latency operation), the pipeline switches to the next ready thread.
// Pseudocode for thread selection
while (true) {
for (int i = 0; i < NUM_THREADS; i++) {
thread = (current_thread + i) % NUM_THREADS;
if (thread_is_ready(thread)) {
execute(thread);
break;
}
}
current_thread = (current_thread + 1) % NUM_THREADS;
}
Coarse-Grained
Each thread runs for a fixed quantum (configurable, default 64 cycles) before yielding the pipeline. Useful for workloads with good instruction-level parallelism.
Thread State
Each thread has its own fully replicated architectural state:
- Register window (16 registers, with window pointer)
- PC and NPC (program counters)
- PSTATE (processor state register)
- TL (trap level, up to 4)
- Interrupt state (pending interrupts, PIL)
Memory Model
The T1/T2 memory model is Total Store Order (TSO) with per-thread write buffers. ReviveSparc implements:
- TSO by default (SPARC V9 default)
- Partial Store Order (PSO) as an alternative
- Relaxed Memory Order (RMO) for maximum performance
Memory Barrier Instructions
| Instruction | Description |
|---|---|
MEMBAR #Sync |
Full memory barrier |
MEMBAR #StoreLoad |
Store before Load ordering |
MEMBAR #StoreStore |
Store before Store ordering |
MEMBAR #LoadLoad |
Load before Load ordering |
MEMBAR #LoadStore |
Load before Store ordering |
MEMBAR #Lookaside |
Invalidate lookaside buffers |
Configuration
CMT parameters are configured at core initialization:
sparc_core_config_t config = {
.num_threads = 4,
.sched_mode = SCHED_FINE_GRAINED,
.pipeline_mode = TSO,
.quantum = 64 // coarse-grained quantum
};
sparc_core_t *core = sparc_core_new(&config);
Or via command line:
./sparc-emu -cores 4 -threads 4 -sched fine kernel.bin
Performance Considerations
- Fine-grained threading works best with high cache-miss rates (database, web serving)
- Coarse-grained threading works best with compute-intensive workloads
- The optimal thread count depends on the L2 cache size and memory latency
- Monitor thread utilization with the
-statsflag:./sparc-emu -stats kernel.bin