← Back to projects

Single-Cycle RV32I Processor

Structural VHDL · CprE 3810 Project Part 1

A complete single-cycle RISC-V (RV32I) processor built from structural and dataflow VHDL. Its instruction-by-instruction trace matches the RARS simulator with zero mismatches on arithmetic, control-flow, and recursive merge-sort test programs.

Timeline
Summer 2026
ISA
RV32I · single-cycle
Team
Two-person team (IS_04)
Verification
0 mismatches vs. RARS

Overview

I built this processor with a teammate for Iowa State's CprE 3810 (Computer Organization and Assembly-Level Programming). The PC drives instruction memory. Instruction fields feed the control unit, immediate generator, and register file. Operand multiplexers choose RS1, PC, or zero as one ALU input, and RS2 or an immediate as the other. The ALU produces arithmetic, logic, comparison, shift, and address results; data memory and a load-extension unit handle word, halfword, and byte loads; and a four-way write-back multiplexer chooses between the ALU result, extended load data, and PC+4.

The PC resets to 0x00400000 to match RARS, and instruction memory is indexed by address bits 11:2.

Schematics & waveforms

Figures are from the project report.

Datapath components

Fetch / next-PC unit

A PC register, a PC+4 adder, a PC+immediate adder, branch-condition logic, and a next-PC multiplexer. The next PC is PC+4 (sequential or untaken branch), PC+immediate (taken branch or JAL), or (rs1 + imm) & ~1 for JALR, which has the highest priority.

Control unit

Purely combinational dataflow logic. It first decodes the opcode class (LUI, AUIPC, JAL, JALR, BRANCH, LOAD, STORE, OP-IMM, OP, SYSTEM), then generates RegWrite, MemWrite, MemToReg, ALUSrcA/B, ALUOp, Branch, Jump, JumpReg, ImmSel, and Halt. Instruction bit 30 separates ADD from SUB and SRL from SRA.

Register file

32 × 32-bit registers built structurally from reg_n primitives, with two asynchronous read ports and one synchronous write port. x0 is hardwired to zero, and sp/gp reset to the same values RARS uses so the traces line up.

Immediate generator

Produces the sign-extended 32-bit immediate for the I, S, B, U, and J formats, chosen by ImmSel, using only bit slicing and sign-bit replication.

32-bit ALU

Computes add/subtract, AND/OR/XOR, signed and unsigned compare, and left/right shift results in parallel, then selects one with ALUOp. Zero is a 32-input NOR of the result. The adder/subtractor is a ripple-carry chain of full adders.

Barrel shifter

A five-stage cascade of 2:1 multiplexers that shifts by 16, 8, 4, 2, and 1. SRL fills with zeros and SRA fills with the sign bit. SLL reuses the same right-shift structure by reversing the bits before and after the shift.

Load-extension unit

Handles LB/LH with sign extension, LBU/LHU with zero extension, and full-word LW.

Write-back

A four-way multiplexer that selects the ALU result, the extended load data, or PC+4 (the link address for JAL/JALR) to write to rd.

Supported instructions

Arithmetic & logic
ADDADDISUBANDANDIORORIXORXORISLTSLTISLTIU
Shifts
SLLSLLISRLSRLISRASRAI
Upper immediate
LUIAUIPC
Memory
LWSWLBLHLBULHU
Branches
BEQBNEBLTBGEBLTUBGEU
Jumps
JALJALR
System
WFI (halt)

Design decisions

  • Comparisons reuse the subtractor. Signed SLT uses sign(A−B) XOR overflow(A−B), which stays correct even when the subtraction overflows. Unsigned SLTU uses the inverted carry-out as a borrow. This avoids separate 32-bit comparators.
  • One ALU for branches and JALR. BEQ/BNE use the ALU's Zero output after subtraction, BLT/BGE/BLTU/BGEU use bit 0 of the set-less-than result, and the ALU also computes the JALR target. This avoids duplicating hardware.
  • Reusable structural blocks. A ripple-carry adder/subtractor, a mux-based barrel shifter, shared operand multiplexers, and a separate load-extension unit.

Verification

26 / 26
ALU testbench vectors passed
17 / 17
Control-unit vectors passed
10 / 10
Barrel-shifter vectors passed

The testbenches check themselves in QuestaSim, and each test program is compared against RARS's architectural trace:

ProgramWhat it exercisesResult
Proj1_base_test.sEvery required non-control-flow instruction, with later instructions using earlier results, plus sign and zero extension of bytes and halfwords0 mismatches (780 ns)
Proj1_cf_test.sAll six conditional branches, JAL, and JALR/RET, including a recursive routine five calls deep that saves and restores ra on the stack0 mismatches (1,300 ns)
Proj1_mergesort.sRecursive merge sort of a 12-element array (65, 12, 10, 89, 11, 70, 67, 5, 9, 45, 90, 7 → sorted)0 mismatches (39,260 ns)

The ALU tests deliberately include cases where signed and unsigned comparison disagree (0xFFFFFFFF vs 0x00000001), signed-overflow boundaries, shift amounts of 0 and 31, and negative operands for arithmetic right shifts.

Synthesis & critical path

22.69 MHz
Maximum frequency
≈ 44.07 ns
Worst-case clock period
−24.07 ns
Slack against the 20 ns (50 MHz) target

The critical path runs through almost the entire single-cycle datapath, as you would expect when every stage has to settle within one clock period:

PCInstruction memoryControl decodeRegister-file read muxALU operand muxALU / shifter selectData memoryWrite-back muxRegister-file input

The report identifies three ways to improve frequency: use FPGA block memories, reduce the depth of the register-file read multiplexers, and simplify the ALU and shifter multiplexing. Pipelining would help the most.

In later parts of the same course, the team used this processor as the baseline for five-stage software-scheduled and hardware-scheduled pipelined versions. Those are separate projects.

Tech

VHDL (structural / dataflow)RISC-V RV32IQuestaSimRARSRISC-V assemblyDigital logicTiming analysis