Processor Fundamentals
Pull back the cover on a computer's actual thinking machinery: the Von Neumann model every modern processor is built on, the registers and buses that move a single instruction through the CPU, the Fetch-Execute cycle that repeats billions of times a second, and the assembly language and bit-level operations that let you write instructions the processor can execute directly. This chapter covers Cambridge 9618 syllabus sections 4.1 to 4.3.
Every program you have ever run, no matter what language it was written in, is eventually reduced to the same thing: a sequence of simple instructions moving through a processor one at a time. This chapter covers Cambridge 9618 syllabus section 4.1, the Von Neumann architecture, registers, buses and the Fetch-Execute cycle that make that happen; section 4.2, assembly language, the low-level instructions that map directly onto machine code; and section 4.3, bit manipulation, the shifting and masking operations that let software control individual bits.
The Von Neumann Model
Almost every general-purpose computer built since the 1940s, from a phone to a supercomputer, follows the same basic blueprint, named after the mathematician John von Neumann. Understanding this one model unlocks the entire rest of this chapter, because the registers, buses and Fetch-Execute cycle you are about to meet all exist purely to make it work.
- Instructions and data live in the same memory, and look identical, just binary patterns
- The CPU has no way to tell, just by looking at a memory location, whether it holds an instruction or data, context is everything
- Instructions are normally executed sequentially, one after another, unless a jump instruction changes that order
- Software can be loaded, changed and replaced without touching hardware
- One general-purpose machine can run any program, rather than needing to be rebuilt for each task
- It is the single idea that makes reprogrammable, general-purpose computing possible at all
The CU, the ALU and the System Clock
Inside the processor, three components do the actual work of executing a von Neumann style program: one component fetches and directs, one component calculates, and one component keeps everything moving in step.
- Fetches instructions from memory and decodes what they mean
- Generates the control signals that direct every other component, telling the ALU when to calculate, memory when to read or write, and registers when to update
- Manages the entire Fetch-Execute cycle, timed against the system clock
- Performs all arithmetic operations: addition, subtraction, multiplication, division
- Performs all logical operations: AND, OR, NOT and comparisons
- Results of every ALU operation are stored back into the Accumulator (ACC)
Special Purpose Registers
Registers are tiny, extremely fast storage locations built directly into the CPU itself, far faster than even cache memory, because there is no bus transfer involved in accessing them. General purpose registers can be used flexibly by a program for whatever it needs; special purpose registers each have one fixed, dedicated job in making the Fetch-Execute cycle work.
| Register | Full name | Purpose |
|---|---|---|
| PC | Program Counter | Holds the address of the next instruction to be fetched |
| MAR | Memory Address Register | Holds the address currently being read from or written to in memory |
| MDR | Memory Data Register | Holds data just read from memory, or about to be written to it |
| ACC | Accumulator | General purpose register that holds the results of ALU operations |
| CIR | Current Instruction Register | Holds the instruction currently being decoded and executed |
| IX | Index Register | Used in indexed addressing, its value is added to a base address |
| Status | Status (Flag) Register | Individual bits flag conditions: overflow, carry, zero, negative, interrupt |
System Buses
Registers and memory are physically separate, so something has to carry signals between them. A bus is a set of parallel wires that carries related signals together, and the CPU uses three, each with a completely different job.
| Bus | Direction | Carries | Effect of width |
|---|---|---|---|
| Address bus | CPU → memory (unidirectional) | The memory address to be read from or written to | More bits means more distinct addresses, so more memory can be addressed |
| Data bus | Bidirectional | The actual data value being transferred | A wider bus moves more bits in a single transfer, improving performance |
| Control bus | Bidirectional | Control signals: read/write, clock timing, interrupt requests | More lines allow more distinct control signals to be sent |
Ports for Peripheral Devices
Peripheral devices, printers, monitors, keyboards, external drives, connect to a computer through physical ports, each following a standard that defines the physical connector and the rules for the data flowing through it.
Factors Affecting Performance
"How fast is this computer" is never answered by a single number. Four separate hardware factors each contribute, and exam questions expect you to explain the mechanism behind each one, not just name it.
| Factor | How it improves performance |
|---|---|
| Clock speed | Measured in GHz (billions of cycles per second). A higher clock speed means more clock cycles occur every second, so more instructions can potentially be executed in the same amount of time. |
| Number of cores | Each core is an independent processing unit capable of executing its own instruction stream. Multiple cores let a CPU genuinely execute several threads simultaneously (true parallelism), rather than just switching rapidly between them. |
| Bus width | A wider data bus transfers more bits in a single operation, so more data moves between the CPU and memory per clock cycle, reducing the number of transfers a task needs. |
| Cache memory | A small amount of very fast SRAM sitting between the CPU and main RAM, organised in levels (L1 fastest and smallest, then L2, then L3, largest and slowest of the three). It stores recently or frequently used data, so the CPU can retrieve it without the far longer delay of accessing main memory. |
The Fetch-Execute Cycle
Everything covered so far, the registers, the buses, the CU and ALU, exists to make one repeating process happen: fetch an instruction, work out what it means, then carry it out. This cycle repeats, billions of times a second, for as long as the computer runs.
- MAR ← [PC], the address to fetch is copied from PC into MAR
- MDR ← [[MAR]], the instruction at that memory address is copied into MDR
- PC ← [PC] + 1, PC is incremented so it now points to the next instruction
- CIR ← [MDR], the fetched instruction moves from MDR into CIR
- The CU decodes the opcode and operand now sitting in CIR
- The CU generates the control signals needed to carry the instruction out
- For example, an ADD instruction sends signals to fetch the operand's value and add it to ACC via the ALU
- Once execution completes, the cycle loops back to fetch, using the new value of PC
Interrupts
A processor cannot simply stop mid-instruction whenever something urgent happens elsewhere in the system, but it also cannot ignore urgent events indefinitely. An interrupt is a signal sent to the CPU requesting that it pause its current task to deal with a higher-priority event, and the CPU checks for pending interrupts at one fixed, predictable point: the end of every Fetch-Execute cycle.
| Category | Examples |
|---|---|
| Hardware interrupt | A keyboard keypress, a mouse click, a printer signalling it has finished, network data arriving |
| Software interrupt | An error such as division by zero, or a program deliberately requesting a service from the operating system |
| Timer interrupt | A regular signal used for process scheduling, letting the operating system switch between running processes |
How an interrupt is handled
Once the Interrupt Service Routine (ISR) finishes running, the CPU restores the saved state (PC and registers) from the stack exactly as it was, and execution resumes from precisely where it left off, as though the interruption had never happened.
Assembly Language & the Two-Pass Assembler
Assembly language is a low-level language with a direct, one-to-one relationship to machine code: every single assembly instruction corresponds to exactly one machine code instruction. This is completely different from a high-level language, where one line of code (a loop, a function call) can expand into dozens of machine code instructions. Assembly uses short mnemonics like ADD or JMP in place of raw binary opcodes purely to make the code readable by humans.
Why a two-pass assembler is needed
An assembler is the program that translates assembly language source code into machine code. The difficulty is that a program frequently jumps forward to a label that has not been defined yet at the point the jump instruction is written, a forward reference, so the assembler cannot always resolve every address on a single read-through.
| Pass | What it does |
|---|---|
| Pass 1 | Reads through the entire program without generating any machine code yet. Its only job is to build a symbol table that maps every label name used in the program to the actual memory address it will occupy. |
| Pass 2 | Reads through the program a second time, this time translating every instruction into its binary machine code equivalent. Whenever an instruction references a label, the assembler looks up its address in the symbol table built during pass 1, resolving the reference correctly no matter which direction it points. |
JMP loop_end written near the top of a program, where loop_end: is defined several lines further down. On a single pass, the assembler would reach the jump instruction before it has ever seen the label loop_end, and would have no address to put in its place. Pass 1 solves this by finding and recording every label's address first, before pass 2 ever needs to use it.
Instructions are grouped by purpose
The full instruction set (covered in detail in the next section) is organised into five functional groups, and exam questions sometimes ask you to classify a given instruction into its correct group.
- LDM, LDD, LDI, LDX, LDR, MOV, STO
- IN, OUT
- ADD, SUB, INC, DEC
- JMP (unconditional); JPE, JPN (conditional, following a compare); CMP, CMI (compare)
Addressing Modes & the Instruction Set
An instruction like LDD 200 and an instruction like LDM 200 look almost identical, but they do completely different things. The addressing mode determines how the operand written after the opcode should actually be interpreted: is it a value, an address, or something more indirect?
| Mode | Operand means | Example | Typical use |
|---|---|---|---|
| Immediate | The value itself, not an address at all | LDM #42 → ACC = 42 | Loading a known constant |
| Direct | The address in memory to read from | LDD 200 → ACC = Memory[200] | Accessing a variable directly |
| Indirect | The given address holds another address, and it is that second address which is read | LDI 200 → ACC = Memory[Memory[200]] | Pointer dereferencing |
| Indexed | The given address plus the current contents of IX | LDX 200 → ACC = Memory[200 + IX] | Stepping through an array inside a loop |
| Relative | An offset counted from the current value of PC | JMP +3 → skip forward 3 instructions | Jumps that stay correct even if the whole program is relocated |
The full instruction set
| Opcode | Operand | Operation |
|---|---|---|
| LDM | #n | Load the immediate value n into ACC |
| LDD | <address> | Load the contents of the given memory address into ACC |
| LDI | <address> | Indirect: load the contents of the address held at the given address into ACC |
| LDX | <address> | Indexed: load the contents of (address + IX) into ACC |
| LDR | #n | Load the immediate value n into IX |
| MOV | <register> | Move the contents of ACC into the given register (e.g. IX) |
| STO | <address> | Store the contents of ACC at the given memory address |
| ADD | <address> / #n | Add the value at the address, or the immediate value n, to ACC |
| SUB | <address> / #n | Subtract the value at the address, or the immediate value n, from ACC |
| INC | ACC / IX | Add 1 to the named register |
| DEC | ACC / IX | Subtract 1 from the named register |
| JMP | <address> | Unconditional jump to the given address |
| CMP | <address> / #n | Compare ACC with the value at the address, or with n, setting the flags |
| CMI | <address> | Indirect compare: compare ACC with the contents of the address held at the given address |
| JPE | <address> | Following a compare, jump to the address only if the comparison was true (equal) |
| JPN | <address> | Following a compare, jump to the address only if the comparison was false (not equal) |
| IN | — | Read a character from the keyboard and store its ASCII value in ACC |
| OUT | — | Output to the screen the character whose ASCII value is stored in ACC |
| END | — | Return control to the operating system |
#n denotes an immediate denary (base 10) number, e.g. #123. A leading B denotes a binary number, e.g. B01001010. A leading & denotes a hexadecimal number, e.g. &4A. An <address> can be written as an absolute number or as a symbolic label, and exam papers assume only one general purpose register, the Accumulator, is available.
This program adds two numbers stored at labelled memory locations x and y, and stores the result.
LDM #0 ; ACC = 0 ADD x ; ACC = ACC + memory[x] ADD y ; ACC = ACC + memory[y] STO result ; memory[result] = ACC END x: 5 ; data declaration y: 3 result: 0
<label>: <opcode> <operand> labels an instruction (so another instruction can jump to it), while <label>: <data> simply gives a symbolic name to a memory location holding a fixed data value, exactly like x:, y: and result: above.
Bit Manipulation
So far every operation has worked on a whole byte or word at once. Sometimes, especially when controlling a device directly, a program needs to work on individual bits within a register, and that is exactly what shifts and bit masking are for.
Binary shifts
| Shift type | Description | Left shift effect | Right shift effect |
|---|---|---|---|
| Logical | All bits move; zeros fill the vacated positions; any bit shifted off the end is lost | ×2 (approximately, if no significant bit is lost) | ÷2, treating the value as unsigned |
| Arithmetic | Behaves like a logical shift, except a right shift preserves the sign bit (the MSB) so the value's sign is not corrupted | ×2 | ÷2, correctly preserving a negative sign |
| Cyclic | Bits shifted off one end wrap around and re-enter at the other end, no bit is ever lost | Rotate left | Rotate right |
Logical left shift by 2 (LSL #2): before 00001101 = 13, after 00110100 = 52 (13 × 4, since each left shift roughly doubles the value).
Logical right shift by 1 (LSR #1): before 00101100 = 44, after 00010110 = 22 (44 ÷ 2).
Bit masking
Bit masking uses a second byte, the mask, together with a logical operation to manipulate specific bits of a target byte while leaving every other bit completely unchanged.
| Operation | Mask bit | Effect on target bit | Purpose |
|---|---|---|---|
| AND with 0 | 0 | Forces the bit to 0 | Clear specific bits |
| AND with 1 | 1 | Preserves the bit unchanged | Test whether a bit is set |
| OR with 1 | 1 | Forces the bit to 1 | Set specific bits |
| XOR with 1 | 1 | Flips the bit to its opposite | Toggle specific bits |
; TEST if bit 3 (counting from 0 on the right) is set in ACC AND B00001000 ; mask keeps only bit 3, clears every other bit CMP #0 JPE bit_clear ; result was 0, so bit 3 was 0 ; SET bit 5 in ACC OR B00100000 ; forces bit 5 to 1, leaves every other bit unchanged ; CLEAR bit 2 in ACC AND B11111011 ; mask has 0 only at bit 2, so only bit 2 is forced to 0
Practice Questions
The stored program concept holds both a program's instructions and the data it works on together, as binary values, in the same main memory, using the same address space. This is important because it means a computer can run a completely different program simply by loading different instructions and data into memory, with no changes to the hardware required, which is what makes general-purpose, reprogrammable computers possible.
MAR (Memory Address Register) holds the address in memory currently being accessed. MDR (Memory Data Register) holds the actual data that has been read from, or is about to be written to, that address.
MAR ← [PC], the address in PC is copied into MAR. MDR ← [[MAR]], the contents of the memory location addressed by MAR are copied into MDR. PC ← [PC] + 1, PC is incremented to point at the next instruction. CIR ← [MDR], the fetched instruction is copied from MDR into CIR ready for decoding.
Cache is very fast SRAM situated between the CPU and main memory. A larger cache can hold more recently or frequently used data and instructions, increasing the chance that the CPU finds what it needs already in cache (a cache hit) rather than having to fetch it from the much slower main memory, reducing the average time spent waiting for data.
AND the register with the mask B00010000. This clears every bit except bit 4, and the result can then be compared with 0: a non-zero result confirms bit 4 was set, while leaving the original register unaffected since the operation is performed on a copy or checked without storing back.
A forward jump references a label that appears later in the program than the jump instruction itself, so on a single read-through the assembler would reach the jump before it has seen where that label is defined, leaving it with no address to translate the instruction with. Pass 1 solves this by scanning the whole program first and recording every label's address in a symbol table, so that pass 2 can look up and correctly resolve every reference, including forward ones, while generating the machine code.
