SYSTEMS LECTURE WED SEP 9th

Fundamentals of LC-3 Assembly Language and Toolchain Workflow

  • Manual Conversion vs. Automated Assembly:

    • Transitioning from hand-converting instructions to hexadecimal machine code to using an automated LC-3 assembler software.

    • The manual conversion process is replaced entirely by writing source code in standard assembly language and allowing the assembler tool to calculate instruction fields, branch offsets, and absolute addresses.

  • Software Toolchain Architecture:

    • Text Editor: Assembly source code is written in plain text and saved with a .ASM file extension.

    • LC-3 Assembler: The .ASM file is opened and processed by the assembler application. The assembler checks syntax, resolves symbols, calculates relative addresses, and produces a compiled machine code file with a .HEX file extension.

    • LC-3 Simulator: The resulting .HEX file is loaded into the simulator software to execute, test, and debug the program.

  • Software Structure & Allowed Syntax Elements:

    • An LC-3 assembly program file must strictly contain only five permitted components:

    1. Assembly Instructions: Valid LC-3 mnemonics (ADD, LD, LEA, BR, etc.).

    2. Pseudo-Operations (Pseudo-Ops): Assembler directives starting with a period (.ORIG, .FILL, .BLKW, .STRINGZ, .END, .SUB).

    3. Labels: Symbolic names defining specific memory locations.

    4. Comments: Prefixed by a semicolon ; or a hash symbol # to document code.

    5. Whitespace/Tabbing: Structural formatting.

    • No arbitrary text or unapproved syntax may exist inside the .ASM file.

  • Variables in Memory vs. Processor Registers:

    • In high-level programming languages (such as Java), declaring a variable (e.g., int x;) allocates a dedicated memory location for x.

    • Executing an assignment like x = x + 1 requires the central processing unit (CPU) to:

    1. Read the current data value from the memory address associated with x.

    2. Load that value into an internal working register.

    3. Perform the addition operation within the register.

    4. Write the updated value from the register back out to memory location x.

    • While LC-3 assembly code can utilize registers (R0R_0 through R7R_7) as direct local storage due to its general-purpose register set, standard hardware architectures (such as x86 Windows processors) possess limited general-purpose registers.

    • Processors with restricted general-purpose registers must constantly transfer data back and forth to main memory. Because main memory access is slow, CPUs utilize high-speed internal cache memory.

    • Thrashing and Caching: To prevent performance degradation (thrashing) caused by frequent main memory bus reads and writes, data loaded from memory is temporarily held in CPU cache. The CPU interacts directly with cache memory and periodically writes modified cache lines back out to primary memory.

    • High-level language compilers targeting assembly ensure safety during conditional branches (e.g., if statements) by storing register contents to memory locations, executing comparison code, and subsequently reloading the original values back into registers.

Assembly Language Syntax Rules and Labels

  • Definition and Behavior of Labels:

    • A label is a human-readable identifier assigned to a specific memory location.

    • Labels do not exist in the final compiled .HEX machine code file. During compilation, the assembler translates every label into a calculated numeric address or offset.

    • Labels serve as the assembly language equivalent of high-level variables.

  • Label Naming Rules:

    • Must consist exclusively of uppercase letters, lowercase letters, numeric digits (00–99), and underscores (_).

    • No spaces, hyphens, or other punctuation characters are permitted.

    • Must begin with a letter or an underscore (_). Starting a label name with a numeric digit is strictly invalid.

    • Cannot match reserved LC-3 instruction mnemonics or pseudo-ops.

  • Column Alignment and Error Detection:

    • Standard assemblers require labels to be aligned strictly in the leftmost column (Column 1) of a line.

    • Indenting non-label code via tabs or spaces allows strict assemblers to detect misspelled instruction mnemonics. For example, if a developer accidentally types JSTR hello indented, a strict assembler checks Column 1, sees no label, evaluates JSTR as an instruction, recognizes it as invalid, and flags a syntax error.

    • Lenient assemblers (such as standard LC-3 implementations) allow labels to appear anywhere on a line. Under lenient rules, typing an invalid instruction like JSTR hello indented causes the assembler to misinterpret JSTR as a valid label name, allowing compilation to proceed without reporting an error.

  • Instruction Compatibility with Labels:

    • Labels can only be used alongside instructions that utilize Program Counter (PC) relative offsets. These include:

    • Memory Load Instructions: LD, LDI, LEA

    • Memory Store Instructions: ST, STI

    • Control Flow Instructions: BR (all conditional branch variants), JSR

    • Instructions that do not use PC-relative offsets cannot accept labels directly:

    • Register-based Load/Store: LDR, STR (these require a base register and a 6-bit offset).

    • Arithmetic/Logical Instructions: ADD, AND, NOT.

    • Common Architectural Syntax Error:

    • Attempting to pass labels directly to arithmetic operations (e.g., ADD R0, A, B) is invalid because ADD hardware encoding does not support PC offsets.

    • Immediate fields for ADD and AND instructions are strictly constrained to 5-bit signed integers ranging from −16-16 to 1515 ([−16,15][-16, 15]).

    • To perform addition on values stored at labeled memory locations A and B, the data must first be explicitly loaded into registers using PC-relative load instructions:       LD R0,A\text{LD } R_0, A       LD R1,B\text{LD } R_1, B       ADD R0,R0,R1\text{ADD } R_0, R_0, R_1

Pseudo-Operations (Pseudo-Ops)

  • .ORIG (Origin):

    • Syntax: .ORIG address (e.g., .ORIG x3000 or .ORIG 3000).

    • Specifies the starting absolute memory address where the compiled machine code will be loaded into LC-3 memory.

    • The origin address is written as the very first 16-bit word inside the generated .HEX output file.

  • .FILL:

    • Syntax: .FILL value (e.g., .FILL x0041 or .FILL 65).

    • Allocates a single 16-bit memory location and initializes its contents to the specified numeric constant or character code.

    • Converts a decimal, hexadecimal, or character representation directly into a 4-digit hexadecimal word in the output file.

  • .BLKW (Block of Words):

    • Syntax: .BLKW N (e.g., .BLKW 5).

    • Allocates a continuous block of NN sequential 16-bit memory locations.

    • Functions identically to declaring a fixed-size array in high-level languages.

    • The LC-3 assembler initializes all NN allocated memory locations to zero (x0000).

  • .STRINGZ:

    • Syntax: .STRINGZ "string_literal" (e.g., .STRINGZ "Hello").

    • Allocates sequential memory locations and stores each ASCII character of the string literal into consecutive memory words.

    • Automatically appends a trailing null-terminator character (x0000) into the final allocated memory location.

  • .END:

    • Syntax: .END

    • Marks the physical end of the assembly source file. Any text appearing below .END is ignored by the assembler.

  • .SUB (Subroutine Declaration):

    • Syntax: .SUB subroutine_label

    • A custom assembler directive required at the top of the assembly file (directly below .ORIG) to declare every subroutine defined within the program.

    • Informs the assembler of subroutine entry points so it can enforce strict modular control flow rules.

Internal Mechanics of a Two-Pass Assembler

  • Location Counter (LCLC):

    • An internal variable maintained by the assembler software during compilation to track the current target memory address.

    • The LCLC is the translation-time equivalent of the execution-time Program Counter (PCPC).

  • Pass 1: Symbol Table Construction:

    1. The assembler reads the source code from top to bottom.

    2. Upon encountering .ORIG N, the assembler initializes the Location Counter to the base address NN (e.g., LC=3000LC = 3000).

    3. As lines are scanned, if a line contains a label, the assembler creates a new entry in its internal Symbol Table mapping the label string to the current value of LCLC:      Symbol Table Mapping: Label Name→Memory Address (LC)\text{Symbol Table Mapping: } \text{Label Name} \rightarrow \text{Memory Address } (LC)

    4. For every line processed, the assembler increments LCLC based on the memory footprint of the directive or instruction:

    • Standard instructions (ADD, LD, LEA, BR, HALT, etc.): LC=LC+1LC = LC + 1

    • .FILL: LC=LC+1LC = LC + 1

    • .BLKW N: LC=LC+NLC = LC + N

    • .STRINGZ "text": LC=LC+(Length of Text)+1LC = LC + (\text{Length of Text}) + 1

    • Comments and blank lines: LC=LC+0LC = LC + 0 (Location Counter remains unchanged).

  • Pass 2: Machine Code Output Generation:

    1. The assembler resets LCLC back to the origin address defined by .ORIG (e.g., LC=3000LC = 3000).

    2. The source file is scanned line-by-line a second time to translate instructions into 16-bit binary/hexadecimal machine code.

    3. When an instruction references a label, the assembler looks up the label's address in the Symbol Table generated during Pass 1.

    4. The assembler automatically calculates the required PC offset using the formula:      PC Offset=Target Address (from Symbol Table)−(LC+1)\text{PC Offset} = \text{Target Address (from Symbol Table)} - (LC + 1)

    5. The calculated offset is encoded directly into the instruction bit fields of the binary machine code.

  • Detailed Assembly Memory Mapping Example:   Consider the following assembly program segment:

  .ORIG x3000
  LEA R6, C
  LD R0, A
  HALT
  A .FILL x0041
  B .FILL x0042
  C .BLKW 5
  D .FILL x0044
  .END
  ```
  - **Pass 1 Symbol Table Generation Tracking**:
    - `.ORIG x3000` →\rightarrow Set initial LC=3000LC = 3000
    - `LEA R6, C` →\rightarrow Takes 1 word. LCLC increments from 30003000 to 30013001
    - `LD R0, A` →\rightarrow Takes 1 word. LCLC increments from 30013001 to 30023002
    - `HALT` →\rightarrow Encodes opcode `xF025`. Takes 1 word. LCLC increments from 30023002 to 30033003
    - `A .FILL x0041` →\rightarrow Label `A` assigned address 30033003. Takes 1 word. LCLC increments from 30033003 to 30043004
    - `B .FILL x0042` →\rightarrow Label `B` assigned address 30043004. Takes 1 word. LCLC increments from 30043004 to 30053005
    - `C .BLKW 5` →\rightarrow Label `C` assigned address 30053005. Allocates 5 words. LCLC increments from 30053005 to 30083008
    - `D .FILL x0044` →\rightarrow Label `D` assigned address 30083008. Takes 1 word. LCLC increments from 30083008 to 30093009
  - **Final Symbol Table Output**:
    - `A` →\rightarrow `3003`
    - `B` →\rightarrow `3004`
    - `C` →\rightarrow `3005`
    - `D` →\rightarrow `3008`
  - **Pass 2 Translation Calculation**:
    - At line `LEA R6, C` (located at memory address 30003000), the updated location counter during execution is PC=3000+1=3001\text{PC} = 3000 + 1 = 3001.
    - Target address of `C` is 30053005.
    - Offset calculation: Offset=3005−3001=4\text{Offset} = 3005 - 3001 = 4
    - Machine instruction output: `xE604`

# Arrays, Pointer Arithmetic, and Memory Errors

- **Label Stacking Error without Data Directives**:
  - A critical student error occurs when defining consecutive labels without allocating memory between them:
    ```assembly
    HALT
    A
    B
    C
    D
    ```
  - Because labels `A`, `B`, `C`, and `D` do not contain data pseudo-ops (`.FILL` or `.BLKW`), the Location Counter LCLC does not increment.
  - Consequently, labels `A`, `B`, `C`, and `D` all resolve to the **exact same memory address** (the address immediately following `HALT`). Loading or storing to any of these labels alters the identical memory location.
  - Correct implementation requires explicit data allocation:
    ```assembly
    HALT
    A .FILL 5
    B .FILL 7
    C .FILL 17
    D .FILL 0
    ```

- **Array Allocation and Indexing Operations**:
  - Declaring `C .BLKW 5` allocates 5 continuous memory locations starting at address 30053005 (spanning 30053005 to 30093009).
  - The label `C` represents a reference holding the base address of the array (address of index 00).
  - Executing high-level array assignment like `C[2] = 5` in LC-3 assembly requires explicit pointer arithmetic:
    1. Load the base address of array `C` into a register using Load Effective Address:
       LEA R0,C(R0←3005)\text{LEA } R_0, C \quad (R_0 \leftarrow 3005)
    2. Add the literal index offset 22 to the base address:
       ADD R0,R0,#2(R0←3005+2=3007)\text{ADD } R_0, R_0, \#2 \quad (R_0 \leftarrow 3005 + 2 = 3007)
    3. Load the source value 55 into a data register (R1R_1):
       AND R1,R1,#0\text{AND } R_1, R_1, \#0
       ADD R1,R1,#5\text{ADD } R_1, R_1, \#5
    4. Store the value into the calculated address using Base+Offset store (`STR`):
       STR R1,R0,#0\text{STR } R_1, R_0, \#0

- **Buffer Overflows and Absence of Hardware Bounds Checking**:
  - High-level managed languages (like Java) enforce strict runtime array bounds checking, raising exceptions if an index exceeds array boundaries.
  - Low-level languages (like C) and LC-3 assembly hardware provide zero memory protection or bounds checking.
  - If code calculates an out-of-bounds address (e.g., applying offset 55 to array `C`, accessing address 3005+5=30103005 + 5 = 3010), the LC-3 CPU executes the memory operation without warning.
  - Writing out-of-bounds silently overwrites adjacent program data or code instructions. Memory corruption errors often manifest far downstream during execution, causing unpredictable program crashes that are difficult to trace.

# Subroutines and Structural Control Flow Enforcements

- **Subroutine Execution Model**:
  - Subroutines are modular code routines invoked via `JSR label` or `JSRR Rbase`.
  - Calling `JSR` automatically saves the return address (location of the instruction immediately following `JSR`) into link register R7R_7.
  - Subroutines complete execution and return control to the caller by executing the return instruction `RET` (which is an alias for `JMP R7`).

- **Subroutine Discipline and Branching Violations**:
  - Code inside a subroutine must never use conditional branch instructions (`BR`) to jump out of the subroutine into main code or into a separate subroutine.
  - Branching across subroutine boundaries breaks return address handling and violates standard function stack frames.
  - Control must enter exclusively via `JSR`/`JSRR` and exit exclusively via `RET`.

- **`.SUB` Pseudo-Op Rules and Custom Compiler Verification**:
  - The custom assembler enforces subroutine boundary rules using the `.SUB` directive.
  - Declarations for all subroutines must be listed at the top of the file directly beneath `.ORIG`:
    ```assembly
    .ORIG x3000
    .SUB MULP
    .SUB DIV
    ```
  - Each subroutine defined in the file must start with its assigned label and contain exactly **one** return instruction (`RET`) marking its boundary end.
  - The assembler parses code within the scope between the subroutine label and its terminating `RET`. If a branch instruction (`BR`) attempts to jump to a target memory address located before the subroutine entry or after the `RET` line, the assembler throws a compilation error and halts build generation.

- **Multi-Subroutine File Template**:

assembly .ORIG x3000 .SUB MULP .SUB DIV

; --- MAIN PROGRAM --- LD R0, VAL1 LD R1, VAL2 JSR MULP HALT

VAL1 .FILL #5 VAL2 .FILL #3

; --- SUBROUTINE: MULP --- MULP ; (Multiplication code logic here) RET

; --- SUBROUTINE: DIV --- DIV ; (Division code logic here) RET

.END   ```

Verification and Attendance

  • Attendance Record:

    • Braden Bear

    • Alex Kiner

    • Dylan Linker

    • CJ Willigan