Computer Organization and Assembly Language - Basic Elements and Instructions

Assembly Language Basic Elements and Statement Syntax

ASSEMBLY language programs are composed of statements that follow a distinct syntax. Every statement consists of four potential fields, separated by at least one blank or tab character. Every statement is categorized as either an Instruction (translated into machine code) or an Assembler Directive (a pseudo-operation that instructs the assembler to perform tasks like memory allocation or procedure creation).

General Syntax

name operation operand(s) comments

  • Name Field:

    • The name field is used by the assembler to translate names into memory addresses.

    • Names can be between 11 and 3131 characters long.

    • Permitted characters include letters, digits, and special characters.

    • If a period (.) is used, it must be the first character.

    • Embedded blanks (spaces) are strictly forbidden.

    • Names cannot begin with a digit.

    • Assembly language is not case-sensitive regarding name fields.

    • Legal Examples: COUNTER_1, @character, .TEST, DONE?.

    • Illegal Examples: TWO WORDS (contains embedded blank), 2abc (begins with a digit), A45.28 (period is not the first character), YOU&ME (contains invalid character).

  • Operation Field:

    • In instructions, this contains a Symbolic Operation Code (Op code) which is translated into a machine language op code (e.g., ADD, MOV, SUB).

    • In assembler directives, this field represents a Pseudo-op code, which is not translated into machine code but guides the assembler (e.g., PROC is used to define a procedure).

  • Operand Field:

    • An instruction may have zero, one, or two operands.

    • In a two-operand instruction, the first operand is the destination and the second operand is the source.

    • For directives, this field provides specific information required by the directive.

    • Examples:

      • NOP : Zero operands; performs no operation.

      • INC AX : One operand; increments the contents of the AX register by 11.

      • ADD AX, 2 : Two operands; adds the value 22 to the contents of AX.

  • Comments Field:

    • Optional data marked by a semicolon (;).

    • Ignored completely by the assembler, but considered good practice for documentation.

Program Data and Representation

Processors operate exclusively on binary data, but assembly language allows programmers to express data in several formats.

Radix and Number Systems

To select a specific number system, a radix symbol (suffix) is used:

  • Binary: Suffix b (e.g., 1011b).

  • Octal: Suffix o (e.g., 32o).

  • Decimal: Suffix d (e.g., 35d); this is the default radix if none is specified.

  • Hexadecimal: Suffix h (e.g., 6A15h).

    • Leading Zero Requirement: Hexadecimal constants must begin with a decimal digit (0−90-9). For example, ABCh must be written as 0ABCh.

    • Constants cannot contain non-digit characters like commas (e.g., 1,234 is illegal).

Character Representation
  • Characters are enclosed in single or double quotes.

  • They are stored as ASCII codes. There is no functional difference between specifying "A" or 41h (the hex ASCII value for 'A').

Variables and Data Definition Directives

Each variable is assigned a memory address and a specific data type. Variables can hold numeric values, string constants, constant expressions, or be left uninitialized using the ? symbol.

Data Definition Syntax

variable_name type initial_value
variable_name type value1, value2, value3 (for multiple values/arrays)

Pseudo-ops for Data Definition

Pseudo-op

Description

Size (Bytes)

Examples

DB

Define Byte

11

var1 DB 'A', array1 DB 10, 20, 30

DW

Define Word

22

var2 DW 1234h, array2 DW 1000, 2000

DD

Define Double Word

44

Var3 DD -214743648

DQ

Define Quad Word

88

Used for larger precision

DT

Define Ten Bytes

1010

Used for extended precision

Note: If a value like 10h is stored in a DW (Word), it is saved in memory as 0010h to fill the 2-byte2\text{-byte} space.

Variable Ranges
  • 8-bit8\text{-bit} Number Range:

    • Signed: −128-128 to 127127

    • Unsigned: 00 to 255255

  • 16-bit16\text{-bit} Number Range:

    • Signed: −32,768-32,768 to 32,76732,767

    • Unsigned: 00 to 65,53565,535

Arrays and String Storage

Arrays

An array is a sequence of memory bytes or words.

  • Example 1 (Byte Array): B_ARRAY DB 10h, 20h, 30h

    • If B_ARRAY starts at offset 0200h:

    • Address 0200h: 10h

    • Address 0201h: 20h

    • Address 0202h: 30h

  • Example 2 (Word Array): W_ARRAY DW 1000, 40, 29887, 329

    • If W_ARRAY starts at offset 0300h:

    • Address 0300h: 1000d

    • Address 0302h: 40d (2 bytes away due to word size2 \text{ bytes away due to word size})

    • Address 0304h: 29887d

High and Low Bytes of a Word

When using WORD1 DW 1234h:

  • Low Byte: 34h, accessed via symbolic address WORD1.

  • High Byte: 12h, accessed via symbolic address WORD1 + 1.

Character Strings
  • LETTERS DB 'ABC' is functionally identical to LETTERS DB 41h, 42h, 43h.

  • UpperCase and LowerCase are differentiated by the assembler.

  • Characters and numbers can be combined in one definition: MSG DB 'HELLO', 0Ah, 0Dh, '$' is equivalent to MSG DB 48h, 45h, 4Ch, 4Ch, 4Fh, 0Ah, 0Dh, 24h.

Named Constants

Named constants use the EQU (Equate) directive to assign a symbolic name to a numeric value or string.

  • Syntax: name EQU constant

  • Example: LF EQU 0Ah (Line Feed constant).

  • Memory Impact: No memory is allocated for named constants; the assembler simply replaces the name with the value during translation.

Essential Data Transfer and Arithmetic Instructions

MOV (Move)

Used to transfer data between registers, between a register and memory, or to move a constant directly to a register or memory location.

  • Syntax: MOV destination, source

  • Legal Combinations:

    • General Register to General Register: YES

    • Memory to General Register: YES

    • Constant to General Register: YES

    • General Register to Memory: YES

    • Register to Segment Register / Segment Register to Memory: YES

    • Memory to Memory: NO (requires an intermediate register).

XCHG (Exchange)

Exchanges the contents of two registers, or a register and a memory location.

  • Syntax: XCHG destination, source

  • Legal Combinations:

    • Register to Register: YES

    • Register to Memory/Memory to Register: YES

    • Memory to Memory: NO

Arithmetic: ADD, SUB, INC, DEC, NEG
  • ADD/SUB: Adds or subtracts source from destination.

    • Legal: Reg-Reg, Memory-Reg, Reg-Memory, Reg-Constant, Memory-Constant.

    • Illegal: Memory-to-Memory (e.g., ADD BYTE1, BYTE2 is illegal; use MOV AL, BYTE2 followed by ADD BYTE1, AL).

  • INC/DEC: Increment or Decrement a register or memory location by 11.

    • Operates on 8-bit8\text{-bit} or 16-bit16\text{-bit} values.

  • NEG: Replaces the contents of the destination with its 2’s complement (negation).

    • Example: If BX = 0002, NEG BX results in FFFE.

Translation of High-Level Language (HLL) to Assembly

HLL Statement

Assembly Translation (Example)

B = A

MOV AX, A
MOV B, AX

A = 5 - A

MOV AX, 5
SUB AX, A
MOV A, AX
OR
NEG A
ADD A, 5

A = B - 2 x A

MOV AX, B
SUB AX, A
SUB AX, A
MOV A, AX

Note: Translation solutions are not unique and must account for whether variables are defined as bytes or words.

Program Segment Structure and Memory Models

Programs operate across three primary segments: Code, Data, and Stack. The assembler translates these segments into memory segments.

Memory Models

Determined by the .MODEL directive:

  • SMALL: One code segment, one data segment.

  • MEDIUM: More than one code segment, one data segment.

  • COMPACT: One code segment, more than one data segment.

  • LARGE: Multiple code and data segments; no array larger than 64KB64KB.

  • HUGE: Multiple code and data segments; arrays may exceed 64KB64KB.

Segment Declarations
  • Data Segment: .DATA contains all variable definitions.

  • Stack Segment: .STACK size specifies the stack area in bytes. If omitted, default is 1KB1KB (1024 bytes1024 \text{ bytes}). Example: .STACK 100h.

  • Code Segment: .CODE [name] contains the instructions. Name is optional and typically omitted in SMALL model.

  • ORG 0100h: Directive used to set the origin of the program code.

Practical Implementation: Adding Two Numbers

The following logic is used in an emu8086 environment using the int 21h interrupt for basic I/O.

I/O Logic and ASCII Adjustment
  • Input (ah=01h): Accepts a character. The ASCII code is stored in AL. Since entering '2' results in ASCII 5050, you must subtract 4848 (or 30h30h) to get the actual numeric value 22.

  • Output (ah=02h): Prints a character. The data to be printed must be moved to DL. To print a numeric result, add 4848 (or 30h30h) to convert it back to its ASCII character representation.

  • Control Characters:

    • 10 (Line Feed): Moves to a new line.

    • 13 (Carriage Return): Moves cursor to the start of the line.

Detailed Program Flow
  1. Initialize: Define .MODEL SMALL, .STACK 100h, and .DATA.

  2. Input 1: Use ah=01h and int 21h. Subtract 4848 from AL. Store result in BL.

  3. New Line: Move 10 to DL, use ah=02h/int 21h. Move 13 to DL, repeat.

  4. Input 2: Repeat input process. Subtract 4848 from AL.

  5. Addition: ADD BL, AL.

  6. Convert to Character: ADD BL, 48 to restore ASCII for display.

  7. Output Sum: Move BL to DL, use ah=02h/int 21h to print.

  8. Terminate: Use MAIN ENDP and END MAIN.