2. Organization & Implemention
2์ฅ Organization & Implementation ๊ฐ์
ARM ํ๋ก์ธ์์ Organization (๊ตฌ์ฑ) ๊ณผ Implementation (๊ตฌํ) ์ ํ์ตํ๋ค. 3-stage pipeline ARM๋ถํฐ 5-stage ARM9๊น์ง์ ์งํ, Principal Components, Data Forwarding, Branch, Memory Bottleneck์ ๋ค๋ฃฌ๋ค.
ARM์ Principal Components
- Register file: 32-bit 16๊ฐ์ ๋ฒ์ฉ ๋ ์ง์คํฐ. ์ผ๋ถ๋ PC(R15), LR(R14) ๋ฑ์ผ๋ก ์ฉ๋ ๊ณ ์
- Barrel shifter: ๋ ์ง์คํฐ ๊ฐ์ cycle ๋ด์์ shift/rotate (ARM์ ํน์ง)
- ALU: Arithmetic & Logic Unit
- Address register/incrementer: ๋ค์ fetch๋ฅผ ์ํด PC ์ฆ๊ฐ
- Data register: ๋ฉ๋ชจ๋ฆฌ์ ์ฃผ๊ณ ๋ฐ์ ๋ฐ์ดํฐ ๋ณด๊ด
- Instruction register: ํ์ฌ ์คํ ์ค์ธ instruction
3-Stage Pipeline ARM (ARM7 ์ด์ )
Fetch โ Decode โ Execute
๊ฐ stage๊ฐ ํ cycle์ฉ ๊ฑธ๋ฆฌ๊ณ CPI๋ 1์ ๊ฐ๊น๋ค. Program counter๋ ํญ์ ํ์ฌ ์คํ ์ค์ธ instruction + 8 (pipelining ๋๋ฌธ).
3 Stages ์ค๋ช
- Fetch: Instruction์ instruction register๋ก ๋ก๋ โ fetch ์๋ฃ
- Decode: Instruction์ ํด์, register ์ฝ๊ธฐ, control signal ์์ฑ
- Execute: ALU ์ฐ์ฐ, register์ ๊ฒฐ๊ณผ ์ ์ฅ

ARM Single-Cycle Instruction Operations
ARM์ single-cycle instruction ์ข ๋ฅ:
- Data operation: ADD, SUB, MUL ๋ฑ ALU ์ฐ์ฐ โ register์ ์ ์ฅ
- Addressing register R shifter: barrel shifter๋ฅผ ํตํด ์ฐ์ฐ (ARM ํน์ง)
- Store instruction: register ๊ฐ์ ๋ฉ๋ชจ๋ฆฌ๋ก (1 cycle)
STR (Store Instruction) โ Address Mode 4, ์๋ ์ฃผ์๋ก ์ฐ๊ธฐ
์: STR r0, [r1, #4] โ r1 + 4 ์ฃผ์์ r0 ์ ์ฅ
How to Improve the Performance
์ฑ๋ฅ ํฅ์์ ๊ธฐ๋ณธ ๊ณต์:
$$T_{prog} = N_{inst} \times CPI \times T_{clk}$$
- $N_{inst}$: instruction ์ (์ปดํ์ผ๋ฌ, ์๊ณ ๋ฆฌ์ฆ์ ์์กด)
- $CPI$: cycles per instruction (pipeline ์ค๊ณ)
- $T_{clk}$: clock ์ฃผ๊ธฐ
โ Pipeline stage ์๋ฅผ ๋๋ ค $T_{clk}$์ ์ค์ด๋ฉด, pipeline ํจ์จ์ ๋ฐ๋ผ ์ฑ๋ฅ ํฅ์ ๊ฐ๋ฅ. ๊ทธ๋ฌ๋ hazard ์ฆ๊ฐ.
ARM์ 3-stage โ 5-stage pipeline์ผ๋ก ํ์ฅ. ๋ ๋น ๋ฅธ clock๊ณผ ๋ ๋ง์ pipeline ๋จ๊ณ๋ก ์ต์ข ์ ์ผ๋ก RISC ๊ธฐ๋ฐ์ ํจ์จ์ pipeline์ ๊ตฌ์ฑ.

Memory Bottleneck (Von Neumann Bottleneck)
- Memory bottleneck์ ์ํ Structural Hazard ๋น๋ฒํ ๋ฐ์
- Instruction์ ์ฝ๋ ๊ฒ๊ณผ Data๋ฅผ ์ฝ๋ ๊ฒ์ด ๋์์ ์ํ๋ ์ ์๊ธฐ์ ํ๋๊ฐ ๋ฉ์ถฐ์ ์ํ โ Hazard ์ ๋ฐ
Solution
Memory bandwidth๋ฅผ ์ฆ๊ฐ์ํด: - ์ผ๋ฐ์ ์ธ ๋ฐฉ๋ฒ: Instruction๊ณผ Data Memory๋ฅผ ๋ถ๋ฆฌ (Harvard Architecture) - 1 cycle ๋น access๊ฐ 32-bit ์ด์์ ์ ๋ฌํด ์ฃผ๋ฉด ๋จ
5-Stage ARM Organization (ARM9)
Fetch โ Decode โ Execute โ Buffer/Data โ Write-Back
- Stage ์์ ์ฆ๊ฐ โ $f_{clk}$ ์ฆ๊ฐ
- Instruction๊ณผ data memory ๋ถ๋ฆฌ โ Memory bandwidth 2๋ฐฐ
- CPI ์ฆ๊ฐ (Stage์ ์๊ฐ ์ฆ๊ฐํ๊ธฐ์ ์ค์ฒฉํด ์ํํ ์ ์๋ Instruction ์ฆ๊ฐ)
- Hazard ๋ฐ์ ํ๋ฅ ์ ์ฆ๊ฐ์ด์ง๋ง ๋๋จธ์ง ์ด์ ์ด ๋ ํฌ๊ธฐ์ ์์ โ $f_{clk}$ ์ฆ๊ฐ, Memory bandwidth ์ฆ๊ฐ
5 Stages ์ค๋ช
- Fetch: I๋ฅผ ์ฝ์ด์์ pipeline์ ๋ฃ์
- Decode: I์ ์ข ๋ฅ๋ฅผ ์๊ณ ์ค๋น (register, Control Signal ์ค๋น)
- Execute: Operand shift, ALU ์ฐ์ฐ, ์ฃผ์ ์ฐ์ฐ
- Buffer/Data: Memory Access ํน์ by-pass๋จ
- Write-Back: ๊ฒฐ๊ณผ๊ฐ register์ ์ ์ฅ
Multiple Load/Store
์ฌ๋ฌ ๊ฐ์ I๋ฅผ ํ๋๋ก ๋์ฒดํ ์ ์๋ I ์ฌ์ฉ ๊ฐ๋ฅ โ MUX ๋ถ๋ถ์ด ์๋ ์ด์
LDR r1, 0(r0) ์ฐ์์ ์ธ ์ฆ๊ฐ
LDR r1, 4(r0) โ ํ๋๋ก ๋์ฒด ๊ฐ๋ฅ
LDR r1, 8(r0)
LDR r1, 12(r0)

Data Forwarding
5-stage pipeline์์๋ throughput โ, I๊ฐ ๋์์ ์ํ.
- Stall์ด ์๋ค๋ฉด data dependency ๊ด๊ณ๊ฐ ์์ (data Hazard)
- Data Forwarding์ผ๋ก ํด๊ฒฐ
reg bank โ Mux โ ALU โ D-cache
โ โ
โโโโโโโ internal forwarding
Mux๋ฅผ ํตํด ๋ฐ๋ก ๋์จ ๋ฐ์ดํฐ๋ฅผ ์ฌ์ฉํ ์ ์์.
Load I์ ๊ฒฝ์ฐ
๊ทธ๋ฌ๋ Load I์ internal forwarding์ ์ด์ฉํด๋ 1 stall์ด ํ์. โ Instruction reordering์ผ๋ก ํด๊ฒฐ ๊ฐ๋ฅ.
์:
LDR r0, 0(r1) LDR r0, 0(r1)
ADD r3, r0, r2 โ ADD r4, r5, r1
ADD r4, r5, r1 ADD r3, r0, r2
(1 stall ํ์) (stall ํ์ X)
Data Processing Instruction
๋ง์ , ๋บ์ , ๊ณฑ์ ๊ณผ ๊ฐ์ I์ ์ฒ๋ฆฌ ๊ณผ์ .
- 2๊ฐ์ operand ํ์
- ํ๋๋ ํญ์ register, ๋ค๋ฅธ ํ๋๋ register๋ immediate value ๊ฐ๋ฅ (8-bit๋ก ํํ ๊ฐ๋ฅ)
- barrel shifter๋ฅผ ํต๊ณผํจ (ARM ํน์ง)
ADD [reg] [reg/imm] (shift ์ต์
ํฌํจ)
- 2๊ฐ์ operand๋ ALU์์ ์ฐ์ฐํด์ register์ ์ ์ฅ (ARM ํน์ง)
- Hazard๊ฐ ์๋ค๋ฉด 1 cycle ๋์ ์ํ
์์
ADD r0, r1, r2 << 1 โ r2 LSL 1 (shift ์ด์ฉ)
- Data ์ฐ์ฐ:
data out์ผ๋ก ๋์ด - Decoder ์๋ต

Data Transfer Instruction
๋ฉ๋ชจ๋ฆฌ์ Register ์ฌ์ด ๋ฐ์ดํฐ ์ด๋ (Load/Store).
- Level 1 (Load I) = Single mem cycle: Stage 1 (1 cycle register access)
- ๋ฐ์ดํฐ ์ฝ๋ ๊ฒฝ์ฐ: register โ memory[address]
- ๋ฐ์ดํฐ ์ฐ๋ ๊ฒฝ์ฐ: memory[address] โ register
Auto-Index (Pre/Post Index)
- Pre-index:
LDR r1, [r2, #4]โ address = r2 + 4, ๋ฉ๋ชจ๋ฆฌ์์ ์ฝ์ (r2๋ ๊ทธ๋๋ก) - Pre-index with writeback:
LDR r1, [r2, #4]!โ address r2 + 4, r2 r2 + 4 (๊ฐฑ์ ) - Post-index:
LDR r1, [r2], #4โ address r2, ์ฝ์ ํ r2 r2 + 4
์ด๋ ARM์ ํน์ง์ธ load/store with auto-indexing์ ์ง์.
Single I โ Auto-index Calculation
ARM์ load/store ์ auto-indexing์ผ๋ก ์ฃผ์ ๊ณ์ฐ์ ํจ์จ์ ์ผ๋ก ์ฒ๋ฆฌ. ๋ค๋ฅธ ISA์์๋ load + add๊ฐ ํ์ํ์ง๋ง ARM์ ํ๋์ instruction์์ ์ฒ๋ฆฌ.
Branch Instruction
PC relative addressing, target address ๊ณ์ฐ.
- PC๋ฅผ ๊ธฐ์ค์ผ๋ก offset (signed) ๊ณ์ฐ โ target address
- Target address = PC + offset
๋จ์ Branch (B label): - 1st cycle: target ๊ณ์ฐ ๋ฐ PC ๊ฐฑ์
Branch with Link (BL label): - 1st cycle: branch target ๊ณ์ฐ - 2nd cycle: return address๋ฅผ R14 (LR) ์ ์ ์ฅ - 3rd cycle: LR ์์ (ํ์ฌ PC์ ๋ค์๋ฒ์ผ๋ก ๋์์์ผ ํ๊ธฐ์) โ R14 = R14 + 4

์ ๋ฆฌ
- Principal Components: Register file, Barrel shifter, ALU, Address/Data register, Instruction register
- 3-stage (ARM7) โ 5-stage (ARM9) ๋ก ์งํ, $f_{clk}$ ์ฆ๊ฐ + Harvard architecture๋ก memory bandwidth 2๋ฐฐ
- Data Forwarding: Mux๋ฅผ ํตํ internal forwarding์ผ๋ก ๋๋ถ๋ถ data hazard ํด๊ฒฐ, Load-use๋ 1 stall ๋๋ reordering
- Data Processing I: Barrel shifter ํต๊ณผ, 2 operand(ํ๋๋ immediate ๊ฐ๋ฅ), ALU ์ฐ์ฐ ํ register ์ ์ฅ
- Data Transfer I: Auto-indexing (pre/post)์ผ๋ก ์ฃผ์ ๊ณ์ฐ ํจ์จํ
- Branch: PC-relative, BL์ 3 cycle ํ์ (target + LR ์ ์ฅ + LR ์์ )
Pipelining์ ๊ธฐ๋ณธ ์์น๊ณผ ARM ํน์ ์ ์ต์ ํ (barrel shifter, auto-indexing, multiple load/store)๋ฅผ ๊ฒฐํฉํด RISC์ ์ ์๋ฅผ ๊ตฌํํ๋ค.
Comments (0)
No comments yet. Be the first to comment!
Please to write a comment.