Architecture

Browse the Architecture posts.

IO Coherence

Cache coherence between CPUs are most explained in textbooks, but IO coherence is not well understood. Recently I’m involved in architecture discussion about IO coherence, and found this paper, “Maintaining I/O Data …

Dec. 2019 verion Full List

List of training content# [RISC-V Architecture Training] Schedule [RISC-V Architecture Training] Introduction of RISC-V Open ISA [RISC-V Architecture Training] Basics & Unprivileged Specification [RISC-V Architecture …

Uncore

– Uncore# CPU core is fun, but uncore is the real work.# Uncore / components# Cache (already discussed) Interrupt controller Network Fabric Debug Interrupt recap# 3 types of interrupts External: peripheral devices …

Schedule

Momentum: 2018 RISC-V Summit# Fun moment: anti-RISCV website by ARM# Schedule# 2-day x 8-hour# Step-by-step# Lecture + demo + DIY# Schedule / Day 1 morning# Schedule and self-introduction Introduction of RISC-V open ISA …

Privileged Architecture

Privileged architecture# Purpose of privileged architecture# To manage and protect shared resources Memory, IO devices, even cores Also needs to decouple implementation details Handle unimplemented operations: software …

Introduction of RISC-V Open ISA

What is ISA?# Contract between software and hardware.# What is RISC?# Reduced instruction set computer# Small set of simple/general instructions + load/store architecture Optimize hardware to be simple and faster …

Computer Architecture with RISC-V Examples

Computer architecture basics# Pipeline / Parallelism / Cache# Three ultimate mechanisms to imporve performance/power Computer architecture basics / pipeline# IF = Instruction Fetch, ID = Instruction Decode, EX = Execute, …

Basics & Unprivileged Specification

RISC-V SPEC# https://riscv.org/specifications (official version v1.10) https://github.com/riscv/riscv-isa-manual (source code) User-level ISA (unpriviledged)# All the basic instructions, and extensions Memory model …

ARM AMBA Protocl

AMBA (ARM Advanced Microcontroller Bus Architecture) 1. AXI# AXI protocol is a point-to-point protocol So no matter what the network channels really use, as long as its ports comply AXI protocol, IP can be connected to …

ARMv8 Architecture

ARM online training note 1. Introduction# What is an architecture?# Instruction set Exception model Memory model Debug ARMv8# AArch32 vs AArch64 AArch32: backward compatible to ARMv7 AArch64: fixed 32-bit instruction, …

ARM Training Cortex Processor Behaviors

Simple Sequential Execution Model# After optimization, the result should be exactly the same with “simple sequential execution model”. Optimization: Instruction Fetching# Fetch multiple instructions from memory Branch …

Interrupts and ARM GIC Architecture

Reference Interrupt Categorization# Hardware vs. Software Hardware: usually caused by peripheral or other processors IRQ: maskable interrupt NMI: non-maskable interrupt For highest priority tasks, like times, especially …

Re-discover Hardware Security in Modern SoC

There is a big difference between how I used to understand hardware security and state-of-the-art security supported by hardware software co-design, after I watched some video talking about SEP (Security Enclave …

ARM's big.LITTLE Architecture

https://developer.arm.com/technologies/big-little big.LITTLE is a practical example of SMP (Symmetric Multiprocessing). It combines high performance CPU cores and low power CPU cores in the same chip, connected using …

Huwcha Accelerator architecture

From the reading of this paper, “The Hwacha Microarchitecture Manual, Version 3.8.1”, I found out that our Pygmy ES1 architecture is almost the same idea, just not as fancy. We don’t have cache coherency, because we …

Note of RISC-V Vector ISA Spec v0.6

Vector regfile# 32 of them, v0 to v31 Each is VLEN bits Each can be divided into several elements The max element width is ELEN CSR vsew maps to SEW (standard element width) controls their width dynamically CSR vl …

Network-on-Chip Notes

NoC# Clustering coefficient: the most intuitive explanation is the number of hops between two random nodes in the network. Layers Physical layer Link layer Transaction protocol: such as AXI Seperated channels like AXI, …

Ariane (PULP series high-performance core)

Ariane Document Architecture note# PC gen stage# The fetching address for i-cache is always word-aligned. Fetch stage# Its fetch stage doesn’t have much decoding work to do, only the necessary one to generate next PC. …

CPU Architecture Notes

Register renaming# To eliminate the false and output data dependency by adding extra physical registers more than architectural registers. Read-after-write (RAW) is true data dependency Write-after-write (WAW) is output …

Cache Coherence Notes

Coherence mechanism# Snooping# Every cache maintain its own cache state. And when it needs to access a shared address space, it sends snooping messages to all the other caches to either update or invalidate them. Write …

Posts without dates

How to Evaluate NoC (Network-on-Chip)?

Modern SoCs heavily relies on NoC to connect interfaces and storage to compute. As the ML models grow larger and larger, the data delivery ability becomes more and more important to overall system performance. While …

CPU Performance Test

Background# i7-12700K = Intel Core i7-12700K (8 big cores each has 2 threads, 4 little cores each has 1 thread) running at 5GHz rpi4 = Raspberry Pi 4 Rev B, with 4x Cortex-A72 running at 1.8GHz (Broadcom BCM2711) rpi5 = …