Cache coherence between CPUs are most explained in textbooks, but IO coherence is not well understood. Recently I’m involved in architecture discussion about IO coherence, and found this paper, “Maintaining I/O Data …
List of training content# [RISC-V Architecture Training] Schedule [RISC-V Architecture Training] Introduction of RISC-V Open ISA [RISC-V Architecture Training] Basics & Unprivileged Specification [RISC-V Architecture …
–
Uncore# CPU core is fun, but uncore is the real work.# Uncore / components# Cache (already discussed) Interrupt controller Network Fabric Debug Interrupt recap# 3 types of interrupts
External: peripheral devices …
Momentum: 2018 RISC-V Summit# Fun moment: anti-RISCV website by ARM# Schedule# 2-day x 8-hour# Step-by-step# Lecture + demo + DIY# Schedule / Day 1 morning# Schedule and self-introduction Introduction of RISC-V open ISA …
Privileged architecture# Purpose of privileged architecture# To manage and protect shared resources
Memory, IO devices, even cores Also needs to decouple implementation details
Handle unimplemented operations: software …
What is ISA?# Contract between software and hardware.# What is RISC?# Reduced instruction set computer# Small set of simple/general instructions + load/store architecture Optimize hardware to be simple and faster …
RISC-V SPEC# https://riscv.org/specifications (official version v1.10)
https://github.com/riscv/riscv-isa-manual (source code)
User-level ISA (unpriviledged)# All the basic instructions, and extensions Memory model …
AMBA (ARM Advanced Microcontroller Bus Architecture)
1. AXI# AXI protocol is a point-to-point protocol So no matter what the network channels really use, as long as its ports comply AXI protocol, IP can be connected to …
ARM online training note
1. Introduction# What is an architecture?# Instruction set Exception model Memory model Debug ARMv8# AArch32 vs AArch64 AArch32: backward compatible to ARMv7 AArch64: fixed 32-bit instruction, …
Simple Sequential Execution Model# After optimization, the result should be exactly the same with “simple sequential execution model”.
Optimization: Instruction Fetching# Fetch multiple instructions from memory Branch …
Reference
Interrupt Categorization# Hardware vs. Software Hardware: usually caused by peripheral or other processors IRQ: maskable interrupt NMI: non-maskable interrupt For highest priority tasks, like times, especially …
There is a big difference between how I used to understand hardware security and state-of-the-art security supported by hardware software co-design, after I watched some video talking about SEP (Security Enclave …
https://developer.arm.com/technologies/big-little
big.LITTLE is a practical example of SMP (Symmetric Multiprocessing). It combines high performance CPU cores and low power CPU cores in the same chip, connected using …
From the reading of this paper, “The Hwacha Microarchitecture Manual, Version 3.8.1”, I found out that our Pygmy ES1 architecture is almost the same idea, just not as fancy.
We don’t have cache coherency, because we …
Vector regfile# 32 of them, v0 to v31 Each is VLEN bits Each can be divided into several elements The max element width is ELEN CSR vsew maps to SEW (standard element width) controls their width dynamically CSR vl …
NoC# Clustering coefficient: the most intuitive explanation is the number of hops between two random nodes in the network. Layers Physical layer Link layer Transaction protocol: such as AXI Seperated channels like AXI, …
Ariane Document
Architecture note# PC gen stage# The fetching address for i-cache is always word-aligned. Fetch stage# Its fetch stage doesn’t have much decoding work to do, only the necessary one to generate next PC. …
Register renaming# To eliminate the false and output data dependency by adding extra physical registers more than architectural registers.
Read-after-write (RAW) is true data dependency Write-after-write (WAW) is output …
Coherence mechanism# Snooping# Every cache maintain its own cache state. And when it needs to access a shared address space, it sends snooping messages to all the other caches to either update or invalidate them.
Write …
Modern SoCs heavily relies on NoC to connect interfaces and storage to compute. As the ML models grow larger and larger, the data delivery ability becomes more and more important to overall system performance.
While …
Background# i7-12700K = Intel Core i7-12700K (8 big cores each has 2 threads, 4 little cores each has 1 thread) running at 5GHz rpi4 = Raspberry Pi 4 Rev B, with 4x Cortex-A72 running at 1.8GHz (Broadcom BCM2711) rpi5 = …