Recently, I’m trying out different HDL generator languages and tools, because
Chisel is used heavily inside SiFive, and they have developed amount of IPs including very complicated CPUs, and a very well maintained …
Cache coherence between CPUs are most explained in textbooks, but IO coherence is not well understood. Recently I’m involved in architecture discussion about IO coherence, and found this paper, “Maintaining I/O Data …
FANG# Frame# Problem statement & background
Assumption# Based on best info
Non-goals# Something NOT trying to solve
Goals# Something trying to solve
Avoid# Too long of framing, not history lession Facts are not …
From Self-evaluation and improvement in engineering leadership
How manage a remote/distributed teams?
Communication is the key. TODO: setup regular communication channel. Knowing the ppl you are working with, the methods …
From Cloning yourself isn’t an option by Camille Fournier
Everyone wants to have clones to help them with certain work. But additive is linear improvements while multiplier is trying to achieve exponential improvements …
List of training content# [RISC-V Architecture Training] Schedule [RISC-V Architecture Training] Introduction of RISC-V Open ISA [RISC-V Architecture Training] Basics & Unprivileged Specification [RISC-V Architecture …
–
Uncore# CPU core is fun, but uncore is the real work.# Uncore / components# Cache (already discussed) Interrupt controller Network Fabric Debug Interrupt recap# 3 types of interrupts
External: peripheral devices …
Momentum: 2018 RISC-V Summit# Fun moment: anti-RISCV website by ARM# Schedule# 2-day x 8-hour# Step-by-step# Lecture + demo + DIY# Schedule / Day 1 morning# Schedule and self-introduction Introduction of RISC-V open ISA …
Privileged architecture# Purpose of privileged architecture# To manage and protect shared resources
Memory, IO devices, even cores Also needs to decouple implementation details
Handle unimplemented operations: software …
What is ISA?# Contract between software and hardware.# What is RISC?# Reduced instruction set computer# Small set of simple/general instructions + load/store architecture Optimize hardware to be simple and faster …
RISC-V SPEC# https://riscv.org/specifications (official version v1.10)
https://github.com/riscv/riscv-isa-manual (source code)
User-level ISA (unpriviledged)# All the basic instructions, and extensions Memory model …
AMBA (ARM Advanced Microcontroller Bus Architecture)
1. AXI# AXI protocol is a point-to-point protocol So no matter what the network channels really use, as long as its ports comply AXI protocol, IP can be connected to …
ARM online training note
1. Introduction# What is an architecture?# Instruction set Exception model Memory model Debug ARMv8# AArch32 vs AArch64 AArch32: backward compatible to ARMv7 AArch64: fixed 32-bit instruction, …
Simple Sequential Execution Model# After optimization, the result should be exactly the same with “simple sequential execution model”.
Optimization: Instruction Fetching# Fetch multiple instructions from memory Branch …
Reference
Interrupt Categorization# Hardware vs. Software Hardware: usually caused by peripheral or other processors IRQ: maskable interrupt NMI: non-maskable interrupt For highest priority tasks, like times, especially …
Typical usage# go to some directory use short alias# > pwd /home/jw > go prj0 cd /work/projects/design/master-branch/ > pwd /work/projects/design/master-branch/run a serial of commands use short alias# > run …
There is a big difference between how I used to understand hardware security and state-of-the-art security supported by hardware software co-design, after I watched some video talking about SEP (Security Enclave …
https://developer.arm.com/technologies/big-little
big.LITTLE is a practical example of SMP (Symmetric Multiprocessing). It combines high performance CPU cores and low power CPU cores in the same chip, connected using …
Because I was job hunting recently, I’ve got lots of appointments, either phone call or face-to-face. There are two HR’s who gave me particularly deep impression, and they are on totally opposite end of professionalism. …
What Are My Strengths?# Feedback analysis# The only way to discover your strengths is through feedback analysis. Whenever you make a key decision or take a key action, write down what you expect will happen. 9 or 12 …
Intro# Device tree: for non-discoverable hardware, included in BSP Source type Old style: C code BSP, files compiled into the kernel New style: device-tree BSP -> device tree blob (load by boot loader) Compilation# …
Trying to find the perfect static site generator. Used to use Pelican, because it’s written in Python. Also tried with Jekyll, the most popular candidate, because it’s used by Github. Their common problems are
Not …
Keynote panel on RISC-V Summit 2018: opportunities and challenges in security for open source hardware# Complex systems tend to have bugs, so making it preparatory will make it more secure from attacks. But open source …
From the reading of this paper, “The Hwacha Microarchitecture Manual, Version 3.8.1”, I found out that our Pygmy ES1 architecture is almost the same idea, just not as fancy.
We don’t have cache coherency, because we …
Vector regfile# 32 of them, v0 to v31 Each is VLEN bits Each can be divided into several elements The max element width is ELEN CSR vsew maps to SEW (standard element width) controls their width dynamically CSR vl …
My notes on RISC-V Summit 2018 at Santa Clara Conventional Center# This year’s summit has many more participants than the last one, which means RISC-V is getting a lot of momentum around the world. Although most of the …
NoC# Clustering coefficient: the most intuitive explanation is the number of hops between two random nodes in the network. Layers Physical layer Link layer Transaction protocol: such as AXI Seperated channels like AXI, …
Ariane Document
Architecture note# PC gen stage# The fetching address for i-cache is always word-aligned. Fetch stage# Its fetch stage doesn’t have much decoding work to do, only the necessary one to generate next PC. …
A piece of very precious memory
Home –> LEGOLAND# 264 miles, 4 hours non-stop (with stop, 7:00AM to 12:00AM)
Lunch at Wendy’s: 5821 Dennis McCarthy Dr, Lebec, CA 93243
Wendy’s is just another burger place. We ended …
Register renaming# To eliminate the false and output data dependency by adding extra physical registers more than architectural registers.
Read-after-write (RAW) is true data dependency Write-after-write (WAW) is output …
Coherence mechanism# Snooping# Every cache maintain its own cache state. And when it needs to access a shared address space, it sends snooping messages to all the other caches to either update or invalidate them.
Write …
1. 晚上睡前戴眼镜片# 在使用“阿托品”之后至少两小时,以保证药物被吸收。 洗手,并且用厨房纸(不掉纸屑)擦干。 滴眼药水,润湿眼球。 灰色眼镜片是右眼的,GRAY has an R for RIGHT;蓝色眼镜片是左眼的,BLUE has an L for LEFT。 将眼药水注满眼镜片内部;低头,拉开上下眼帘;将眼镜慢慢地水平地放入。 确认眼镜片在眼球正中,否则可以闭眼后,用手指在眼帘外轻推调整。 确认没有大的可见气泡,否则会影响 …
When using a counter to divide a clock, don’t reset the counter, especially when you are using synchronous reset. It will make the clock quiet while reset. And if it’s used along with sync reset, then those flip-flop …
Price of GCP# Persistance Disk# Can be used to put all the data/eda/os on it.
Price (per month) Price (per GB per month) SSD 50GB $8.50 $0.17 SSD 1TB $174.08 $0.17 HDD 50GB $2.00 $0.04 HDD 200GB $8.00 $0.04 HDD 1TB …
If your design needs to switch from one clock source to another, there is high possibility of harmful clock glitches while switching. Normally you need to stop this clock during the switching process, but what if you …
The course is on Coursera
understanding design patterns# what? well-known solutions for recurring problems why? don’t reinvent wheels reuse best practices characteristics language neutral dynamic incomplete by …
Scala introduction course on LinkedIn. Not very useful, if not using it in real project.
introduction# short for Scalable language object-oriented + functional programming everything is object including numbers and …
The following is my notes of GENUS training course on Cadence’s training module
Module 03: genus fundamentals# common UI vs legacy mode# unified commands with Tempus common us: set_db & get_db legacy mode: …
Day 1 Morning# Paper from microsoft Microsoft has a IC design team? Apparently it does. Grey code: even when metastability happends, it falls to adjecent states, instead of unknow states, it’s acceptable in some cases. …
To start gnome-terminal on WSL (Windows Subsystem for Linux)# After upgrade to Windows 10 Creators Update, reinstall WSL will have Ubuntu 16.04.2 LTS on Windows.
To reinstall WSL you should do:
> lxrun /uninstall …
In general, things like Anaconda Server are designed to make this sort of workflow easier.
Some suggested workarounds:
Reproduce your install on another machine with internet (save conda list –export to a file and conda …
Question: how to control the clock skew between a group of clocks to be minimum, say less than 30ps, instead of utilizing useful skew? This case happens to our hard macros.
A: in Innovus, use skew group
set …
The following is my notes of INNOVUS training course on Cadence’s training module
Module 02: overview# “gift” directory contains lots of useful scripts to help productivity Independent “viewlog” utility or …
AI and ML# Artificial intellegence vs human intellegence The imitation game, eugen Goosman passed the Turing Test, 2014 Alpha Go, deepmind 2015
Introduction to deep learning# Improve on task T with respect to performance …
Start-up: python -> enterprise: C/Java/Scala, more engineers, faster Research: quick result and prototyping
GPU? Data movement between GPU and CPU is important
[ ] fast.ai: class (high school math)
infrastructure: …
By Jon Shlens and George Toderici from Google Research @ 2017-01-20 Fri
History
Convolutional NN: old tech, why suddenly it works?
Scale: 60M parameters At least 60M +1 data point to fit these parameters
SIMD hardware …
Some parameterized example RTL code for register-based SRAM read circuit using “generate” feature
parameter d = 32; // FIFO depth parameter w = 64; // FIFO data bit-width logic [w-1:0] mem [d-1:0]; // FIFO memory array …
This is my reading note of book “SystemVerilog for Design (2nd edition)". As a non-full-time RTL designer, it has opened my mind. But still, I’m sad about the antient tool that we are using to design …
As shown in the schematic, we have some clock divider that divide root clock by half. While in scan mode, these flip-flops will be bypassed and treated as normal flip-flop that need to be inserted into the scan chain …
Recent consolidation progress is going so crazy, mostly because of the IC industry is becoming less and less profitable in every individual application fields. My perspective is that this is not just a consolidation of …
More organizations are starting to adopt a remote work culture. But how do managers stay in sync with what their teams are doing when they can’t see them? While it’s important to define clear goals early on, you should …
Intel IoT platform Intel IoT Platform Sensors and things Arduino Gateway transfer data between different types of networks some data processing as well moving quickly: different types of data flexibility is important …
Bohr: moore’s law for 50 years
$mm^2$ is increasing since 130nm heterogeneous intergration: 3D chip is not a replacement of moore’s law (quite opposite with Sehat) refocus on general purpose design Hill: 21 century …
This was my second time to read this book. You cannot imagine how shock I was when I first read this book on Kindle. So I again bought a hardcopy and want to read it again.
The topic about “present” was the first chapter …
Background# This is the summary of my experience from project LBRAM in Marvell in the year of 2014.
The first thing: discuss timing/area/power SPEC’s in details# Most of the time, because custom design takes lots of …
Let’s consider a face recognition system, where we’ve got facial images from a list of known persons and the system input is a camera image. We need to figure out if there are people in this camera image from …
#software #ai #llm #open-source
How to train a transformer based model, like LLaMA2, from scratch? Andrej Karpathy has open-sourced llama2.c project on GitHub. My learning process is “duplicate and rewrite”, …
#software #ai #llm #open-source
I was involved in the early ExecuTorch definition phase and had used its predecessor Lite Interpreter extensively in work. I really like this idea and its design. This is a great effort …
#software #ai #llm #open-source
https://github.com/jimwang99/understanding-llama2/tree/main/pytorch
Above GitHub repo is an implementation of LLaMA2 and test-case use TinyStories in PyTorch
Example output
------ …
#software #ai #llm #open-source
Following up with Understanding LLaMA2 Part 1 Model Architecture, this diagram explains LLaMA model architecture with KV Cache support. We follow the same legend as well as the …
#software #ai #llm #open-source
Here I’m capturing the details of llama’s model architecture use PlantUML component diagram, using the following 2 GitHub repos as references …
DFS (depth first search)
DFS can be done easily with recursive method, because you can treat the subtrees as new trees and use the same function to traverse them.
BFS (breadth first search)
BFS will need extra space of a …
Mark all nodes as not visited. Create an empty ready bucket to hold nodes that has been visited itself but not its neighbors yet. Find a starting node, mark it as visited and put into the ready bucket. Get an node from …
After working on enabling use-cases for edge AI accelerator for more than 2 years now, here is some of my thinking about system performance.
Thoughts: Much More Beyond Model# One of the biggest challenges we’ve …
“Software is king”# It’s a hard statement to make, as a veteran hardware engineer. Lots of pride and ego to swallow. However, it’s truly based on my observations in the industry. Allow me to …
#software #accelerator #ai #on-device #WIP
In this part of the tutorial, we will learn how to run a pretrained model from TorchVision: ResNeXt50, which is a model architecture built upon the concepts of ResNet. …
#software #accelerator #ai #on-device
SNPE = Snapdragon Neural Processing Engine
In this tutorial we assume that Qualcomm SNPE has been successfully installed use QPM. Follow “Qualcomm Package Manager 1.0” …
Attributes# “Min heap”, where index 0 is the smallest item APIs# heapq.heapify(iterable) -> None: Create a heap queue in-place heapq.heappush(heap, item) -> None: Add a new item heapq.heappop(heap) …
1. Have a Vision: Without a clear vision outlining the final goals, how can we measure our progress? And it’s essential to clearly communicate this vision with all team members.
2. Break Down into Tasks and …
TL;DR
Nvidia’s DIGITS offers an on-premise AI solution aimed at smaller organizations that require strict data privacy. While it is cost-effective for small businesses and college research labs, it may be less suitable …
Here is my notes from JPMorgan’s “Eye on Market” 2024 April issue
The emergence and integration of large language models (LLMs) into various professional sectors have marked a significant milestone in …
Research projects often come with a lot of unknowns. Many engineers find this unsettling because it’s less straightforward than math or digital realm. However, I think we should welcome these uncertainties. They reflect …
All the credit goes to ByteByteGo.com
Do you believe that Google, Meta, Uber, and Airbnb put almost all of their code in one repository?
This practice is called a monorepo.
Monorepo vs. Microrepo. Which is the best? Why …
To optimize the efficiency of training or executing an ML model, whether implemented locally on a device or hosted in the cloud, parallelization plays a critical role, akin to other computational challenges.
Utilizing …
LLM is a memory bound problem. This inspired me to look at different memory technologies. In this article, I’m going to summarize my research these days, especially about HBM and its impact on AI applications. …
Modern SoCs heavily relies on NoC to connect interfaces and storage to compute. As the ML models grow larger and larger, the data delivery ability becomes more and more important to overall system performance.
While …
DeepSeek has created a huge wave of discussion and panic in the market.
I think it’s great overall.
Innovation From a technology standpoint, DeepSeek is truly pushing the envelope in both training methods and model …
Background# i7-12700K = Intel Core i7-12700K (8 big cores each has 2 threads, 4 little cores each has 1 thread) running at 5GHz rpi4 = Raspberry Pi 4 Rev B, with 4x Cortex-A72 running at 1.8GHz (Broadcom BCM2711) rpi5 = …
#hardware #accelerator #chisel #open-source
I’ve tried to learn Chisel 5 years ago, but gave up and went back to SystemVerilog to design our RISC-V CPU + AI custom instructions from scratch. After 5 years, both …
Taking my home lab Ceph distributed storage system on Raspberry Pi to the next level: make it useful by creating a block device interface so that Linux system can mount it and use it.
Enable client# Install Ceph package …
Trying to follow this tutorial to create a Ceph distributed storage system in my home lab, because I’m sick of the NFS performance of my Synology NAS. …
#note #learn-with-chatgpt #algorithm
NOTE: this is a note from ChatGPT
The Boyer-Moore Voting Algorithm, also known simply as the Voting Algorithm, is a method for finding a majority element in a sequence of elements. …
https://github.com/jimwang99/jimon
While experimenting my little distributed storage system built with Raspberry Pi SBCs and Ceph, some Raspberry Pi machines die now and then. My suspicion is heat, because I built them …
In this Github project, I created a simple application that can do text to image and image to image search, using open-source transformer model. Details can be found in the repo and its docs directory.
Here is a screen …