Posts

All posts are listed below. Jump to posts without dates, or choose a topic from the sidebar.

Status of current HDL generator

Recently, I’m trying out different HDL generator languages and tools, because Chisel is used heavily inside SiFive, and they have developed amount of IPs including very complicated CPUs, and a very well maintained …

IO Coherence

Cache coherence between CPUs are most explained in textbooks, but IO coherence is not well understood. Recently I’m involved in architecture discussion about IO coherence, and found this paper, “Maintaining I/O Data …

FANG for decision making

FANG# Frame# Problem statement & background Assumption# Based on best info Non-goals# Something NOT trying to solve Goals# Something trying to solve Avoid# Too long of framing, not history lession Facts are not …

Improve in engineering leadership

From Self-evaluation and improvement in engineering leadership How manage a remote/distributed teams? Communication is the key. TODO: setup regular communication channel. Knowing the ppl you are working with, the methods …

How to become a multiplier

From Cloning yourself isn’t an option by Camille Fournier Everyone wants to have clones to help them with certain work. But additive is linear improvements while multiplier is trying to achieve exponential improvements …

Dec. 2019 verion Full List

List of training content# [RISC-V Architecture Training] Schedule [RISC-V Architecture Training] Introduction of RISC-V Open ISA [RISC-V Architecture Training] Basics & Unprivileged Specification [RISC-V Architecture …

Uncore

– Uncore# CPU core is fun, but uncore is the real work.# Uncore / components# Cache (already discussed) Interrupt controller Network Fabric Debug Interrupt recap# 3 types of interrupts External: peripheral devices …

Schedule

Momentum: 2018 RISC-V Summit# Fun moment: anti-RISCV website by ARM# Schedule# 2-day x 8-hour# Step-by-step# Lecture + demo + DIY# Schedule / Day 1 morning# Schedule and self-introduction Introduction of RISC-V open ISA …

Privileged Architecture

Privileged architecture# Purpose of privileged architecture# To manage and protect shared resources Memory, IO devices, even cores Also needs to decouple implementation details Handle unimplemented operations: software …

Introduction of RISC-V Open ISA

What is ISA?# Contract between software and hardware.# What is RISC?# Reduced instruction set computer# Small set of simple/general instructions + load/store architecture Optimize hardware to be simple and faster …

Computer Architecture with RISC-V Examples

Computer architecture basics# Pipeline / Parallelism / Cache# Three ultimate mechanisms to imporve performance/power Computer architecture basics / pipeline# IF = Instruction Fetch, ID = Instruction Decode, EX = Execute, …

Basics & Unprivileged Specification

RISC-V SPEC# https://riscv.org/specifications (official version v1.10) https://github.com/riscv/riscv-isa-manual (source code) User-level ISA (unpriviledged)# All the basic instructions, and extensions Memory model …

ARM AMBA Protocl

AMBA (ARM Advanced Microcontroller Bus Architecture) 1. AXI# AXI protocol is a point-to-point protocol So no matter what the network channels really use, as long as its ports comply AXI protocol, IP can be connected to …

ARMv8 Architecture

ARM online training note 1. Introduction# What is an architecture?# Instruction set Exception model Memory model Debug ARMv8# AArch32 vs AArch64 AArch32: backward compatible to ARMv7 AArch64: fixed 32-bit instruction, …

ARM Training Cortex Processor Behaviors

Simple Sequential Execution Model# After optimization, the result should be exactly the same with “simple sequential execution model”. Optimization: Instruction Fetching# Fetch multiple instructions from memory Branch …

Interrupts and ARM GIC Architecture

Reference Interrupt Categorization# Hardware vs. Software Hardware: usually caused by peripheral or other processors IRQ: maskable interrupt NMI: non-maskable interrupt For highest priority tasks, like times, especially …

shortcut.py A Command Line Script to Provide Macro

Typical usage# go to some directory use short alias# > pwd /home/jw > go prj0 cd /work/projects/design/master-branch/ > pwd /work/projects/design/master-branch/run a serial of commands use short alias# > run …

Re-discover Hardware Security in Modern SoC

There is a big difference between how I used to understand hardware security and state-of-the-art security supported by hardware software co-design, after I watched some video talking about SEP (Security Enclave …

ARM's big.LITTLE Architecture

https://developer.arm.com/technologies/big-little big.LITTLE is a practical example of SMP (Symmetric Multiprocessing). It combines high performance CPU cores and low power CPU cores in the same chip, connected using …

On Time Is Professional

Because I was job hunting recently, I’ve got lots of appointments, either phone call or face-to-face. There are two HR’s who gave me particularly deep impression, and they are on totally opposite end of professionalism. …

HBR's 10 Must Read Managing Oneself Note

What Are My Strengths?# Feedback analysis# The only way to discover your strengths is through feedback analysis. Whenever you make a key decision or take a key action, write down what you expect will happen. 9 or 12 …

Working with Device Tree (DOULOS)

Intro# Device tree: for non-discoverable hardware, included in BSP Source type Old style: C code BSP, files compiled into the kernel New style: device-tree BSP -> device tree blob (load by boot loader) Compilation# …

Static Site Generator

Trying to find the perfect static site generator. Used to use Pelican, because it’s written in Python. Also tried with Jekyll, the most popular candidate, because it’s used by Github. Their common problems are Not …

Hardware Security

Keynote panel on RISC-V Summit 2018: opportunities and challenges in security for open source hardware# Complex systems tend to have bugs, so making it preparatory will make it more secure from attacks. But open source …

Huwcha Accelerator architecture

From the reading of this paper, “The Hwacha Microarchitecture Manual, Version 3.8.1”, I found out that our Pygmy ES1 architecture is almost the same idea, just not as fancy. We don’t have cache coherency, because we …

Note of RISC-V Vector ISA Spec v0.6

Vector regfile# 32 of them, v0 to v31 Each is VLEN bits Each can be divided into several elements The max element width is ELEN CSR vsew maps to SEW (standard element width) controls their width dynamically CSR vl …

SystemC Tutorial

// Some simple example #include <systemc.h> SC_MODULE (seq_and2 ) { // sequential AND2 sc_in< sc_uint<8> > a; sc_in< sc_unit<8> > b; sc_out< sc_uint<8> > f; sc_in<bool> …

RISC-V Summit 2018

My notes on RISC-V Summit 2018 at Santa Clara Conventional Center# This year’s summit has many more participants than the last one, which means RISC-V is getting a lot of momentum around the world. Although most of the …

Network-on-Chip Notes

NoC# Clustering coefficient: the most intuitive explanation is the number of hops between two random nodes in the network. Layers Physical layer Link layer Transaction protocol: such as AXI Seperated channels like AXI, …

Ariane (PULP series high-performance core)

Ariane Document Architecture note# PC gen stage# The fetching address for i-cache is always word-aligned. Fetch stage# Its fetch stage doesn’t have much decoding work to do, only the necessary one to generate next PC. …

Family trip to LEGOLAND

A piece of very precious memory Home –> LEGOLAND# 264 miles, 4 hours non-stop (with stop, 7:00AM to 12:00AM) Lunch at Wendy’s: 5821 Dennis McCarthy Dr, Lebec, CA 93243 Wendy’s is just another burger place. We ended …

CPU Architecture Notes

Register renaming# To eliminate the false and output data dependency by adding extra physical registers more than architectural registers. Read-after-write (RAW) is true data dependency Write-after-write (WAW) is output …

Cache Coherence Notes

Coherence mechanism# Snooping# Every cache maintain its own cache state. And when it needs to access a shared address space, it sends snooping messages to all the other caches to either update or invalidate them. Write …

OK镜使用和保养方法

1. 晚上睡前戴眼镜片# 在使用“阿托品”之后至少两小时,以保证药物被吸收。 洗手,并且用厨房纸(不掉纸屑)擦干。 滴眼药水,润湿眼球。 灰色眼镜片是右眼的,GRAY has an R for RIGHT;蓝色眼镜片是左眼的,BLUE has an L for LEFT。 将眼药水注满眼镜片内部;低头,拉开上下眼帘;将眼镜慢慢地水平地放入。 确认眼镜片在眼球正中,否则可以闭眼后,用手指在眼帘外轻推调整。 确认没有大的可见气泡,否则会影响 …

Case Study Clock Divider with Synchronous Reset

When using a counter to divide a clock, don’t reset the counter, especially when you are using synchronous reset. It will make the clock quiet while reset. And if it’s used along with sync reset, then those flip-flop …

FPGA Solution for LiDAR Project

1-stop solution: Zynq UltraScale+ RFSoC ZCU111 Evaluation Kit (https://www.xilinx.com/products/boards-and-kits/zcu111.html) Features: XCZU28DR-2FFVG1517E: high-end RFSoC 12-bit 4GSPS ADC x8, 14-bit 6.5GSPS DAC x8 (all …

GCP (Google Cloud Platform)

Price of GCP# Persistance Disk# Can be used to put all the data/eda/os on it. Price (per month) Price (per GB per month) SSD 50GB $8.50 $0.17 SSD 1TB $174.08 $0.17 HDD 50GB $2.00 $0.04 HDD 200GB $8.00 $0.04 HDD 1TB …

Docker for EDA

Dockerfile# FROM ubuntu:16.04 COPY ./boot.sh /tmp COPY ./hello.sh /tmp RUN /bin/bash /tmp/boot.sh RUN /bin/bash /tmp/hello.shboot.sh# apt-get update; apt-get install -y make autoconf g++ flex bison wget cd /tmp wget …

Case Study Glitch Free Clock Mux

If your design needs to switch from one clock source to another, there is high possibility of harmful clock glitches while switching. Normally you need to stop this clock during the switching process, but what if you …

Course Note of Python Design Patterns

The course is on Coursera understanding design patterns# what? well-known solutions for recurring problems why? don’t reinvent wheels reuse best practices characteristics language neutral dynamic incomplete by …

Scala First Look

Scala introduction course on LinkedIn. Not very useful, if not using it in real project. introduction# short for Scalable language object-oriented + functional programming everything is object including numbers and …

GENUS Training Notes

The following is my notes of GENUS training course on Cadence’s training module Module 03: genus fundamentals# common UI vs legacy mode# unified commands with Tempus common us: set_db & get_db legacy mode: …

SNUG Silicon Valley 2017 at Santa Clara Conventional Center

Day 1 Morning# Paper from microsoft Microsoft has a IC design team? Apparently it does. Grey code: even when metastability happends, it falls to adjecent states, instead of unknow states, it’s acceptable in some cases. …

WSL (Windows Subsystem Linux) Tips

To start gnome-terminal on WSL (Windows Subsystem for Linux)# After upgrade to Windows 10 Creators Update, reinstall WSL will have Ubuntu 16.04.2 LTS on Windows. To reinstall WSL you should do: > lxrun /uninstall …

Install Python Offline

In general, things like Anaconda Server are designed to make this sort of workflow easier. Some suggested workarounds: Reproduce your install on another machine with internet (save conda list –export to a file and conda …

Case Study Clock Skew Control

Question: how to control the clock skew between a group of clocks to be minimum, say less than 30ps, instead of utilizing useful skew? This case happens to our hard macros. A: in Innovus, use skew group set …

INNOVUS Training Notes

The following is my notes of INNOVUS training course on Cadence’s training module Module 02: overview# “gift” directory contains lots of useful scripts to help productivity Independent “viewlog” utility or …

Deep learning and Siri by Alex Acero @ Apple

AI and ML# Artificial intellegence vs human intellegence The imitation game, eugen Goosman passed the Turing Test, 2014 Alpha Go, deepmind 2015 Introduction to deep learning# Improve on task T with respect to performance …

Deep learning with GPUs in production

Start-up: python -> enterprise: C/Java/Scala, more engineers, faster Research: quick result and prototyping GPU? Data movement between GPU and CPU is important [ ] fast.ai: class (high school math) infrastructure: …

未来的AR应该是什么样子的?

今天去参加一个AI的meetup,碰到了一个连续创业者。他介绍了自己正在做的事情:“改变现有的输入方式,不应该是人给机器输入指令,而应该是机器预测人的需求并作出相应的动作。”,这才是机器的未来。同时他提到输出界面应该AR(增强现实)这种类型的,而不应该是一个显示屏幕。

智能用电的革命

一场涉及普通消费者的智能用电革命正在悄然发生。加州政府近年来努力推动这项能源节约革命,在近三年来取得了快速的进步,得到了能源公司和电器制造商的广泛支持。 Prosumer的概念# 普通家庭以往是以单一的电力消费者出现的。但是近年来由于太阳能发电设备的推广力度不断加大,得到了大量消费者的欢迎和支持,催生了prosumer的概念,即producer + consumer。通过政府大力支持的贷款在自家屋顶安装太阳能板,并将发出的电力上网卖给电 …

Register-based SRAM Read Circuit RTL Example using generate

Some parameterized example RTL code for register-based SRAM read circuit using “generate” feature parameter d = 32; // FIFO depth parameter w = 64; // FIFO data bit-width logic [w-1:0] mem [d-1:0]; // FIFO memory array …

SystemVerilog for Design Note

This is my reading note of book “SystemVerilog for Design (2nd edition)". As a non-full-time RTL designer, it has opened my mind. But still, I’m sad about the antient tool that we are using to design …

Scan Chain Problem of Clock Generator Flip-Flops

As shown in the schematic, we have some clock divider that divide root clock by half. While in scan mode, these flip-flops will be bypassed and treated as normal flip-flop that need to be inserted into the scan chain …

《人类简史》读书笔记

对本书的评价:从理性客观中立的工科生思维出发,阐述事实,以科学的语言解释人类发展历程中的众多抉择和现象的成因。 认知革命# 智人脱颖而出的原因:imagination,并且用同样的想象出来的概念团结更多的群体,以达到合作的目的。 智人的认知革命使得文化演化取代基因演化,突破了后者变化缓慢的障碍,超越了其他人类和物种。 联合国、国家、人权这样的概念都是想象出来的,但是又是现代社会所有人都认同的“事实”,并不被认为是谎言。 共同的虚构概念建 …

“中文学校:ABC成长的烦恼”

很有意思的一篇文章。转载自纽约时报中文网 中文学校:ABC成长的烦恼 我大学第一堂中文课,老师问了一个让我念念不忘的问题:“告诉我,华裔在美国受歧视吗?” 我在加州大学伯克利分校(University of California, Berkeley)上一个专门给华裔设计的班,学生都是ABC(即美国出生的华人)。我们八点钟来上课,还在半梦半醒状态中。有的学生在啃面包,有的趴在课桌上,但我们差不多都摇头表示否认。 老师笑了。“嗯,我知道, …

Bluetooth hub from Cassia Networks

今天的CES meetup是一个清华85级的学长来宣传他们自己公司的bluetooth hub,Cassia Networks。如果简单从名字上来看并没有什么高端的感觉,毕竟bluetooth和hub两个东西都不是什么稀奇的东西。但是它获得了2016年CES大会最佳产品奖项,可不是浪得虚名。 它解决了bluetooth的痛点,而且是用现有的商业化芯片,在系统集成和固件上下功夫,通过对协议的深刻理解,得到了超出世人的性能:不需要 …

Exercise Log

运动笔记,已经好久没有记了…… 2016-01-03 Sun# 搬了新家。San Antonio附近有个不错的公园,开车过去很近,也很好停车。周五白天跑了一次,感觉很累。看了一篇知乎的帖子提到了“爱上跑步的十三周”这本书,也讲到了循序渐进的法子。其中有一个评论提到了“跑步控”这个APP,就下载到手机上,晚上去实践了一下。 因为基础较弱,所以以“9周跑5公里”为目标。 第一次训练从较长的热身开始,我以快步走的方式开始热身。五分钟。然后就 …

《爱上跑步的十三周》读书笔记

为跑步做好准备# 休息并不是避免运动那么简单,它是一个能够让你的身体从疲劳中恢复过来的合理周期。 13周跑步行走计划的技巧摘要# “拖着脚慢跑”:挺起胸,手臂摆动范围小一些,用小碎步跑,不要太高膝关节,尽量不要跳起。把重心放在脚掌的中前部。从步行到跑步要过渡的非常平稳。 每周进行三次训练。如果因为一些原因而不能完成当周的训练,最好在下一周重复这周的课程,然后再继续。 坚持记日志可以帮助你找到任何受伤的根源(1)记录每一次训练的感受,以及 …

Optical Interconnection in Google Data Centers

今天的Meetup主要讲的是Google的Data Center中optical interconnection的应用。 Presentation之后问了一个问题:Google的Data Center里面是否HDD正在被SSD取代?回答是,并没有,因为虽然SSD在读写速度上有明显的优势,但是由于存储容量的性价比非常低,所以仅是作为缓存使用。也就是说在HDD和DRAM之间再加一层SSD来提升存取速度。看来老板Sehat的FLC其实是基于市 …

Python and Freemind ElementTree

我曾经自己用Python写过一个小工具来parse Freemind文件(XML格式)然后生成RestructureText和Latex格式。这个小工具的目的是为了实践我的“源文件唯一”的理念。因为我的简历(包括既往项目总结)需要保存为两种不同格式,一个为了放在个人网站上所以是HTML,另一个当然是PDF格式。 今天瞎琢磨的时候发现,其实Python对于XML文件格式支持的极好。The ElementTree XML API这个模块 …

Ask Remote Teams to Create Daily Goals (HBR)

More organizations are starting to adopt a remote work culture. But how do managers stay in sync with what their teams are doing when they can’t see them? While it’s important to define clear goals early on, you should …

DAC 2015

汽车设计正在成为EDA公司的讨论热点。我个人觉得ANSYS应该算是其中翘楚,因为各种物理模拟器和散热模拟器是他们的传统强项。而诸如Synospsys和Mentor这样的软件厂商就更多的focus在电子系统上,硬件软件全部都有。 From Mindy: EDA厂商每年都要收购大量的小公司来保证自己的创造力。 ARM推出了专门针对Bluetooth 4.0的IP库和demo 有一家叫做flexlogic的公司,专门做FPGA阵列的IP,用 …

《看见》读后感

再次证实了一个观点:一个人的成熟(见识的增长、能力的提高)是由他的经历决定的,而不是年龄。第一次看到这种观点是在黄铁鹰的《海底捞你学不会》。其中讲述了几位年纪轻轻就担任非常重要职位的海底捞的干部。他们学历不高,但是随着海底捞成长,并一路承担了大量的责任。同样的,柴静年纪轻轻,因为调查记者的身份而经历了大量不同案例,结合自身的思考和领悟,从而获得了不符合其年龄的思想深度。她的勇气,对社会事件的深入思考,对民主社会的深刻观点,都令我自愧不 …

Be Present - Book Note of The Passionate Programmer

This was my second time to read this book. You cannot imagine how shock I was when I first read this book on Kindle. So I again bought a hardcopy and want to read it again. The topic about “present” was the first chapter …

My experience with custom digital design

Background# This is the summary of my experience from project LBRAM in Marvell in the year of 2014. The first thing: discuss timing/area/power SPEC’s in details# Most of the time, because custom design takes lots of …

Survey of Low Power Design

从2017年初的观点来看,这篇报告的部分内容过时了,但是整体结构还是比较适合的。希望今年有时间能够出一版更新的版本。 低功耗设计的最根本驱动力是集成电路芯片的功耗随着工艺的进步不仅没有下降反而不断上涨。因为晶体管速度和集成度的上升速度超过了电路单次翻转所消耗能量的下降速度,所以单位面积芯片的功耗在迅速上升。而根据ITRS的预测,固定电源供电设备和移动设备中芯片的功耗发展趋势如图表 1所示。从中我们不难看出,各类芯片的各种功耗都在不断飞 …

Posts without dates

思考:过度沟通(Over-Communication)

沟通的重要性自不必多言。对于任何规模的组织来说,沟通对于该组织的成功都是必不可少的。因为人类的合作是基于沟通的,至少在脑机接口和人类之间的心电感应成熟之前。正确的沟通能够让信息流通,并因此产生正向的合力。 有很多关于沟通的文章,我在此不再赘述如何正确沟通。在此我想提出另一个概念称为“过度沟通”。“过度沟通”中的“过度”不是贬义,而是指在正常程度以上再进一步多加一些努力,多一些重复。 为什么需要在“正常”程度以上再做努力?因为人性。 人是 …

思考:研究与工程(Thoughts about Research and Engineering)

在一个尚处于快速发展和迭代的领域,例如AI,由于各种新方向的不确定性,研究和工程常常很难互相配合甚至互相纠结掣肘。从而导致,要么工程实现犹犹豫豫、方向摇摆不定;要么研究做得差,无法做到引领方向、开拓新领域的作用。 工程实现的主要矛盾是执行效率。所以它不能够常常进行大范围的功能变更,同时要以最终产品(deliverables)作为主要导向。而且最终产品的功能(functionality)和质量(quality)是其主要量化衡量标准。因此工 …

思考:建立信任的几种方式 (How to gain trust?)

这一段时间在学习B站上资深HR专家王新宇的沟通课,其中有一个主题正好切入到了我工作中遇到的困惑。如何建立信任? 为什么?# 作为团队中的Senior IC(individual contributor),我的日常工作中百分之五十以上的时间需要跟不同团队的manager或者engineer打交道。为的是获得他们的支持或者是合作。而其中很多时候是需要跟陌生的面孔打交道。 如何能够快速获得其他团队的成员的支持常常基于最基本的利益以及两者之间的 …

Understanding LLaMA2 Part 5 Training with TinyStories

#software #ai #llm #open-source How to train a transformer based model, like LLaMA2, from scratch? Andrej Karpathy has open-sourced llama2.c project on GitHub. My learning process is “duplicate and rewrite”, …

Understanding LLaMA2 Part 4 ExecuTorch Runtime

#software #ai #llm #open-source I was involved in the early ExecuTorch definition phase and had used its predecessor Lite Interpreter extensively in work. I really like this idea and its design. This is a great effort …

Understanding LLaMA2 Part 3 PyTorch Implementation

#software #ai #llm #open-source https://github.com/jimwang99/understanding-llama2/tree/main/pytorch Above GitHub repo is an implementation of LLaMA2 and test-case use TinyStories in PyTorch Example output ------ …

Understanding LLaMA2 Part 2 KV Cache

#software #ai #llm #open-source Following up with Understanding LLaMA2 Part 1 Model Architecture, this diagram explains LLaMA model architecture with KV Cache support. We follow the same legend as well as the …

Traverse a tree

DFS (depth first search) DFS can be done easily with recursive method, because you can treat the subtrees as new trees and use the same function to traverse them. BFS (breadth first search) BFS will need extra space of a …

Traverse a graph

Mark all nodes as not visited. Create an empty ready bucket to hold nodes that has been visited itself but not its neighbors yet. Find a starting node, mark it as visited and put into the ready bucket. Get an node from …

Software is King, Especially in AI

“Software is king”# It’s a hard statement to make, as a veteran hardware engineer. Lots of pride and ego to swallow. However, it’s truly based on my observations in the industry. Allow me to …

Run ML Models Use Qualcomm SNPE Part 1 Canned Example

#software #accelerator #ai #on-device SNPE = Snapdragon Neural Processing Engine In this tutorial we assume that Qualcomm SNPE has been successfully installed use QPM. Follow “Qualcomm Package Manager 1.0” …

Python `heapq` Priority queue (heap queue)

Attributes# “Min heap”, where index 0 is the smallest item APIs# heapq.heapify(iterable) -> None: Create a heap queue in-place heapq.heappush(heap, item) -> None: Add a new item heapq.heappop(heap) …

Project Planning Process

1. Have a Vision: Without a clear vision outlining the final goals, how can we measure our progress? And it’s essential to clearly communicate this vision with all team members. 2. Break Down into Tasks and …

NV's DIGITS for On-Premise AI

TL;DR Nvidia’s DIGITS offers an on-premise AI solution aimed at smaller organizations that require strict data privacy. While it is cost-effective for small businesses and college research labs, it may be less suitable …

Navigating the Landscape of Large Language Models

Here is my notes from JPMorgan’s “Eye on Market” 2024 April issue The emergence and integration of large language models (LLMs) into various professional sectors have marked a significant milestone in …

My Tactics with Research Projects

Research projects often come with a lot of unknowns. Many engineers find this unsettling because it’s less straightforward than math or digital realm. However, I think we should welcome these uncertainties. They reflect …

Monorepo vs. Microrepo (ByteByteGo)

All the credit goes to ByteByteGo.com Do you believe that Google, Meta, Uber, and Airbnb put almost all of their code in one repository? This practice is called a monorepo. Monorepo vs. Microrepo. Which is the best? Why …

ML System Architecture (1) System Level Parallelization

To optimize the efficiency of training or executing an ML model, whether implemented locally on a device or hosted in the cloud, parallelization plays a critical role, akin to other computational challenges. Utilizing …

Memory Interfaces for LLM

LLM is a memory bound problem. This inspired me to look at different memory technologies. In this article, I’m going to summarize my research these days, especially about HBM and its impact on AI applications. …

How to Evaluate NoC (Network-on-Chip)?

Modern SoCs heavily relies on NoC to connect interfaces and storage to compute. As the ML models grow larger and larger, the data delivery ability becomes more and more important to overall system performance. While …

DeepSeek

DeepSeek has created a huge wave of discussion and panic in the market. I think it’s great overall. Innovation From a technology standpoint, DeepSeek is truly pushing the envelope in both training methods and model …

CPU Performance Test

Background# i7-12700K = Intel Core i7-12700K (8 big cores each has 2 threads, 4 little cores each has 1 thread) running at 5GHz rpi4 = Raspberry Pi 4 Rev B, with 4x Cortex-A72 running at 1.8GHz (Broadcom BCM2711) rpi5 = …

Chisel3 Systolic Array Generator

#hardware #accelerator #chisel #open-source I’ve tried to learn Chisel 5 years ago, but gave up and went back to SystemVerilog to design our RISC-V CPU + AI custom instructions from scratch. After 5 years, both …

Ceph on Raspberry Pi (2) Create Block Device

Taking my home lab Ceph distributed storage system on Raspberry Pi to the next level: make it useful by creating a block device interface so that Linux system can mount it and use it. Enable client# Install Ceph package …

Boyer-Moore Voting Algorithm

#note #learn-with-chatgpt #algorithm NOTE: this is a note from ChatGPT The Boyer-Moore Voting Algorithm, also known simply as the Voting Algorithm, is a method for finding a majority element in a sequence of elements. …

AI-Acceleration

Not Found# File Publish/🌟AI-Acceleration.md does not exist.

A simple system metric collector

https://github.com/jimwang99/jimon While experimenting my little distributed storage system built with Raspberry Pi SBCs and Ceph, some Raspberry Pi machines die now and then. My suspicion is heat, because I built them …

A Demo for Image Search

In this Github project, I created a simple application that can do text to image and image to image search, using open-source transformer model. Details can be found in the repo and its docs directory. Here is a screen …