Machine Learning

Browse the Machine Learning posts.

Posts without dates

Understanding LLaMA2 Part 5 Training with TinyStories

#software #ai #llm #open-source How to train a transformer based model, like LLaMA2, from scratch? Andrej Karpathy has open-sourced llama2.c project on GitHub. My learning process is “duplicate and rewrite”, …

Understanding LLaMA2 Part 4 ExecuTorch Runtime

#software #ai #llm #open-source I was involved in the early ExecuTorch definition phase and had used its predecessor Lite Interpreter extensively in work. I really like this idea and its design. This is a great effort …

Understanding LLaMA2 Part 3 PyTorch Implementation

#software #ai #llm #open-source https://github.com/jimwang99/understanding-llama2/tree/main/pytorch Above GitHub repo is an implementation of LLaMA2 and test-case use TinyStories in PyTorch Example output ------ …

Understanding LLaMA2 Part 2 KV Cache

#software #ai #llm #open-source Following up with Understanding LLaMA2 Part 1 Model Architecture, this diagram explains LLaMA model architecture with KV Cache support. We follow the same legend as well as the …

Run ML Models Use Qualcomm SNPE Part 1 Canned Example

#software #accelerator #ai #on-device SNPE = Snapdragon Neural Processing Engine In this tutorial we assume that Qualcomm SNPE has been successfully installed use QPM. Follow “Qualcomm Package Manager 1.0” …

NV's DIGITS for On-Premise AI

TL;DR Nvidia’s DIGITS offers an on-premise AI solution aimed at smaller organizations that require strict data privacy. While it is cost-effective for small businesses and college research labs, it may be less suitable …

Navigating the Landscape of Large Language Models

Here is my notes from JPMorgan’s “Eye on Market” 2024 April issue The emergence and integration of large language models (LLMs) into various professional sectors have marked a significant milestone in …

ML System Architecture (1) System Level Parallelization

To optimize the efficiency of training or executing an ML model, whether implemented locally on a device or hosted in the cloud, parallelization plays a critical role, akin to other computational challenges. Utilizing …

Memory Interfaces for LLM

LLM is a memory bound problem. This inspired me to look at different memory technologies. In this article, I’m going to summarize my research these days, especially about HBM and its impact on AI applications. …

DeepSeek

DeepSeek has created a huge wave of discussion and panic in the market. I think it’s great overall. Innovation From a technology standpoint, DeepSeek is truly pushing the envelope in both training methods and model …

AI-Acceleration

Not Found# File Publish/🌟AI-Acceleration.md does not exist.

A Demo for Image Search

In this Github project, I created a simple application that can do text to image and image to image search, using open-source transformer model. Details can be found in the repo and its docs directory. Here is a screen …