<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Understanding LLaMA2 on When Moore's Law Ends</title><link>https://jimwang99.github.io/posts/machine-learning/understanding-llama2/</link><description>Recent content in Understanding LLaMA2 on When Moore's Law Ends</description><generator>Hugo</generator><language>en-us</language><atom:link href="https://jimwang99.github.io/posts/machine-learning/understanding-llama2/index.xml" rel="self" type="application/rss+xml"/><item><title>Understanding LLaMA2 Part 1 Model Architecture</title><link>https://jimwang99.github.io/posts/machine-learning/understanding-llama2/part-1-model-architecture/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/machine-learning/understanding-llama2/part-1-model-architecture/</guid><description>&lt;p>#software #ai #llm #open-source&lt;/p>
&lt;p>Here I&amp;rsquo;m capturing the details of llama&amp;rsquo;s model architecture use PlantUML component diagram, using the following 2 GitHub repos as references&lt;/p>
&lt;ul>
&lt;li>&lt;a href="https://github.com/facebookresearch/llama/blob/main/llama/model.py">https://github.com/facebookresearch/llama/blob/main/llama/model.py&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://github.com/karpathy/llama2.c/blob/master/model.py">https://github.com/karpathy/llama2.c/blob/master/model.py&lt;/a>&lt;/li>
&lt;/ul>
&lt;p>In the following diagram&lt;/p>
&lt;ul>
&lt;li>The &amp;ldquo;gray boxes&amp;rdquo; are tensors with their shapes in parentheses&lt;/li>
&lt;li>The &amp;ldquo;colored round dots&amp;rdquo; are operations with their categories include
&lt;ul>
&lt;li>&amp;ldquo;memops&amp;rdquo;: memory operations&lt;/li>
&lt;li>&amp;ldquo;module&amp;rdquo;: PyTorch module operators&lt;/li>
&lt;li>&amp;ldquo;compute&amp;rdquo;: non-PyTorch computational operators&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>The &amp;ldquo;grouping boxes&amp;rdquo; are to make the architecture modularized close to the original source code&lt;/li>
&lt;li>The &amp;ldquo;yellow boxes&amp;rdquo; are my in-depth comments and understanding of different parts of the architecture&lt;/li>
&lt;/ul>
&lt;p>In order to make the diagram more readable, here I choose to use abbreviations of parameters and hyper-parameters, which are explained on the upper-right corner of the diagram.&lt;/p></description></item><item><title>Understanding LLaMA2 Part 2 KV Cache</title><link>https://jimwang99.github.io/posts/machine-learning/understanding-llama2/part-2-kv-cache/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/machine-learning/understanding-llama2/part-2-kv-cache/</guid><description>&lt;p>#software #ai #llm #open-source&lt;/p>
&lt;p>Following up with &lt;a href="https://jimwang99.github.io/posts/machine-learning/understanding-llama2/part-1-model-architecture/">Understanding LLaMA2 Part 1 Model Architecture&lt;/a>, this diagram explains LLaMA model architecture with KV Cache support. We follow the same legend as well as the abbreviations.&lt;/p>
&lt;p>&lt;img src="https://jimwang99.github.io/legacy-media/llama2_architecture_kvcache.png" alt="llama2_architecture_kvcache" />&lt;/p></description></item><item><title>Understanding LLaMA2 Part 3 PyTorch Implementation</title><link>https://jimwang99.github.io/posts/machine-learning/understanding-llama2/part-3-pytorch-implementation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/machine-learning/understanding-llama2/part-3-pytorch-implementation/</guid><description>&lt;p>#software #ai #llm #open-source&lt;/p>
&lt;p>&lt;img src="https://upload.wikimedia.org/wikipedia/commons/9/96/Pytorch_logo.png" alt="pytorch-logo|300" />&lt;/p>
&lt;p>&lt;a href="https://github.com/jimwang99/understanding-llama2/tree/main/pytorch">https://github.com/jimwang99/understanding-llama2/tree/main/pytorch&lt;/a>&lt;/p>
&lt;p>Above GitHub repo is an implementation of LLaMA2 and test-case use TinyStories in PyTorch&lt;/p>
&lt;p>Example output&lt;/p>
&lt;pre tabindex="0">&lt;code>------ tinystories110M ------
Once upon a time, there was a little girl named Lily. She loved to play with her toys, especially her favorite teddy bear. One day, Lily&amp;#39;s mommy gave her a big box to play with. Lily was so happy and she jumped up and down with joy.
Lily&amp;#39;s mommy said, &amp;#34;Be careful, Lily. Don&amp;#39;t break the box.&amp;#34; Lily nodded her head and promised to be careful. She opened the box and saw that it was empty. &amp;#34;Mommy, can we fill the box with my toys?&amp;#34; Lily asked.
&amp;#34;Of course, sweetie,&amp;#34; her mommy replied. They started to fill the box with all of Lily&amp;#39;s toys. Suddenly, Lily&amp;#39;s teddy bear fell out of the box and onto the floor. &amp;#34;Oh no! Teddy fell out of the box!&amp;#34; Lily cried.
&amp;#34;Don&amp;#39;t worry, Lily. We&amp;#39;ll pick him up and put him back in the box,&amp;#34; her mommy said. They picked up the teddy bear and put him back in the box. Lily was so happy that her toys were safe in the box and she could play with them again.
------- the end ------&lt;/code>&lt;/pre></description></item><item><title>Understanding LLaMA2 Part 4 ExecuTorch Runtime</title><link>https://jimwang99.github.io/posts/machine-learning/understanding-llama2/part-4-executorch-runtime/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/machine-learning/understanding-llama2/part-4-executorch-runtime/</guid><description>&lt;p>#software #ai #llm #open-source&lt;/p>
&lt;blockquote class='book-hint '>
&lt;p>I was involved in the early ExecuTorch definition phase and had used its predecessor Lite Interpreter extensively in work. I really like this idea and its design. This is a great effort among multiple industry leading companies. So I&amp;rsquo;m trying the open-source version with LLaMA2.&lt;/p>&lt;/blockquote>&lt;h2 id="executorch-overview">&lt;a href="https://pytorch.org/executorch-overview">ExecuTorch Overview&lt;/a>&lt;a class="anchor" href="#executorch-overview">#&lt;/a>&lt;/h2>
&lt;h3 id="what-is-executorch">What is ExecuTorch?&lt;a class="anchor" href="#what-is-executorch">#&lt;/a>&lt;/h3>
&lt;p>ExecuTorch is an end-to-end solution for enabling on-device inference capabilities across mobile and edge devices including wearables, embedded devices and microcontrollers. It is part of the PyTorch Edge ecosystem and enables efficient deployment of PyTorch models to edge devices. Key value propositions of ExecuTorch are:&lt;/p></description></item><item><title>Understanding LLaMA2 Part 5 Training with TinyStories</title><link>https://jimwang99.github.io/posts/machine-learning/understanding-llama2/part-5-training-with-tinystories/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/machine-learning/understanding-llama2/part-5-training-with-tinystories/</guid><description>&lt;link rel="stylesheet" href="https://jimwang99.github.io/katex/katex.min.css" />&lt;script defer src="https://jimwang99.github.io/katex/katex.min.js">&lt;/script>&lt;script defer src="https://jimwang99.github.io/katex/auto-render.min.js" onload="renderMathInElement(document.body, {&amp;#34;delimiters&amp;#34;:[{&amp;#34;left&amp;#34;:&amp;#34;$$&amp;#34;,&amp;#34;right&amp;#34;:&amp;#34;$$&amp;#34;,&amp;#34;display&amp;#34;:true},{&amp;#34;left&amp;#34;:&amp;#34;$&amp;#34;,&amp;#34;right&amp;#34;:&amp;#34;$&amp;#34;,&amp;#34;display&amp;#34;:false},{&amp;#34;left&amp;#34;:&amp;#34;\\(&amp;#34;,&amp;#34;right&amp;#34;:&amp;#34;\\)&amp;#34;,&amp;#34;display&amp;#34;:false},{&amp;#34;left&amp;#34;:&amp;#34;\\[&amp;#34;,&amp;#34;right&amp;#34;:&amp;#34;\\]&amp;#34;,&amp;#34;display&amp;#34;:true}]});">&lt;/script>
&lt;p>#software #ai #llm #open-source&lt;/p>
&lt;p>How to train a transformer based model, like LLaMA2, from scratch? Andrej Karpathy has open-sourced &lt;a href="https://github.com/karpathy/llama2.c">llama2.c project on GitHub&lt;/a>. My learning process is &amp;ldquo;duplicate and rewrite&amp;rdquo;, which involves following the example but rewrite the code completely in my own coding style and language. I did the same and along the way I&amp;rsquo;ve learned something new. In this blog post, I&amp;rsquo;m going to break down the training code rewritten by myself, line by line, explain what and most importantly &lt;strong>why&lt;/strong>.&lt;/p></description></item><item><title>Understanding LLaMA2 Part 6 Positional Encoding</title><link>https://jimwang99.github.io/posts/machine-learning/understanding-llama2/part-6-positional-encoding/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/machine-learning/understanding-llama2/part-6-positional-encoding/</guid><description>&lt;p>#software #ai #llm #open-source&lt;/p></description></item></channel></rss>