<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Architecture on When Moore's Law Ends</title><link>https://jimwang99.github.io/posts/architecture/</link><description>Recent content in Architecture on When Moore's Law Ends</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Mon, 18 May 2020 00:00:00 +0000</lastBuildDate><atom:link href="https://jimwang99.github.io/posts/architecture/index.xml" rel="self" type="application/rss+xml"/><item><title>IO Coherence</title><link>https://jimwang99.github.io/posts/architecture/io-coherence/</link><pubDate>Mon, 18 May 2020 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/architecture/io-coherence/</guid><description>&lt;p>Cache coherence between CPUs are most explained in textbooks, but IO coherence is not well understood. Recently I’m involved in architecture discussion about IO coherence, and found this paper, “Maintaining I/O Data Coherence in Embedded Multicore Systems” by Thomas B. Berg 2009, very useful coming to explain what is IO coherence and how to implement it in embedded system.&lt;/p>
&lt;h2 id="io-coherence">I/O Coherence&lt;a class="anchor" href="#io-coherence">#&lt;/a>&lt;/h2>
&lt;h3 id="producer-consumer-model">Producer-consumer model&lt;a class="anchor" href="#producer-consumer-model">#&lt;/a>&lt;/h3>
&lt;p>Most mechanisms for passing data between IO device and CPU, either CPU -&amp;gt; IO or IO -&amp;gt; CPU, use the classic producer-cosumer model.&lt;/p></description></item><item><title>ARM AMBA Protocl</title><link>https://jimwang99.github.io/posts/architecture/arm-amba-protocl/</link><pubDate>Mon, 25 Mar 2019 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/architecture/arm-amba-protocl/</guid><description>&lt;p>AMBA (ARM Advanced Microcontroller Bus Architecture)&lt;/p>
&lt;h2 id="1-axi">1. AXI&lt;a class="anchor" href="#1-axi">#&lt;/a>&lt;/h2>
&lt;ul>
&lt;li>AXI protocol is a &lt;strong>point-to-point&lt;/strong> protocol
&lt;ul>
&lt;li>So no matter what the network channels really use, as long as its ports comply AXI protocol, IP can be connected to them&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Main features
&lt;ul>
&lt;li>Separate read/write channels
&lt;ul>
&lt;li>Improve bandwidth&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Multiple outstanding requests&lt;/li>
&lt;li>No strict timing relationship between address and data phases&lt;/li>
&lt;li>Unaligned data transfer&lt;/li>
&lt;li>Out-of-order transaction completion&lt;/li>
&lt;li>Implicitly incremental address for burst transfer&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Transfer vs transaction
&lt;ul>
&lt;li>Transfer is a handshake: valid/ready&lt;/li>
&lt;li>Transaction is composed of multiple transfers
&lt;ul>
&lt;li>Active transaction: already started but not finished.&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Detail signals
&lt;ul>
&lt;li>&lt;code>AxBURST [1:0]&lt;/code>
&lt;ul>
&lt;li>&lt;code>00&lt;/code> = FIXED (same address) = for FIFOs&lt;/li>
&lt;li>&lt;code>01&lt;/code> = INCR = for block transfer&lt;/li>
&lt;li>&lt;code>10&lt;/code> = WRAP = suitable for cache line, critical word first
&lt;ul>
&lt;li>Must be aligned&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;code>11&lt;/code> = reserved&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;code>AxPROT [2:0]&lt;/code> defines 3 levels of access protection privileges
&lt;ul>
&lt;li>Bit 0: Privileged or not&lt;/li>
&lt;li>Bit 1: Secure or not&lt;/li>
&lt;li>Bit 2 : instruction or not (hint only)&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;code>AxCACHE [3:0]&lt;/code>
&lt;ul>
&lt;li>Bit 0: bufferable or not&lt;/li>
&lt;li>Bit 1: modifiable or not
&lt;ul>
&lt;li>Merge-able or split-able&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Bit 2: read allocate (hint only)&lt;/li>
&lt;li>Bit 3: write allocate (hint only)
&lt;ul>
&lt;li>If both RA and WA are disabled, then can skip cache and directly pass to the memory controller for access&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;code>xRESP [1:0]&lt;/code>
&lt;ul>
&lt;li>For read, each transfer of the burst has a &lt;code>RRESP&lt;/code>&lt;/li>
&lt;li>For write, at completion of the burst &lt;code>BRESP&lt;/code> is issued&lt;/li>
&lt;li>&lt;code>00&lt;/code> = normal access success = excluesive access failed&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;code>WSTRB [n-1:0]&lt;/code>: byte valid&lt;/li>
&lt;li>&lt;code>AxLOCK&lt;/code> for atomic access
&lt;ul>
&lt;li>Lock access is removed from AXI3 to AXI4&lt;/li>
&lt;li>In AXI4 only exclusive is supported for better network fabric performance, to enable the implementation of semaphore type operations without requiring the bus to remain locked&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;code>AxQOS [3:0]&lt;/code> defines the priority of a transaction, larger number means higher priority
&lt;ul>
&lt;li>Priority is guaranteed by arbiters&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;code>AxREGION [3:0]&lt;/code> for a single physical slave that provides multiple logical interfaces in different address region of the whole memory space&lt;/li>
&lt;li>&lt;code>AxUSER&lt;/code> implemenation defined width, optional. So can be imcompatible&lt;/li>
&lt;li>Dependency
&lt;ul>
&lt;li>&lt;code>WVALID&lt;/code> can assert before &lt;code>AWVALID&lt;/code>&lt;/li>
&lt;li>&lt;code>WLAST&lt;/code> must complete before &lt;code>BVALID&lt;/code>&lt;/li>
&lt;li>&lt;code>RVALID&lt;/code> cannot be asserted until &lt;code>ARADDR&lt;/code> has been transferred&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Atomic access
&lt;ul>
&lt;li>AXI3 has “locked access” which blocks all other masters to locked slave, while AXI4 impoved it to “exclusive access” which only blocks access to particular region.&lt;/li>
&lt;li>Locked transaction
&lt;ul>
&lt;li>Ensure no outstanding transaction&lt;/li>
&lt;li>Initial lock transfer, complete with an unlocked transfer&lt;/li>
&lt;li>Performance impact is huge&lt;/li>
&lt;li>Lock access is enforced by network fabric&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Exclusive access
&lt;ul>
&lt;li>Semaphore style READ &amp;amp; WRITE&lt;/li>
&lt;li>Requires slave hardware support
&lt;ul>
&lt;li>“Exclusive access monitor” to save exclusive operation source ID and target memory address when exclusive READ access happens, remove entry when exclusive WRITE access happens&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;code>xRESP&lt;/code>
&lt;ul>
&lt;li>&lt;code>EXOKAY&lt;/code> means successful, but &lt;code>OKAY&lt;/code> means exclusive fail&lt;/li>
&lt;li>When exclusive write reply with OKAY, the memory won’t get updated&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Ordering
&lt;ul>
&lt;li>Write
&lt;ul>
&lt;li>&lt;code>W&lt;/code> channel must follow the same order of &lt;code>AW&lt;/code> channel&lt;/li>
&lt;li>Different transaction IDs on &lt;code>W&lt;/code> can be interleaved, but same ID must in order even for different transactions&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Read
&lt;ul>
&lt;li>No ordering between &lt;code>R&lt;/code> and &lt;code>AR&lt;/code>&lt;/li>
&lt;li>Different transaction IDs on &lt;code>R&lt;/code> can be interleaved, but same ID must in order&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Read and write don’t have ordering between each other&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Alignment
&lt;ul>
&lt;li>Unaligned start address only affects the first transfer in a transaction, all following transfer is aligned to &lt;code>AxSIZE&lt;/code> to relax requirement on slave side&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Interface attributes
&lt;ul>
&lt;li>Write
&lt;ul>
&lt;li>Issuing capability: a master can issue how many outstanding transactions&lt;/li>
&lt;li>Interleave depth: a slave can receive how many outstanding trasactions&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Read
&lt;ul>
&lt;li>Issuing capability: how many transactions a master can issue&lt;/li>
&lt;li>Acceptance capability: how many transactions a slave can accept&lt;/li>
&lt;li>Reordering depth: how many transactions a slave can transmit data&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul></description></item><item><title>ARM Training Cortex Processor Behaviors</title><link>https://jimwang99.github.io/posts/architecture/arm-training-cortex-processor-behaviors/</link><pubDate>Mon, 18 Mar 2019 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/architecture/arm-training-cortex-processor-behaviors/</guid><description>&lt;h2 id="simple-sequential-execution-model">Simple Sequential Execution Model&lt;a class="anchor" href="#simple-sequential-execution-model">#&lt;/a>&lt;/h2>
&lt;p>After optimization, the result should be exactly the same with “simple sequential execution model”.&lt;/p>
&lt;h2 id="optimization-instruction-fetching">Optimization: Instruction Fetching&lt;a class="anchor" href="#optimization-instruction-fetching">#&lt;/a>&lt;/h2>
&lt;ul>
&lt;li>Fetch multiple instructions from memory&lt;/li>
&lt;li>Branch
&lt;ul>
&lt;li>Predictive fetch and execution&lt;/li>
&lt;li>Branch caches&lt;/li>
&lt;li>Return stack
&lt;ul>
&lt;li>The “link register” value is pushed into the stack, and used/popped when return&lt;/li>
&lt;li>E.g. 4 entries of return stack, but my call stack is 5, then miss will happend. Then what to do? Discard all of the content probably will be the most obvious answer.&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Conditional branch prediction
&lt;ul>
&lt;li>Shift register of branch execution history, taken is 1’b1 and non-taken is 1’b0&lt;/li>
&lt;li>Use the pattern to predict the next branch&lt;/li>
&lt;li>Use branch monitor’s result to update this prediction table&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Misprediction
&lt;ul>
&lt;li>Predict into non-executable memory region, MMU needs to react&lt;/li>
&lt;li>Some architecures require control bits to disable prediction&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;p>&lt;img src="https://jimwang99.github.io/legacy-media/feedback-branch-prediction.png" alt="Feedback branch prediction" />&lt;/p></description></item><item><title>ARMv8 Architecture</title><link>https://jimwang99.github.io/posts/architecture/armv8-architecture/</link><pubDate>Mon, 18 Mar 2019 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/architecture/armv8-architecture/</guid><description>&lt;blockquote class='book-hint '>
&lt;p>ARM online training note&lt;/p>&lt;/blockquote>&lt;h2 id="1-introduction">1. Introduction&lt;a class="anchor" href="#1-introduction">#&lt;/a>&lt;/h2>
&lt;h3 id="what-is-an-architecture">What is an architecture?&lt;a class="anchor" href="#what-is-an-architecture">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>Instruction set&lt;/li>
&lt;li>Exception model&lt;/li>
&lt;li>Memory model&lt;/li>
&lt;li>Debug&lt;/li>
&lt;/ul>
&lt;h3 id="armv8">ARMv8&lt;a class="anchor" href="#armv8">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>AArch32 vs AArch64
&lt;ul>
&lt;li>AArch32: backward compatible to ARMv7&lt;/li>
&lt;li>AArch64: fixed 32-bit instruction, new exception model, 64-bit virtual address&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h4 id="priviledge-and-security-model">Priviledge and security model&lt;a class="anchor" href="#priviledge-and-security-model">#&lt;/a>&lt;/h4>
&lt;p>&lt;img src="https://jimwang99.github.io/legacy-media/aarch64-priviledge-security-model.png" alt="AArch64 Priviledge and Security Model" />&lt;/p>
&lt;ul>
&lt;li>4-level of privilege
&lt;ul>
&lt;li>EL0 &amp;lt; EL1 &amp;lt; EL2 &amp;lt; EL3, larger the higher privilege&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>2 security modes&lt;/li>
&lt;/ul>
&lt;h4 id="mixture-of-aarch32-and-aarch64">Mixture of AArch32 and AArch64&lt;a class="anchor" href="#mixture-of-aarch32-and-aarch64">#&lt;/a>&lt;/h4>
&lt;ul>
&lt;li>Only 64-bit OS can host a mix of 32-bit and 64-bit apps&lt;/li>
&lt;li>32-bit app can only be on lower EL level&lt;/li>
&lt;/ul>
&lt;h2 id="2-isa">2. ISA&lt;a class="anchor" href="#2-isa">#&lt;/a>&lt;/h2>
&lt;h3 id="register">Register&lt;a class="anchor" href="#register">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>X0 to X30: 31 general purpose registers
&lt;ul>
&lt;li>W0 to W30 are their 32-bit form&lt;/li>
&lt;li>Zero register: XZR and WZR&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>V0 to V31: floating point, SIMD, crypto operations
&lt;ul>
&lt;li>Multiple view: B(8), H(16), S(32), D(64), Q(128)&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>System registers
&lt;ul>
&lt;li>MSR / MRS: move from/to system register to/from generator purpose register&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="data-processing">Data processing&lt;a class="anchor" href="#data-processing">#&lt;/a>&lt;/h3>
&lt;h3 id="flow-control">Flow control&lt;a class="anchor" href="#flow-control">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>ARM also have implicitly flag registers that are results of comparisons&lt;/li>
&lt;/ul>
&lt;h3 id="pcs-proceduare-call-standard">PCS (proceduare call standard)&lt;a class="anchor" href="#pcs-proceduare-call-standard">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>Parameter pass in: X0 - X7&lt;/li>
&lt;li>Return value: X0 - X1&lt;/li>
&lt;li>Must preserve: X19 - X29&lt;/li>
&lt;li>Can corrupt: X0 - X18&lt;/li>
&lt;li>Return address (LR): X30&lt;/li>
&lt;/ul>
&lt;h3 id="load-and-store">Load and store&lt;a class="anchor" href="#load-and-store">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>&lt;code>LDR / STR W0, [X1, #12]&lt;/code> (X1 is not changed)
&lt;ul>
&lt;li>Pre-index: &lt;code>[X1, #12]!&lt;/code> (X1 is changed then used)&lt;/li>
&lt;li>Post-index: &lt;code>[X1], #12&lt;/code> (X1 is used, then changed)&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="floating-point">Floating point&lt;a class="anchor" href="#floating-point">#&lt;/a>&lt;/h3>
&lt;h3 id="simd">SIMD&lt;a class="anchor" href="#simd">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>Lane = whole Vx register &amp;amp; element&lt;/li>
&lt;li>Neon&lt;/li>
&lt;/ul>
&lt;h3 id="vectors">Vectors&lt;a class="anchor" href="#vectors">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>&lt;code>Vn.xy&lt;/code>
&lt;ul>
&lt;li>n = register number&lt;/li>
&lt;li>x = number of elements&lt;/li>
&lt;li>y = size of the elements (B/H/S/D/Q)&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Total vector length = 128-bit / 64-bit for instructions to work on a whole vector&lt;/li>
&lt;li>Special instructions work on individual elements&lt;/li>
&lt;/ul>
&lt;h2 id="3-exception">3. Exception&lt;a class="anchor" href="#3-exception">#&lt;/a>&lt;/h2>
&lt;ul>
&lt;li>Synchronous = exception&lt;/li>
&lt;li>Asynchronous = interrupt&lt;/li>
&lt;/ul>
&lt;h3 id="exception-level-el">Exception level (EL)&lt;a class="anchor" href="#exception-level-el">#&lt;/a>&lt;/h3>
&lt;h3 id="pstate--spsr">PSTATE &amp;lt;=&amp;gt; SPSR&lt;a class="anchor" href="#pstate--spsr">#&lt;/a>&lt;/h3>
&lt;p>PSTATE is the current state of the processor, and SPSR is the registers to save the PSTATE.&lt;/p></description></item><item><title>Interrupts and ARM GIC Architecture</title><link>https://jimwang99.github.io/posts/architecture/interrupts-and-arm-gic-architecture/</link><pubDate>Mon, 11 Mar 2019 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/architecture/interrupts-and-arm-gic-architecture/</guid><description>&lt;p>Reference&lt;/p>
&lt;ul>
&lt;li>&lt;a href="https://en.m.wikipedia.org/w/index.php?title=Interrupt">Interrupt&lt;/a>&lt;/li>
&lt;/ul>
&lt;h2 id="categorization">Categorization&lt;a class="anchor" href="#categorization">#&lt;/a>&lt;/h2>
&lt;ul>
&lt;li>Hardware vs. Software
&lt;ul>
&lt;li>Hardware: usually caused by peripheral or other processors
&lt;ul>
&lt;li>IRQ: maskable interrupt&lt;/li>
&lt;li>NMI: non-maskable interrupt
&lt;ul>
&lt;li>For highest priority tasks, like times, especially wathdog timers
&lt;ul>
&lt;li>&lt;strong>&lt;a href="https://en.m.wikipedia.org/wiki/Watchdog_timer">Wathdog timer&lt;/a>&lt;/strong>: a timer has to be reset by software on purpose periodically, otherwise it means the software has gone into some hanging situation and will trigger watchdog routine to recover or reboot.&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>IPI: inter-processor interrupt&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Software: caused by exception or special instructions that used to implement system calls&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Interrupt vs inter-process communication signal
&lt;ul>
&lt;li>Interrupt: mediated by the processor (hardware); handled by the kernel&lt;/li>
&lt;li>Signal: mediated by the kernel (through systeam call); handled by processes
&lt;ul>
&lt;li>Such as: SIGSEGV, SIGBUS, SIGILL, SIGFPE&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Precise vs imprecise interrupt
&lt;ul>
&lt;li>Precise interrupts has
&lt;ul>
&lt;li>PC and other architecture states are saved, so after interrupt handler is done the current process can resume&lt;/li>
&lt;li>All instructions before the time point have fully executed, and no instructions beyond has been executed (or they are killed)&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Triggering methods
&lt;ul>
&lt;li>Physical interrupt
&lt;ul>
&lt;li>Level-triggered vs. Edge-triggered&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Message-signaled interrupts (or message-based interrupt as in ARM’s term)
&lt;ul>
&lt;li>Supported by PCI 2.2 and PCI-Express&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h2 id="msi-message-signaled-interrupts">MSI (Message Signaled Interrupts)&lt;a class="anchor" href="#msi-message-signaled-interrupts">#&lt;/a>&lt;/h2>
&lt;ul>
&lt;li>Triggerred by write to a memory address&lt;/li>
&lt;li>Can be converted from/to physical interrupt&lt;/li>
&lt;li>In-band vs. out-of-band
&lt;ul>
&lt;li>Dedicated interrupts wires are considered out-of-band, while MSI is in-band.&lt;/li>
&lt;li>Need hardware to convert MSI to physical interupts&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Although use in-band network to transmit interrupt info, but it can only be used to descript the interrupt (such as source, priority) while cannot be used to carry data.&lt;/li>
&lt;li>Pros &amp;amp; cons
&lt;ul>
&lt;li>More scalable
&lt;ul>
&lt;li>Multi-sources to multi-processors&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Simpler, cheaper, more reliable (no interference noises)&lt;/li>
&lt;li>Not compatible with devices that need physical interrupts&lt;/li>
&lt;li>Need software support&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h2 id="performance-issue">Performance issue&lt;a class="anchor" href="#performance-issue">#&lt;/a>&lt;/h2>
&lt;ul>
&lt;li>Livelocks&lt;/li>
&lt;/ul>
&lt;h2 id="arm-gic">ARM GIC&lt;a class="anchor" href="#arm-gic">#&lt;/a>&lt;/h2>
&lt;h3 id="categories">Categories&lt;a class="anchor" href="#categories">#&lt;/a>&lt;/h3>
&lt;table>
 &lt;thead>
 &lt;tr>
 &lt;th>LPI (locality-specific peripheral interrupt)&lt;/th>
 &lt;th>PPI (private peripheral interrupt)&lt;/th>
 &lt;th>SPI (shared peripheral interrupt)&lt;/th>
 &lt;th>SGI (software generated interrupt)&lt;/th>
 &lt;/tr>
 &lt;/thead>
 &lt;tbody>
 &lt;tr>
 &lt;td>From peripheral to local PE (processing element)&lt;/td>
 &lt;td>From peripheral to a single, specific PE&lt;/td>
 &lt;td>From peripheral to distributor, then to PE that can accept this type of interrupt&lt;/td>
 &lt;td>From PE to PE, typically used for inter-processor communication&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>Edge-triggered behavior, message-based (???)&lt;/td>
 &lt;td>Edge-/level-triggered, need explicit deactivation&lt;/td>
 &lt;td>edge-/level-triggered, need explicit deactivation&lt;/td>
 &lt;td>Edge-triggered, need explicit deactivation&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>Can be routed with ITS&lt;/td>
 &lt;td>Cannot be routed with ITS&lt;/td>
 &lt;td>Cannot be routed with ITS&lt;/td>
 &lt;td>Cannot be routed with ITS&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>Non-secure&lt;/td>
 &lt;td>Secure or non-secure&lt;/td>
 &lt;td>Secure or non-secure&lt;/td>
 &lt;td>Secure or non-secure&lt;/td>
 &lt;/tr>
 &lt;/tbody>
&lt;/table>
&lt;p>&lt;strong>Q: if need deactivation, then there has to be some ackknowledge mechanisms. what are they???&lt;/strong>&lt;/p></description></item><item><title>ARM's big.LITTLE Architecture</title><link>https://jimwang99.github.io/posts/architecture/arm-s-big-little-architecture/</link><pubDate>Wed, 30 Jan 2019 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/architecture/arm-s-big-little-architecture/</guid><description>&lt;p>&lt;a href="https://developer.arm.com/technologies/big-little">https://developer.arm.com/technologies/big-little&lt;/a>&lt;/p>
&lt;p>big.LITTLE is a practical example of SMP (Symmetric Multiprocessing). It combines high performance CPU cores and low power CPU cores in the same chip, connected using cache coherent interconnect, to achieve &lt;strong>high peak performance within thermal bounds of the system&lt;/strong> when intense computational power is needed, as well as &lt;strong>maximum energy efficiency&lt;/strong> when the device is in light usage mode most of the time. It’s a particular adaption to mobile devices usage.&lt;/p></description></item><item><title>Re-discover Hardware Security in Modern SoC</title><link>https://jimwang99.github.io/posts/architecture/re-discover-hardware-security-in-modern-soc/</link><pubDate>Wed, 30 Jan 2019 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/architecture/re-discover-hardware-security-in-modern-soc/</guid><description>&lt;p>There is a big difference between how I used to understand hardware security and state-of-the-art security supported by hardware software co-design, after I watched some video talking about SEP (Security Enclave Processor) by Apple. It’s a key component in current iPhone to protect user data and password from being observed in any kind of hacking, including traditional side channel attack such as DPA (Dynamic Power Attack), debug channel attack, normal network attack, and etc.&lt;/p></description></item><item><title>Huwcha Accelerator architecture</title><link>https://jimwang99.github.io/posts/architecture/huwcha-accelerator-architecture/</link><pubDate>Thu, 03 Jan 2019 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/architecture/huwcha-accelerator-architecture/</guid><description>&lt;p>From the reading of this paper, “The Hwacha Microarchitecture Manual, Version 3.8.1”, I found out that our Pygmy ES1 architecture is almost the same idea, just not as fancy.&lt;/p>
&lt;ul>
&lt;li>We don’t have cache coherency, because we operate on unified physical memory space&lt;/li>
&lt;li>The vector execution unit is not as fancy, just useless multithreading, no systolic&lt;/li>
&lt;li>The prefetch is supposed to be done by DMA engine that is manually controlled by software&lt;/li>
&lt;/ul>
&lt;h2 id="system-architecture">System architecture&lt;a class="anchor" href="#system-architecture">#&lt;/a>&lt;/h2>
&lt;ul>
&lt;li>The vector accelerator only has L1 I$, no D$
&lt;ul>
&lt;li>Don’t need to maintain cache coherency&lt;/li>
&lt;li>Lots of vector registers, 512 in total, each is 64x2x4=512-bit&lt;/li>
&lt;li>Wide bus connection to L2$, to provide higher bandwidth&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Uncached TileLink between L1 I$ (both scalar processor and vector processors) and L2$&lt;/li>
&lt;li>Cached TileLink between L1 D$ (only in scalar processor)
&lt;ul>
&lt;li>L2$ maintains directory bits, which determines the states of corresponding cache line (JW: maybe something like MESI bits)&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Operations of L2$, supported by TileLink protocol (“Productive Design of Extensible On-Chip Memory Hierarchies”)
&lt;ul>
&lt;li>Sub-cache-block accesses&lt;/li>
&lt;li>Data prefetch requests
&lt;ul>
&lt;li>Read from DDR, don’t need to send back to the requester&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Atomic memory operations
&lt;ul>
&lt;li>ALU inside L2 cache banks&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h2 id="decoupling">Decoupling&lt;a class="anchor" href="#decoupling">#&lt;/a>&lt;/h2>
&lt;ul>
&lt;li>Access/execute decoupling&lt;/li>
&lt;li>Decoupled vector arch&lt;/li>
&lt;li>Cache refill/access decoupling&lt;/li>
&lt;/ul>
&lt;h2 id="vector-command-queue-vcmdq">Vector Command Queue (VCMDQ)&lt;a class="anchor" href="#vector-command-queue-vcmdq">#&lt;/a>&lt;/h2>
&lt;ul>
&lt;li>Instruction fetch is handled by scalar processor, and then sent to VCMDQ
&lt;ul>
&lt;li>There is explicity defined start &lt;code>vf&lt;/code> and stop &lt;code>vstop&lt;/code> instructions that flags the begin and end of vector instructions
&lt;ul>
&lt;li>JW: why is that necessary?&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h2 id="vector-execution-unit-vxu">Vector Execution Unit (VXU)&lt;a class="anchor" href="#vector-execution-unit-vxu">#&lt;/a>&lt;/h2>
&lt;ul>
&lt;li>In a &lt;strong>systolic&lt;/strong> style
&lt;ul>
&lt;li>4 banks in total, each bank has 256-entry 2x64-bit vector register file, as well as ALUs&lt;/li>
&lt;li>from the block diagram, these ALUs are only add/subtration/shift operations. Multi-cycle integer multiplier/divider, and floating point operations are outside of each bank and shared by all 4 banks via a crossbar.
&lt;ul>
&lt;li>JW: I don’t think this is good for ML applications. They are trying to make it more generic, so to improve on this, we could take the same micro-architecture but simplify it as well as putting more ML (basically MAC) into it.&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>As operation flows through the banks, natually chain different operations together.
&lt;ul>
&lt;li>JW: the chaining is useful because we have limited shared function unit, such as IMUL/IDIV. For function units that are exclusive to bank, I don’t see the benifit.&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Predicate is used for simple branch and etc.&lt;/li>
&lt;/ul>
&lt;h2 id="vector-memory-unit-vmu">Vector Memory Unit (VMU)&lt;a class="anchor" href="#vector-memory-unit-vmu">#&lt;/a>&lt;/h2>
&lt;ul>
&lt;li>It’s based TileLink protocol&lt;/li>
&lt;li>Vector Load Unit (VLU)
&lt;ul>
&lt;li>Opportunistic writeback mechanism: return as soon as the data is back, no re-order buffer.&lt;/li>
&lt;li>Simultaneously manage multiple operations to avoid artificial throttling of successive loads: too many requests will drive performance down&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Vector Store Unit (VSU)
&lt;ul>
&lt;li>Vector Store Data Queue (VSDQ)&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h2 id="vector-runahead-unit-vru">Vector Runahead Unit (VRU)&lt;a class="anchor" href="#vector-runahead-unit-vru">#&lt;/a>&lt;/h2>
&lt;ul>
&lt;li>This block process all the vector load/store instructions in the prefetch queue from the scalar processor, generate address and send prefetch requests to L2$.
&lt;ul>
&lt;li>Prefetch request doesn’t return data to requester.&lt;/li>
&lt;li>In ideal world, it will hide the memory access latency.&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>When from a cold start, it skips the first few load/store to run ahead of the VMU, until it reaches too far ahead, determined by the number of load/store operations that hasn’t been processed by VMU yet, it pauses.
&lt;ul>
&lt;li>Cannot run too close, otherwise the memory access latency cannot be hide.&lt;/li>
&lt;li>Cannot run too further ahead, otherwise newly loaded data will force evict the data that’s been currently using by the VXU and cause even worse performance panity.&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Also to limit using too much of the L2$ bandwidth, prefetch in L2$ only uses up 1/3 of available outstanding access (JW: I think it more likely the MSHR entries which handles any cache miss).&lt;/li>
&lt;/ul>
&lt;h2 id="multilane">Multilane&lt;a class="anchor" href="#multilane">#&lt;/a>&lt;/h2>
&lt;ul>
&lt;li>Hwacha support paramerized number of identical lanes, and these lanes work entirely decoupled from one another.&lt;/li>
&lt;li>If the memory needed by each lane aligns to the cache line size, it will be the best, otherwise, there will be some waste of bandwidth to load unnecessary data to the lane.&lt;/li>
&lt;li>JW: curious, how do they manage multilane?
&lt;ul>
&lt;li>There is a &lt;strong>master sequencer&lt;/strong> who knows the number of lanes out there as well as the vector length. It will dispatch the job to these lanes evenly.&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>JW: as it discussed in the paper, the lanes work on strip size, which is a unit-stride fasion. If there is any non-unit stride, not only it will waste the bandwidth, but also it will waste the computation resource. &lt;em>So I think it would be a simple dispatch based on unit-stride, the master sequencer will have to consider the stride as well.&lt;/em>&lt;/li>
&lt;/ul></description></item><item><title>Note of RISC-V Vector ISA Spec v0.6</title><link>https://jimwang99.github.io/posts/architecture/note-of-risc-v-vector-isa-spec-v0-6/</link><pubDate>Sat, 15 Dec 2018 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/architecture/note-of-risc-v-vector-isa-spec-v0-6/</guid><description>&lt;h2 id="vector-regfile">Vector regfile&lt;a class="anchor" href="#vector-regfile">#&lt;/a>&lt;/h2>
&lt;ul>
&lt;li>32 of them, v0 to v31&lt;/li>
&lt;li>Each is &lt;code>VLEN&lt;/code> bits&lt;/li>
&lt;li>Each can be divided into several elements
&lt;ul>
&lt;li>The max element width is &lt;code>ELEN&lt;/code>&lt;/li>
&lt;li>CSR &lt;code>vsew&lt;/code> maps to &lt;code>SEW&lt;/code> (standard element width) controls their width dynamically&lt;/li>
&lt;li>CSR &lt;code>vl&lt;/code> controls the number of elements to operate on for vector instructions&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Packing of shorter vector
&lt;ul>
&lt;li>when &lt;code>SEW&lt;/code> is smaller than &lt;code>ELEN&lt;/code>, multiple SEW will be packed into one &lt;code>ELEN&lt;/code> unit
&lt;ul>
&lt;li>Following little-endian rule&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;code>ELEN&lt;/code> units are packed into &lt;code>VLEN&lt;/code> register also
&lt;ul>
&lt;li>Following little-endiam rule&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Storage of longer vector
&lt;ul>
&lt;li>If operand longer than &lt;code>SEW&lt;/code> is needed, then
&lt;ul>
&lt;li>Even-numbered vector register holds the even-numbered elements&lt;/li>
&lt;li>Odd-numbered vector register holds the odd-numberred elements&lt;/li>
&lt;li>&lt;strong>WHY?&lt;/strong> the author said it’s designed to simplify data path alignment&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Register grouping via &lt;code>vlmul&lt;/code> field
&lt;ul>
&lt;li>Any operation applies on one vector register &lt;code>n&lt;/code> will apply on all the other vector registers in the same group&lt;/li>
&lt;li>Depending on the &lt;code>vlmul&lt;/code>, 00 = no group, 01 = v[n+16], 02 = v[n+8], 03 = v[n+4]
&lt;ul>
&lt;li>&lt;code>VLMAX&lt;/code> will change accordingly, by double, quarod, 8 times&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Reuse for floating-point regfile
&lt;ul>
&lt;li>Floating-point registers reside at the LSB &lt;code>FLEN&lt;/code> of the vector registers&lt;/li>
&lt;li>Lower precision floating-point types are NaN-boxed (MSB bits are set to &lt;code>'b1&lt;/code>)&lt;/li>
&lt;li>Loading floating-point data doesn’t change the upper bits where &lt;code>VLEN&lt;/code> &amp;gt; &lt;code>FLEN&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h2 id="vector-operation">Vector operation&lt;a class="anchor" href="#vector-operation">#&lt;/a>&lt;/h2>
&lt;ul>
&lt;li>Masking
&lt;ul>
&lt;li>Masked off elements are skipped, and will not generate exceptions&lt;/li>
&lt;li>Use the LSB of each SEW in &lt;code>v0&lt;/code> as the masking bit&lt;/li>
&lt;li>4 types in &lt;code>vm[1:0]&lt;/code>: true, false, scalar, no-masking&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Vector load/store
&lt;ul>
&lt;li>Unit-stride
&lt;ul>
&lt;li>Starting address = base address in RS1 + imm offset&lt;/li>
&lt;li>Fault-first version
&lt;ul>
&lt;li>&lt;strong>WHY?&lt;/strong>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Constant-stride
&lt;ul>
&lt;li>Base in RS1, stride in RS2, by an unit of byte&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Indexed (scatter-gather)
&lt;ul>
&lt;li>Base in RS1, indices in VRS2 (signed integers)
&lt;ul>
&lt;li>If &lt;code>SEW&lt;/code> is larger than &lt;code>XLEN&lt;/code>, LSB bits are taken&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>For store, there is an extra &lt;em>unordered-indexed load&lt;/em>, in contrast to &lt;em>ordered-indexed load&lt;/em>
&lt;ul>
&lt;li>JW: we can make &lt;em>unordered&lt;/em> version to be higher performance&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Vector arithmetic
&lt;ul>
&lt;li>Widening operation with &lt;code>w&lt;/code> suffix puts 2xSEW wide destination into an even-odd vector register pair&lt;/li>
&lt;li>Merge: use the mask field to merge 2 source operands&lt;/li>
&lt;li>Narrowing: convert multi-width vector into single-width vector&lt;/li>
&lt;li>Reduction: scalar &amp;lt;= vector op scalar&lt;/li>
&lt;li>Matrix multiplication support
&lt;ul>
&lt;li>Fused-multiply-add + reduction&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Mask
&lt;ul>
&lt;li>Mask population count&lt;/li>
&lt;li>Find-first-set mask bit&lt;/li>
&lt;li>&lt;code>viota&lt;/code>: count the number of elements in &lt;code>v0&lt;/code> that’s masked or un-masked
&lt;ul>
&lt;li>Can be combined with scatter/gather instructions to perform vector compress/expand instructions&lt;/li>
&lt;li>They are still discussing if this should be able to take any registers, not only &lt;code>v0&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Set-before-frst mask bit&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Permutation
&lt;ul>
&lt;li>insert/extract&lt;/li>
&lt;li>slide&lt;/li>
&lt;li>gather&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h2 id="exception">Exception&lt;a class="anchor" href="#exception">#&lt;/a>&lt;/h2>
&lt;ul>
&lt;li>CSR &lt;code>progress&lt;/code> logs the element index that caused the trap
&lt;ul>
&lt;li>With this, failed operations can resume from where exactly it traps exception&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Precise vs imprecise traps
&lt;ul>
&lt;li>Precise trap: slow, but support debug&lt;/li>
&lt;li>Imprecise trap: fast, but possibly obscure error conditions&lt;/li>
&lt;li>Implementations can choose&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul></description></item><item><title>Ariane (PULP series high-performance core)</title><link>https://jimwang99.github.io/posts/architecture/ariane-pulp-series-high-performance-core/</link><pubDate>Sat, 01 Dec 2018 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/architecture/ariane-pulp-series-high-performance-core/</guid><description>&lt;p>&lt;a href="https://pulp-platform.github.io/ariane/docs/home/">Ariane Document&lt;/a>&lt;/p>
&lt;p>&lt;img src="https://jimwang99.github.io/legacy-media/ariane_overview.png" alt="Ariane Block Diagram" />&lt;/p>
&lt;h2 id="architecture-note">Architecture note&lt;a class="anchor" href="#architecture-note">#&lt;/a>&lt;/h2>
&lt;h3 id="pc-gen-stage">PC gen stage&lt;a class="anchor" href="#pc-gen-stage">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>The fetching address for i-cache is always word-aligned.&lt;/li>
&lt;/ul>
&lt;h3 id="fetch-stage">Fetch stage&lt;a class="anchor" href="#fetch-stage">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>
&lt;p>Its fetch stage doesn’t have much decoding work to do, only the necessary one to generate next PC. And it relies on its branch prediction to give out next PC.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>There is an internal FIFO with 2 entries to log the PC (and other meta-info) that was sent to i-cache, while waiting for its response.&lt;/p></description></item><item><title>Network-on-Chip Notes</title><link>https://jimwang99.github.io/posts/architecture/network-on-chip-notes/</link><pubDate>Sat, 01 Dec 2018 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/architecture/network-on-chip-notes/</guid><description>&lt;h2 id="noc">NoC&lt;a class="anchor" href="#noc">#&lt;/a>&lt;/h2>
&lt;ul>
&lt;li>Clustering coefficient: the most intuitive explanation is the number of hops between two random nodes in the network.&lt;/li>
&lt;li>Layers
&lt;ul>
&lt;li>Physical layer&lt;/li>
&lt;li>Link layer
&lt;ul>
&lt;li>Transaction protocol: such as AXI&lt;/li>
&lt;li>Seperated channels like AXI, to avoid dead-lock caused by depency problem&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Transport layer
&lt;ul>
&lt;li>Packet: header &amp;amp; payload&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Flow control
&lt;ul>
&lt;li>On link-level or end-to-end level
&lt;ul>
&lt;li>Link-level: every hop there is a notification, just like valid-ready protocol&lt;/li>
&lt;li>End-to-end level: every transaction of data from sender to receiver has to have some kind of notification that the receiver notify the sender it has received the data successfully.&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Advantages:
&lt;ul>
&lt;li>Multiple transactions can complete in parallel in the network to utilize the routing resource.&lt;/li>
&lt;li>Layered structure makes it more clear, and each layer can be implemented independently.&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul></description></item><item><title>Cache Coherence Notes</title><link>https://jimwang99.github.io/posts/architecture/cache-coherence-notes/</link><pubDate>Thu, 01 Nov 2018 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/architecture/cache-coherence-notes/</guid><description>&lt;h2 id="coherence-mechanism">Coherence mechanism&lt;a class="anchor" href="#coherence-mechanism">#&lt;/a>&lt;/h2>
&lt;h3 id="snooping">Snooping&lt;a class="anchor" href="#snooping">#&lt;/a>&lt;/h3>
&lt;p>Every cache maintain its own cache state. And when it needs to access a shared address space, it sends snooping messages to all the other caches to either update or invalidate them.&lt;/p>
&lt;ul>
&lt;li>Write invalidate: write operation will invalidate all the other shared copies. Others will have to read again from the next level of cache to use it again.&lt;/li>
&lt;li>Write update: write operation will give the written data to the shared copies and update them accordingly.&lt;/li>
&lt;/ul>
&lt;p>One could add a snooping filter to filter out the exesive snooping traffic that doesn’t belong to current cache.&lt;/p></description></item><item><title>CPU Architecture Notes</title><link>https://jimwang99.github.io/posts/architecture/cpu-architecture-notes/</link><pubDate>Thu, 01 Nov 2018 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/architecture/cpu-architecture-notes/</guid><description>&lt;h2 id="register-renaming">Register renaming&lt;a class="anchor" href="#register-renaming">#&lt;/a>&lt;/h2>
&lt;p>To eliminate the &lt;strong>false and output data dependency&lt;/strong> by adding extra physical registers more than architectural registers.&lt;/p>
&lt;ul>
&lt;li>Read-after-write (RAW) is &lt;strong>true data dependency&lt;/strong>&lt;/li>
&lt;li>Write-after-write (WAW) is &lt;strong>output data dependency&lt;/strong>&lt;/li>
&lt;li>Write-after-read (WAR) is &lt;strong>false data dependency&lt;/strong>&lt;/li>
&lt;/ul>
&lt;h2 id="superscalar">Superscalar&lt;a class="anchor" href="#superscalar">#&lt;/a>&lt;/h2>
&lt;p>Dynamically issue multiple instructions in each cycle to increase IPC.&lt;/p>
&lt;ul>
&lt;li>Normally need multi-port register files and ALU to avoid structural hazard.&lt;/li>
&lt;li>Can be in-order or out-of-order&lt;/li>
&lt;/ul>
&lt;h2 id="re-order-buffer">Re-order buffer&lt;a class="anchor" href="#re-order-buffer">#&lt;/a>&lt;/h2>
&lt;p>For out-of-order execution CPU architecture, results are put into re-order buffer waiting for commit. These result can be forwarded to later instructions.&lt;/p></description></item><item><title>CPU Performance Test</title><link>https://jimwang99.github.io/posts/architecture/cpu-performance-test/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/architecture/cpu-performance-test/</guid><description>&lt;h2 id="background">Background&lt;a class="anchor" href="#background">#&lt;/a>&lt;/h2>
&lt;p>i7-12700K = Intel Core i7-12700K (8 big cores each has 2 threads, 4 little cores each has 1 thread) running at 5GHz
rpi4 = Raspberry Pi 4 Rev B, with 4x Cortex-A72 running at 1.8GHz (Broadcom BCM2711)
rpi5 = Raspberry Pi 5, with 4x Coretex-A76 running at 2.4GHz (Broadcom BCM2712)
am69 = TI AM69 starter kit, with 8x Cortex-A72 running at 2.0GHz (TI AM69)
lpi4a = LiCheePi 4A, with 4x RISC-V RV64GCV running at 1.85GHz (Alibaba TH1520)&lt;/p></description></item><item><title>How to Evaluate NoC (Network-on-Chip)?</title><link>https://jimwang99.github.io/posts/architecture/how-to-evaluate-noc-network-on-chip/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/architecture/how-to-evaluate-noc-network-on-chip/</guid><description>&lt;p>Modern SoCs heavily relies on NoC to connect interfaces and storage to compute. As the ML models grow larger and larger, the data delivery ability becomes more and more important to overall system performance.&lt;/p>
&lt;p>While everyone is talking about computing resource such as systolic array and vector processors, NoC&amp;rsquo;s critical role sometimes is overseen. In this blog, I&amp;rsquo;m going to talk about the metrics to evaluate NoC in engineering practice.&lt;/p></description></item></channel></rss>