<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Chip Design on When Moore's Law Ends</title><link>https://jimwang99.github.io/posts/chip-design/</link><description>Recent content in Chip Design on When Moore's Law Ends</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Mon, 21 Jan 2019 00:00:00 +0000</lastBuildDate><atom:link href="https://jimwang99.github.io/posts/chip-design/index.xml" rel="self" type="application/rss+xml"/><item><title>Working with Device Tree (DOULOS)</title><link>https://jimwang99.github.io/posts/chip-design/working-with-device-tree-doulos/</link><pubDate>Mon, 21 Jan 2019 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/chip-design/working-with-device-tree-doulos/</guid><description>&lt;h2 id="intro">Intro&lt;a class="anchor" href="#intro">#&lt;/a>&lt;/h2>
&lt;ul>
&lt;li>Device tree: for non-discoverable hardware, included in BSP&lt;/li>
&lt;li>Source type
&lt;ul>
&lt;li>Old style: C code BSP, files compiled into the kernel&lt;/li>
&lt;li>New style: device-tree BSP -&amp;gt; device tree blob (load by boot loader)&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h2 id="compilation">Compilation&lt;a class="anchor" href="#compilation">#&lt;/a>&lt;/h2>
&lt;p>In-tree vs. out-of-tree&lt;/p>
&lt;p>&lt;code>dtc&lt;/code> command&lt;/p>
&lt;ul>
&lt;li>convert .dts to .dtb, and backwards&lt;/li>
&lt;/ul>
&lt;h2 id="device-tree-syntax">Device tree syntax&lt;a class="anchor" href="#device-tree-syntax">#&lt;/a>&lt;/h2>
&lt;ul>
&lt;li>
&lt;p>devicetree.org&lt;/p>
&lt;/li>
&lt;li>
&lt;p>Nodes&lt;/p>
&lt;/li>
&lt;li>
&lt;p>Properties&lt;/p>
&lt;ul>
&lt;li>Values&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>
&lt;p>Root node = &lt;code>/&lt;/code>&lt;/p>
&lt;/li>
&lt;li>
&lt;p>Values format&lt;/p>
&lt;ul>
&lt;li>int, string, list of string, phandle (reference to other node)&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>
&lt;p>Standard properties&lt;/p></description></item><item><title>SystemC Tutorial</title><link>https://jimwang99.github.io/posts/chip-design/systemc-tutorial/</link><pubDate>Fri, 07 Dec 2018 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/chip-design/systemc-tutorial/</guid><description>&lt;pre tabindex="0">&lt;code>// Some simple example

#include &amp;lt;systemc.h&amp;gt;

SC_MODULE (seq_and2 ) { // sequential AND2
 sc_in&amp;lt; sc_uint&amp;lt;8&amp;gt; &amp;gt;		a;
 sc_in&amp;lt; sc_unit&amp;lt;8&amp;gt; &amp;gt;		b;
 sc_out&amp;lt; sc_uint&amp;lt;8&amp;gt; &amp;gt;	f;
 sc_in&amp;lt;bool&amp;gt;				clk;

 void func() {
 f.write( a.read() &amp;amp; b.read() );
 }

 SC_CTOR ( seq_and2 ) {
 SC_CTHREAD(func);
 sensitive &amp;lt;&amp;lt; clk.neg();
 }
}&lt;/code>&lt;/pre>&lt;h2 id="port--signal">Port &amp;amp; signal&lt;a class="anchor" href="#port--signal">#&lt;/a>&lt;/h2>
&lt;ul>
&lt;li>Port
&lt;ul>
&lt;li>&lt;code>sc_in&lt;/code> &amp;amp; &lt;code>sc_out&lt;/code>&lt;/li>
&lt;li>&lt;code>.read()&lt;/code> &amp;amp; &lt;code>.write()&lt;/code> functions&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Signal
&lt;ul>
&lt;li>&lt;code>sc_signal&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h2 id="threads">Threads&lt;a class="anchor" href="#threads">#&lt;/a>&lt;/h2>
&lt;ul>
&lt;li>&lt;code>SC_METHOD()&lt;/code>
&lt;ul>
&lt;li>Just like &lt;code>always_comb&lt;/code> in Verilog, but you have to list the sensitive list&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;code>SC_THREAD()&lt;/code>
&lt;ul>
&lt;li>Not commonly used&lt;/li>
&lt;li>Behavior like &lt;code>initial&lt;/code> in Verilog&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;code>SC_CTHREAD(function name, clock sensitive)&lt;/code>
&lt;ul>
&lt;li>Most commonly used&lt;/li>
&lt;li>Only sensitive to clock edge, just like &lt;code>always_ff&lt;/code> in Verilog&lt;/li>
&lt;li>Not limited to one cycle&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;code>sensitive&lt;/code> keyword to define the sensitive list&lt;/li>
&lt;/ul>
&lt;h2 id="datatypes">Datatypes&lt;a class="anchor" href="#datatypes">#&lt;/a>&lt;/h2>
&lt;ul>
&lt;li>
&lt;p>Integers&lt;/p></description></item><item><title>Case Study Clock Divider with Synchronous Reset</title><link>https://jimwang99.github.io/posts/chip-design/case-study-clock-divider-with-synchronous-reset/</link><pubDate>Thu, 25 Oct 2018 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/chip-design/case-study-clock-divider-with-synchronous-reset/</guid><description>&lt;p>When using a counter to divide a clock, don’t reset the counter, especially when you are using synchronous reset. It will make the clock quiet while reset. And if it’s used along with sync reset, then those flip-flop won’t be reset at all.&lt;/p>
&lt;p>But if without reset, the counter will be “X” in simulation.&lt;/p>
&lt;pre tabindex="0">&lt;code>logic [1:0] cntr;

`ifndef SYNTHESIS
initial begin
 cntr = 2&amp;#39;b00;
end
`endif

always_ff @ (posedge clk) begin
	cntr &amp;lt;= cntr + 1;
end&lt;/code>&lt;/pre></description></item><item><title>FPGA Solution for LiDAR Project</title><link>https://jimwang99.github.io/posts/chip-design/fpga-solution-for-lidar-project/</link><pubDate>Tue, 23 Oct 2018 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/chip-design/fpga-solution-for-lidar-project/</guid><description>&lt;ul>
&lt;li>
&lt;p>1-stop solution: Zynq UltraScale+ RFSoC ZCU111 Evaluation Kit (&lt;a href="https://www.xilinx.com/products/boards-and-kits/zcu111.html">https://www.xilinx.com/products/boards-and-kits/zcu111.html&lt;/a>)&lt;/p>
&lt;ul>
&lt;li>Features:
&lt;ul>
&lt;li>XCZU28DR-2FFVG1517E: high-end RFSoC&lt;/li>
&lt;li>12-bit 4GSPS ADC x8, 14-bit 6.5GSPS DAC x8 (all RFSoC has the same type of ADC/DAC, no higher speed ones)&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Pros: 1-stop with everything we need for bench-top demo&lt;/li>
&lt;li>Cons: expensive $9K, need secondary solution for backup; overkill for second step product&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>
&lt;p>FMC daughter board with high-speed ADC&lt;/p>
&lt;ul>
&lt;li>FMC163 (&lt;a href="https://www.abaco.com/products/fmc163-fpga-mezzanine-card">https://www.abaco.com/products/fmc163-fpga-mezzanine-card&lt;/a>)
&lt;ul>
&lt;li>1x 12-bit ADC, 4.0 GSPS at single channel, or 2GSPS at dual channel, LVDS (TI’s ADC12D2000RF)&lt;/li>
&lt;li>1x 14-bit DAC, 5.7 GSPS, LVDS (ADI’s AD9129)&lt;/li>
&lt;li>Question: will it work with our backup dev boards?&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>AD-FMCDAQ2-EBZ (&lt;a href="https://www.analog.com/en/design-center/evaluation-hardware-and-software/evaluation-boards-kits/eval-ad-fmcdaq2-ebz.html">https://www.analog.com/en/design-center/evaluation-hardware-and-software/evaluation-boards-kits/eval-ad-fmcdaq2-ebz.html&lt;/a>)
&lt;ul>
&lt;li>2x 14-bit ADC, 1.0 GSPS, JESD204B (ADI’s AD9680)&lt;/li>
&lt;li>4x 16-bit DAC, 2.8 GSPS, JESD204B (ADI’s AD9144)&lt;/li>
&lt;li>$1495, buy directly from ADI&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>EVAL-FMCDAQ3-EBZ (&lt;a href="https://www.analog.com/en/design-center/evaluation-hardware-and-software/evaluation-boards-kits/eval-fmcdaq3-ebz.html">https://www.analog.com/en/design-center/evaluation-hardware-and-software/evaluation-boards-kits/eval-fmcdaq3-ebz.html&lt;/a>)
&lt;ul>
&lt;li>2x 14-bit ADC, 1.25 GSPS, JESD204B (ADI’s AD9680)
&lt;ul>
&lt;li>Question: why the same device, here is 1.25G but it’s 1.0G previously&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>2x 16-bit DAC, 2.5 GSPS, JESD204B (ADI’s AD9152)&lt;/li>
&lt;li>$1495, buy directly from ADI
&lt;ul>
&lt;li>ADI provides the whole package, including RTL for FPGA, dev board schematic and etc.&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>ADC12D1800RF Reference board (&lt;a href="http://www.ti.com/tool/ADC12D1800RFRB?keyMatch=adc12d1800rfrb">http://www.ti.com/tool/ADC12D1800RFRB?keyMatch=adc12d1800rfrb&lt;/a>)
&lt;ul>
&lt;li>12-bit, dual 1.8 GSPS or single 3.6 GSPS, LVDS (TI’s ADC12D1800RF)&lt;/li>
&lt;li>$2999, buy directly from TI&lt;/li>
&lt;li>It has a Xilinx Virtex-? on board, but I doubt it can be programmed&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul></description></item><item><title>Case Study Glitch Free Clock Mux</title><link>https://jimwang99.github.io/posts/chip-design/case-study-glitch-free-clock-mux/</link><pubDate>Sat, 31 Mar 2018 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/chip-design/case-study-glitch-free-clock-mux/</guid><description>&lt;p>If your design needs to switch from one clock source to another, there is high possibility of harmful clock glitches while switching. Normally you need to stop this clock during the switching process, but what if you design relies on non-stop clock? Here is the circuit proven to work on silicon.&lt;/p>
&lt;p>&lt;a href="https://www.eetimes.com/document.asp?doc_id=1202359">Techniques to make clock switching glitch free&lt;/a>&lt;/p>
&lt;p>&lt;img src="https://jimwang99.github.io/legacy-media/mahmud3.jpg" alt="img" />&lt;/p>
&lt;p>We used it in our 28nm TSMC HPC+ chip, after we carefully simulated it with HSPICE. If possible, make it a hard macro. But if you cannot do that, ask the back-end engineers to&lt;/p></description></item><item><title>GENUS Training Notes</title><link>https://jimwang99.github.io/posts/chip-design/genus-training-notes/</link><pubDate>Fri, 05 May 2017 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/chip-design/genus-training-notes/</guid><description>&lt;blockquote class='book-hint '>
&lt;p>The following is my notes of GENUS training course on Cadence&amp;rsquo;s training module&lt;/p>&lt;/blockquote>&lt;h2 id="module-03-genus-fundamentals">Module 03: genus fundamentals&lt;a class="anchor" href="#module-03-genus-fundamentals">#&lt;/a>&lt;/h2>
&lt;h3 id="common-ui-vs-legacy-mode">common UI vs legacy mode&lt;a class="anchor" href="#common-ui-vs-legacy-mode">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>unified commands with Tempus&lt;/li>
&lt;li>common us: &lt;code>set_db&lt;/code> &amp;amp; &lt;code>get_db&lt;/code>&lt;/li>
&lt;li>legacy mode: &lt;code>set_attribute&lt;/code> &amp;amp; &lt;code>get_attribute&lt;/code>
&lt;ul>
&lt;li>&lt;code>.synth_init&lt;/code> file: setup info, auto load when start legacy UI, can be skipped with &lt;code>-no_custom&lt;/code> command line option&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="explore-design-hier-in-legacy-ui">explore design hier in legacy UI&lt;a class="anchor" href="#explore-design-hier-in-legacy-ui">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>virtual directory structure
&lt;ul>
&lt;li>&lt;code>/&lt;/code>: root dir
&lt;ul>
&lt;li>designs
&lt;ul>
&lt;li>top_module
&lt;ul>
&lt;li>instances_hier: current module&amp;rsquo;s hier instances&lt;/li>
&lt;li>instances_seq: current module&amp;rsquo;s sequential instances&lt;/li>
&lt;li>instances_cmb: current module&amp;rsquo;s combinational instancs&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>libraries&lt;/li>
&lt;li>hdl_libraries&lt;/li>
&lt;li>flows&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>use &lt;code>find&lt;/code> to locate objects
&lt;ul>
&lt;li>ex. find all the pins &lt;code>find /designs/* -pin *&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>use &lt;code>ls&lt;/code> + &lt;code>cd&lt;/code> to navigate through this virtual directory structure
&lt;ul>
&lt;li>even &lt;code>rm&lt;/code>, &lt;code>mv&lt;/code>, &lt;code>pushd&lt;/code>, &lt;code>popd&lt;/code>&lt;/li>
&lt;li>report all related attributes associated for all the pins: &lt;code>ls -la [find /designs/* -pin *]&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>navigate UNIX disk
&lt;ul>
&lt;li>&lt;code>lpwd&lt;/code>, &lt;code>lcd&lt;/code>, &lt;code>lls&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="attributes">attributes&lt;a class="anchor" href="#attributes">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>&lt;code>set_attribute &amp;lt;attr_name&amp;gt; &amp;lt;value&amp;gt; &amp;lt;object&amp;gt;&lt;/code>&lt;/li>
&lt;li>&lt;code>get_attribute &amp;lt;attr_name&amp;gt; &amp;lt;object&amp;gt;&lt;/code>
&lt;ul>
&lt;li>works on single object only&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>get help
&lt;ul>
&lt;li>&lt;code>get_attribute -h &amp;lt;attr_name&amp;gt; [&amp;lt;object_type&amp;gt;]&lt;/code>
&lt;ul>
&lt;li>get help on attribute&lt;/li>
&lt;li>&lt;code>&amp;lt;attr_name&amp;gt;&lt;/code> can include wildcards&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;code>set_attribute -h&lt;/code>: reports writable attr&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>attr are dependent on the stage of synthesis flow&lt;/li>
&lt;/ul>
&lt;h3 id="input-and-output">input and output&lt;a class="anchor" href="#input-and-output">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>input: RTL + constraint + library + power intent + physical&lt;/li>
&lt;li>output: netlist + LEC dofile + ATPG, scanDEF + constraints + physical design input files&lt;/li>
&lt;/ul>
&lt;h3 id="template-script">template script&lt;a class="anchor" href="#template-script">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>&lt;code>write_template&lt;/code>&lt;/li>
&lt;/ul>
&lt;h3 id="flow">flow&lt;a class="anchor" href="#flow">#&lt;/a>&lt;/h3>
&lt;ol>
&lt;li>setup libraries&lt;/li>
&lt;/ol>
&lt;ul>
&lt;li>&lt;code>set_attribute init_lib_search_path &amp;lt;path&amp;gt; /&lt;/code>&lt;/li>
&lt;li>&lt;code>set_attribute library $ls_lib&lt;/code>&lt;/li>
&lt;li>library domain for low-power design (if not included in CPF)
&lt;ul>
&lt;li>&lt;code>create_library_domain {lib_domain1 lib_domain2}&lt;/code>&lt;/li>
&lt;li>&lt;code>set_attribute library $ls_lib1 lib_domain1 power_domain1&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>dont use
&lt;ul>
&lt;li>&lt;code>set_attribute avoid &amp;lt;1/0&amp;gt; &amp;lt;cell_names&amp;gt;&lt;/code>&lt;/li>
&lt;li>or use &lt;code>set_dont_use &amp;lt;cell_names&amp;gt;&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>(optional) setup physical layout estimation (PLE)
&lt;ul>
&lt;li>dynamically calculates wire delays for different logic structures&lt;/li>
&lt;li>vs Genus-Physical
&lt;ul>
&lt;li>floorplan DEF is optional&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;code>set_attribute lef_library &amp;lt;lef_header&amp;gt;&lt;/code>&lt;/li>
&lt;li>&lt;code>set_attribute qrc_tech_file &amp;lt;qrc_tech_file_path&amp;gt;&lt;/code>&lt;/li>
&lt;li>&lt;code>set_attribute interconnect_mode ple&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;ol start="2">
&lt;li>read HDL&lt;/li>
&lt;/ol>
&lt;ul>
&lt;li>&lt;code>set_attr init_hdl_search_path &amp;lt;path&amp;gt; /&lt;/code>&lt;/li>
&lt;li>&lt;code>read_hdl&lt;/code>&lt;/li>
&lt;/ul>
&lt;ol start="3">
&lt;li>elaborate&lt;/li>
&lt;/ol>
&lt;ul>
&lt;li>what
&lt;ul>
&lt;li>build data structure, infer registers&lt;/li>
&lt;li>high-level HDL opt, remove dead code&lt;/li>
&lt;li>identify clock gating candidates&lt;/li>
&lt;li>overwrite parameters for diff modules&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;code>elaborate&lt;/code>
&lt;ul>
&lt;li>after elaboration, the &lt;code>/designs&lt;/code> is populated&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;code>check_design -all&lt;/code>
&lt;ul>
&lt;li>must: unresolved references&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;ol start="4">
&lt;li>read constraints&lt;/li>
&lt;/ol>
&lt;ul>
&lt;li>&lt;code>read_sdc&lt;/code> (preferred)&lt;/li>
&lt;li>&lt;code>echo $::dc::sdc_failed_commands &amp;gt; failed.sdc&lt;/code>&lt;/li>
&lt;li>&lt;code>check_timing_intent -verbose&lt;/code>
&lt;ul>
&lt;li>check failed commands and errors&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;ol start="5">
&lt;li>opt directives&lt;/li>
&lt;/ol>
&lt;ul>
&lt;li>preserve instances and subdesign (dont touch)
&lt;ul>
&lt;li>&lt;code>set_attr preserve &amp;lt;option&amp;gt;&lt;/code> (more options than &lt;code>set_dont_touch&lt;/code>)
&lt;ul>
&lt;li>false/true&lt;/li>
&lt;li>delete_ok&lt;/li>
&lt;li>const_prop_delete_ok&lt;/li>
&lt;li>const_prop_size_delete_ok&lt;/li>
&lt;li>size_ok&lt;/li>
&lt;li>map_size_ok&lt;/li>
&lt;li>size_delete_ok&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>grouping/ungrouping hierarchy
&lt;ul>
&lt;li>&lt;code>group -group_name &amp;lt;name&amp;gt; &amp;lt;ls_inst&amp;gt;&lt;/code>&lt;/li>
&lt;li>&lt;code>ungroup &amp;lt;hier&amp;gt;&lt;/code>&lt;/li>
&lt;li>disable ungrouping by &lt;code>set_attr ungroup_ok false &amp;lt;inst&amp;gt;&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>boundary opt (default performed)
&lt;ul>
&lt;li>disable by &lt;code>set_attr boundary_opto false &amp;lt;sub_design&amp;gt;&lt;/code>&lt;/li>
&lt;li>use dynamic hierarchical check to verify boundary opt in conformal LEC&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>opt sequential logic (default performed)
&lt;ul>
&lt;li>remove unused flops that is not driving an output port&lt;/li>
&lt;li>disable by
&lt;ul>
&lt;li>&lt;code>set_attr hdl_preserve_unused_register true /&lt;/code>&lt;/li>
&lt;li>&lt;code>set_attr delete_unloaded_seqs false /&lt;/code>&lt;/li>
&lt;li>&lt;code>set_attr optimize_constant_0_flops false /&lt;/code>&lt;/li>
&lt;li>&lt;code>set_attr optimize_constant_1_flops false /&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>same thing to combinational logic that drives unloaded hier pins
&lt;ul>
&lt;li>disable by &lt;code>set_attr prune_unused_logic false &amp;lt;pins&amp;gt;&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>merge sequential logic (default performed)
&lt;ul>
&lt;li>combine flops and latches that are equivalent in the same hierarchy&lt;/li>
&lt;li>disable by
&lt;ul>
&lt;li>&lt;code>set_attr optimize_merge_flops false /&lt;/code>&lt;/li>
&lt;li>&lt;code>set_attr optimize_merge_latches false /&lt;/code>&lt;/li>
&lt;li>&lt;code>set_attr optimize_merge_seq false &amp;lt;inst&amp;gt;&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>multibit cell inference (MBCI)
&lt;ul>
&lt;li>flops/tri-state cell/MUX/inverters/…&lt;/li>
&lt;li>share clock to reduce power/improve reliability&lt;/li>
&lt;li>LEC support&lt;/li>
&lt;li>can control naming style (for verification)&lt;/li>
&lt;li>&lt;code>set_attr use_multibit_cells true&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>other opt
&lt;ul>
&lt;li>opt async reset logic
&lt;ul>
&lt;li>&lt;code>set_attr time_recovery_arcs true /&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>auto ungrouping
&lt;ul>
&lt;li>&lt;code>set_attr auto_ungroup {none | both}&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>keep the synchronous feedback logic immediately in front of the sequential elements (?)
&lt;ul>
&lt;li>&lt;code>set_attr hdl_ff_keep_feedback&lt;/code>&lt;/li>
&lt;li>affect how enable logic of a flop is implemented&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>opt TNS other than WNS
&lt;ul>
&lt;li>&lt;code>set_attr tns_opto true /&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;ol start="6">
&lt;li>synthesis&lt;/li>
&lt;/ol>
&lt;ul>
&lt;li>1st level: &lt;code>syn_generic &amp;lt;-physical&amp;gt;&lt;/code>
&lt;ul>
&lt;li>tech independent RTL opt
&lt;ul>
&lt;li>can skip for netlist-to-netlist synthesis&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;code>set_attr syn_generic_effort&lt;/code>
&lt;ul>
&lt;li>medium by default&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>2nd level: &lt;code>syn_map &amp;lt;-physical&amp;gt;&lt;/code>
&lt;ul>
&lt;li>mapping to lib, and logic opt
&lt;ul>
&lt;li>initial structuring
&lt;ul>
&lt;li>constant propagation, clock gating&lt;/li>
&lt;li>structuring for best delay&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>target info
&lt;ul>
&lt;li>estimate timing&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>global mapping
&lt;ul>
&lt;li>mapping to meet target&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>global incremental
&lt;ul>
&lt;li>net/drive opt&lt;/li>
&lt;li>timing tuning&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;code>set_attr syn_map_effort&lt;/code>
&lt;ul>
&lt;li>high by default&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>check the slack, if too negative, check the constraint/design&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>3rd level: &lt;code>syn_opt &amp;lt;-physical&amp;gt; &amp;lt;-spatial&amp;gt; &amp;lt;-incr&amp;gt;&lt;/code>
&lt;ul>
&lt;li>opt gates
&lt;ul>
&lt;li>fix drc, cleanup area, cleanup timing&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;code>set_attr syn_opt_effort&lt;/code>
&lt;ul>
&lt;li>high by default&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>global effort
&lt;ul>
&lt;li>&lt;code>set_attr syn_global_effort&lt;/code>&lt;/li>
&lt;li>set to express while explore flow
&lt;ul>
&lt;li>accept not clean design&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;ol start="7">
&lt;li>analyze and report&lt;/li>
&lt;/ol>
&lt;ul>
&lt;li>after elaboration
&lt;ul>
&lt;li>&lt;code>check_design unresolved&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>constraint
&lt;ul>
&lt;li>&lt;code>check_timing_intent&lt;/code>&lt;/li>
&lt;li>use Conformal Constraint Designer (CCD) tool to validate timing constraint
&lt;ul>
&lt;li>&lt;code>write_to_ccd validate -sdc &amp;gt; dofile&lt;/code> generate dofile used in CCD&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>check &lt;code>preserve&lt;/code> attributes, remove those that are not needed&lt;/li>
&lt;li>ungrouping small blocks can improve timing/area&lt;/li>
&lt;li>reports
&lt;ul>
&lt;li>report_area&lt;/li>
&lt;li>report_dp (datapath)&lt;/li>
&lt;li>report_design_rules (drc)&lt;/li>
&lt;li>report_messages&lt;/li>
&lt;li>report_power&lt;/li>
&lt;li>report_qor&lt;/li>
&lt;li>report_timing&lt;/li>
&lt;li>report_summary&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>from GUI
&lt;ul>
&lt;li>timing -&amp;gt; timing lint: gives a thorough&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;ol start="8">
&lt;li>gen outputs&lt;/li>
&lt;/ol>
&lt;ul>
&lt;li>&lt;code>write_hdl &amp;gt; filename&lt;/code>&lt;/li>
&lt;li>&lt;code>write_sdc &amp;gt; filename&lt;/code>&lt;/li>
&lt;li>&lt;code>write_design -innovus&lt;/code>&lt;/li>
&lt;/ul>
&lt;h3 id="command-help">command help&lt;a class="anchor" href="#command-help">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>&lt;code>setenv MANPATH $CDN_SYNTH_ROOT/share/synth/man&lt;/code> to view man pages from UNIX shell&lt;/li>
&lt;/ul>
&lt;h2 id="module-04-datapath">Module 04: datapath&lt;a class="anchor" href="#module-04-datapath">#&lt;/a>&lt;/h2>
&lt;h3 id="datapath-info-in-virtual-file-system">datapath info in virtual file system&lt;a class="anchor" href="#datapath-info-in-virtual-file-system">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>/hdl_libraries/
&lt;ul>
&lt;li>/hdl_libraries/CW (chipware)&lt;/li>
&lt;li>/hdl_libraries/DW (designware)&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="datapath-operation">datapath operation&lt;a class="anchor" href="#datapath-operation">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>architecture selection&lt;/li>
&lt;li>sharing and speculation (unsharing)&lt;/li>
&lt;li>carry-save arithmetic (CSA)&lt;/li>
&lt;li>…&lt;/li>
&lt;/ul>
&lt;h3 id="datapath-directives">datapath directives&lt;a class="anchor" href="#datapath-directives">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>CSA
&lt;ul>
&lt;li>&lt;code>set_attr dp_csa {inherited|basic|none} &amp;lt;design&amp;gt;&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>sharing and speculation
&lt;ul>
&lt;li>sharing: improve area&lt;/li>
&lt;li>&lt;code>set_attr dp_sharing&lt;/code>&lt;/li>
&lt;li>&lt;code>set_attr dp_speculation&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>arch selection
&lt;ul>
&lt;li>manually control datapath arch selection (not recommended)
&lt;ul>
&lt;li>&lt;code>set_attr user_speed_grade [find /designs* -subdesign &amp;lt;name&amp;gt;]&lt;/code> while speed can be ver_fast|fast|medium|slow|very_slow&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>reordering (reorder input to opt critical path)&lt;/li>
&lt;li>ChipWare (CW)
&lt;ul>
&lt;li>also maps DesignWare components in RTL to CW&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="opt-in-syn_generic">opt in syn_generic&lt;a class="anchor" href="#opt-in-syn_generic">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>constant propagation&lt;/li>
&lt;li>resource sharing&lt;/li>
&lt;li>logic speculation&lt;/li>
&lt;li>MUX opt&lt;/li>
&lt;li>CSA opt&lt;/li>
&lt;li>datapath rewriting
&lt;ul>
&lt;li>QoR driven RTL code rewrite&lt;/li>
&lt;li>by default during &lt;code>syn_generic&lt;/code> with high effort level&lt;/li>
&lt;li>no LEC impact&lt;/li>
&lt;li>ex.&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;pre tabindex="0">&lt;code>assign p = a - b;
assign q = a + b;
assign y = s ? p : q;

# better timing, smaller area
assign t = {16{s}} ^ b;
assign y = a + t + s;&lt;/code>&lt;/pre>&lt;h3 id="report">report&lt;a class="anchor" href="#report">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>&lt;code>set_attr hdl_track_filename_row_col true /&lt;/code> before &lt;code>read_hdl&lt;/code>&lt;/li>
&lt;li>&lt;code>report_dp&lt;/code> after every stages: elaboration/syn_gen/syn_map/syn_opt to track datapath components changes&lt;/li>
&lt;/ul>
&lt;h2 id="module-05-debug-design-scenarios">Module 05: debug design scenarios&lt;a class="anchor" href="#module-05-debug-design-scenarios">#&lt;/a>&lt;/h2>
&lt;h3 id="problem-with-sdc">problem with sdc&lt;a class="anchor" href="#problem-with-sdc">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>check the log file for errors and warnings&lt;/li>
&lt;li>check constraint consistency by &lt;code>check_timing_intent -verbose&lt;/code> before synthesis&lt;/li>
&lt;/ul>
&lt;h3 id="path-grouping">path grouping&lt;a class="anchor" href="#path-grouping">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>cost group: opt cost groups simultaneously according to their weight, to minimize their WNS for each group&lt;/li>
&lt;li>path group -&amp;gt; cost group&lt;/li>
&lt;/ul>
&lt;h3 id="tightenrelax-constraint">tighten/relax constraint&lt;a class="anchor" href="#tightenrelax-constraint">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>emphasize some paths in opt without impacting output SDC&lt;/li>
&lt;li>&lt;code>path_adjust -from &amp;lt;obj&amp;gt; -to &amp;lt;obj&amp;gt; -delay &amp;lt;delta_slack_ps&amp;gt;&lt;/code>
&lt;ul>
&lt;li>if delta_slack_ps &amp;lt; 0, tighten the path&lt;/li>
&lt;li>if delta_slack_ps &amp;gt; 0, loosen the path&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>use &lt;code>rm [find /des* -exceptions pa_*]&lt;/code> before report timing to get normal timing reports
&lt;ul>
&lt;li>the adjustment will be in the timing report if not removed&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="bottom-up-design-flow">bottom-up design flow&lt;a class="anchor" href="#bottom-up-design-flow">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>promote submodule
&lt;ul>
&lt;li>&lt;code>create_derived_design&lt;/code> promote submodule to top-level module&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h2 id="module-06-physical-synthesis">Module 06: physical synthesis&lt;a class="anchor" href="#module-06-physical-synthesis">#&lt;/a>&lt;/h2>
&lt;h3 id="why">why?&lt;a class="anchor" href="#why">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>for synthesis: all wires of fanout=n are the same&lt;/li>
&lt;li>for physical: each wire is unique
&lt;ul>
&lt;li>80% to 90% of wires are local, the rest are big problems&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>old tricks don&amp;rsquo;t work: over-constraint&lt;/li>
&lt;/ul>
&lt;h3 id="how">how?&lt;a class="anchor" href="#how">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>incremental congestion prevetion&lt;/li>
&lt;li>structural datapath&lt;/li>
&lt;li>physical aware clock gating/logic structuring/mapping&lt;/li>
&lt;li>use floorplan as bridge to close pre and post layout gap
&lt;ul>
&lt;li>def file: must define die size; macro locations, fences/guides/regions are better to have (impact timing)&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>genus vs innovus: 5% timing &amp;amp; wirelength diff&lt;/li>
&lt;/ul>
&lt;h3 id="spatial-flow">spatial flow&lt;a class="anchor" href="#spatial-flow">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>if backend is going to run full place_opt, instead of &lt;code>place_opt -incr&lt;/code> with genus-physical outputs as inputs, then no need to waste time on the final syn_opt stage&lt;/li>
&lt;li>use &lt;code>syn_opt -spatial&lt;/code> instead of &lt;code>syn_opt -physical&lt;/code>&lt;/li>
&lt;/ul>
&lt;h3 id="pam-physical-aware-mapping--pas-physical-aware-structuring">PAM (physical-aware mapping) &amp;amp; PAS (physical-aware structuring)&lt;a class="anchor" href="#pam-physical-aware-mapping--pas-physical-aware-structuring">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>automatically turned on with &lt;code>-physical&lt;/code>&lt;/li>
&lt;/ul>
&lt;h3 id="useful-attributes">useful attributes&lt;a class="anchor" href="#useful-attributes">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>&lt;code>invs_enable_useful_skew&lt;/code>&lt;/li>
&lt;li>&lt;code>phys_ignore_nets&lt;/code>&lt;/li>
&lt;li>&lt;code>pqos_ignore_msv&lt;/code>
&lt;ul>
&lt;li>whether to pass lib or power domain info to INVS&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;code>invs_user_constraint_file&lt;/code>
&lt;ul>
&lt;li>sourced during INVS session&lt;/li>
&lt;li>&lt;code>invs_preload_script&lt;/code> &amp;amp; &lt;code>invs_postload_script&lt;/code> &amp;amp; &lt;code>invs_preexport_script&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;code>number_of_routing_layers&lt;/code>
&lt;ul>
&lt;li>important to have&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;code>invs_pre_place_opt&lt;/code>&lt;/li>
&lt;li>&lt;code>pqos_placement_effort&lt;/code>
&lt;ul>
&lt;li>congestion effort&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;code>invs_gzip_interface_file&lt;/code>&lt;/li>
&lt;li>&lt;code>invs_temp_dir&lt;/code>&lt;/li>
&lt;/ul>
&lt;h3 id="correlation-between-genus-phys-and-invs">correlation between genus-phys and invs&lt;a class="anchor" href="#correlation-between-genus-phys-and-invs">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>ensure NDR and layer-promotion info is passed to innvous&lt;/li>
&lt;li>assure wirelength has good correlation&lt;/li>
&lt;/ul>
&lt;h3 id="early-stage-physical-analysis">early stage physical analysis&lt;a class="anchor" href="#early-stage-physical-analysis">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>at generic physical synthesis stage&lt;/li>
&lt;li>why?
&lt;ul>
&lt;li>analyze hier&lt;/li>
&lt;li>hard macro locations&lt;/li>
&lt;li>floor plan constraints&lt;/li>
&lt;li>timing debug with gui&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="check-placement-legality">check placement legality&lt;a class="anchor" href="#check-placement-legality">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>&lt;code>check_placement&lt;/code>&lt;/li>
&lt;/ul>
&lt;h3 id="edit-floorplan-in-genus-gui">edit floorplan in Genus GUI&lt;a class="anchor" href="#edit-floorplan-in-genus-gui">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>go into edit mode&lt;/li>
&lt;/ul>
&lt;h3 id="report-1">report&lt;a class="anchor" href="#report-1">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>&lt;code>write_report&lt;/code>
&lt;ul>
&lt;li>wrote QoS statistics&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;code>report_summary&lt;/code>
&lt;ul>
&lt;li>write summary table&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;code>write_snapshot&lt;/code>
&lt;ul>
&lt;li>design database and reports&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="faq">FAQ&lt;a class="anchor" href="#faq">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>recommended flow
&lt;ul>
&lt;li>after synthesis with physical, &lt;code>write_design -innovus&lt;/code>&lt;/li>
&lt;li>then in innovus load the output data, and &lt;code>place_opt_design -incr&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>what is under the hood of &lt;code>syn_opt -phy&lt;/code>?
&lt;ul>
&lt;li>it calls &lt;code>place_opt -phy_syn&lt;/code> in INVS, and load back the result and do low effort TNS/WNS opt&lt;/li>
&lt;li>so the engine is the same between genus-phy and invs&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>is it possible to do CTS in genus?
&lt;ul>
&lt;li>No. but simple CTS will be enabled in coming versions&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="debug-with-common-ui">debug with common ui&lt;a class="anchor" href="#debug-with-common-ui">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>timing debug
&lt;ul>
&lt;li>timing -&amp;gt; debug timing&lt;/li>
&lt;li>diff path groups histogram&lt;/li>
&lt;li>highlight violating path&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h2 id="module-07-low-power-opt">Module 07: low power opt&lt;a class="anchor" href="#module-07-low-power-opt">#&lt;/a>&lt;/h2>
&lt;ul>
&lt;li>low power opt impacts timing a lot
&lt;ul>
&lt;li>trade-off&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="flow-1">flow&lt;a class="anchor" href="#flow-1">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>enable clock gating&lt;/li>
&lt;li>annotate switching activities with TCF/SAIF/VCD&lt;/li>
&lt;li>apply clock-gating directives&lt;/li>
&lt;li>apply leakage/dynamic power constraints&lt;/li>
&lt;li>synthesis with clock gating insertion/power opt&lt;/li>
&lt;li>analyze&lt;/li>
&lt;/ul>
&lt;h3 id="multi-vth-lib">multi-Vth lib&lt;a class="anchor" href="#multi-vth-lib">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>low VT on timing critical path, high VT on non-critical path&lt;/li>
&lt;/ul>
&lt;h3 id="clock-gating">clock gating&lt;a class="anchor" href="#clock-gating">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>&lt;code>set attr lp_insert_clock_taing true /&lt;/code>&lt;/li>
&lt;li>specify clock gating cell
&lt;ul>
&lt;li>customied: &lt;code>lp_clock_gating_module&lt;/code> attr&lt;/li>
&lt;li>select from library: &lt;code>lp_clock_gating_cell&lt;/code> attr&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>disable clock gating: &lt;code>lp_clock_gating_exclude&lt;/code>&lt;/li>
&lt;li>control fanout of CGC: &lt;code>lp_clock_gating_*_flops&lt;/code>&lt;/li>
&lt;li>common enable: &lt;code>lp_clock_gating_extract_common_enable&lt;/code>&lt;/li>
&lt;li>clock gating for sync reset&lt;/li>
&lt;/ul>
&lt;h3 id="backannotate-switching-activity">backannotate switching activity&lt;a class="anchor" href="#backannotate-switching-activity">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>&lt;code>read_tcf&lt;/code> (toggle count format)&lt;/li>
&lt;li>&lt;code>read_saif&lt;/code> (converted to TCF internally)&lt;/li>
&lt;li>&lt;code>read_vcd&lt;/code>&lt;/li>
&lt;li>manipulate activity with &lt;code>lp_toggle_*&lt;/code> attr&lt;/li>
&lt;/ul>
&lt;h3 id="joules-rtl-power-estimation">Joules: RTL power estimation&lt;a class="anchor" href="#joules-rtl-power-estimation">#&lt;/a>&lt;/h3>
&lt;h3 id="effort">effort&lt;a class="anchor" href="#effort">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>&lt;code>leakage_power_effort&lt;/code> attr
&lt;ul>
&lt;li>{none | low | high}&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>disable leakage power opt
&lt;ul>
&lt;li>&lt;code>max_leakage_power&lt;/code> must not be set while &lt;code>leakge_power_effort&lt;/code> set to none&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>dynamic vs leakage
&lt;ul>
&lt;li>&lt;code>lp_power_optimization_weight&lt;/code> attr: power = weight * leakage + (1 - weight) * dynamic
&lt;ul>
&lt;li>normally, weight is close to 1&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>POPT-501&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="report-2">report&lt;a class="anchor" href="#report-2">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>&lt;code>report_clock_gating&lt;/code>&lt;/li>
&lt;li>&lt;code>report_power&lt;/code>&lt;/li>
&lt;li>get power-related info
&lt;ul>
&lt;li>&lt;code>lp_internal/leakage/net_power&lt;/code>&lt;/li>
&lt;li>&lt;code>lp_default_toggle_rate&lt;/code>&lt;/li>
&lt;li>&lt;code>lp_default_probability&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="useful-attr">useful attr&lt;a class="anchor" href="#useful-attr">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>&lt;code>lp_clock_gating_exceptions_aware&lt;/code>&lt;/li>
&lt;li>&lt;code>declone/share/split/merge_clock_gate&lt;/code>&lt;/li>
&lt;/ul>
&lt;h2 id="module-08-design-for-test">Module 08: design for test&lt;a class="anchor" href="#module-08-design-for-test">#&lt;/a>&lt;/h2>
&lt;h3 id="flow-2">flow&lt;a class="anchor" href="#flow-2">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>setup DFT rule, and check
&lt;ul>
&lt;li>shift enable&lt;/li>
&lt;li>test mode&lt;/li>
&lt;li>prevent scan mapping of flops&lt;/li>
&lt;li>internal clock as test clock&lt;/li>
&lt;li>DFT controllable constraints&lt;/li>
&lt;li>abstract scan segment&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>add test logic
&lt;ul>
&lt;li>insert test point&lt;/li>
&lt;li>insert shadow logic&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>synthesis&lt;/li>
&lt;li>setup DFT config, and preview scan chains
&lt;ul>
&lt;li>scan chain: number, length&lt;/li>
&lt;li>control data lockup elements&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>connect scan chains&lt;/li>
&lt;li>incremental opt&lt;/li>
&lt;/ul>
&lt;h3 id="dft-in-virtual-file-structure">DFT in virtual file structure&lt;a class="anchor" href="#dft-in-virtual-file-structure">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>/designs/dft&lt;/li>
&lt;/ul>
&lt;h3 id="dft-constraint">DFT constraint&lt;a class="anchor" href="#dft-constraint">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>2 scan styles: controlled by &lt;code>dft_scan_style&lt;/code> attr&lt;/li>
&lt;/ul>
&lt;ol>
&lt;li>muxed style (muxed_scan) (most commonly used)&lt;/li>
&lt;li>clocked LSSD (clocked_lssd_scan) (1 system clock, and 2 scan clocks)&lt;/li>
&lt;/ol>
&lt;ul>
&lt;li>define shift enable signal
&lt;ul>
&lt;li>for muxed style: &lt;code>define_shift_enable&lt;/code>
&lt;ul>
&lt;li>default one for common usage, or each chain has its own enable signal&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>for LSSD style: &lt;code>define_lssd_scan_clock_a/b&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>define test mode signal: &lt;code>define_test_mode&lt;/code>
&lt;ul>
&lt;li>put circuit in test mode so that gated clocks are all activated&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>define test clock domains: &lt;code>define_test_clock -name &amp;lt;name&amp;gt; -domain &amp;lt;domain&amp;gt; &amp;lt;pin_name&amp;gt;&lt;/code>
&lt;ul>
&lt;li>due to unbalanced clock tree, create separate test clock domains to prevent timing issues&lt;/li>
&lt;li>lock-up latches (auto added) for crossing test clocks in the same domain, if more than 1 test clocks are defined in one domain&lt;/li>
&lt;li>by default, in the same test clock domain use the same clock edge (controlled by &lt;code>dft_mix_clock_edges_in_scan_chain&lt;/code> attr)&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>define scan segment
&lt;ul>
&lt;li>&lt;code>define_scan_abstract/fixed/floating/preserved_segment&lt;/code>&lt;/li>
&lt;li>&lt;code>define_scan_shift_register_segment&lt;/code>&lt;/li>
&lt;li>&lt;code>define_jtag_boundary_scan_segment&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>preserve nonscan flops
&lt;ul>
&lt;li>set &lt;code>dft_scan_map_mode&lt;/code> attr to preserve&lt;/li>
&lt;li>set &lt;code>dft_dont_scan&lt;/code> attr to true&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>control the length and number
&lt;ul>
&lt;li>by default, no max length for scan chain&lt;/li>
&lt;li>&lt;code>dft_min_number_of_scan_chains&lt;/code>&lt;/li>
&lt;li>&lt;code>dft_max_length_of_scan_chains&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="dft-rule-check">DFT rule check&lt;a class="anchor" href="#dft-rule-check">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>uncontrollable clock nets&lt;/li>
&lt;li>uncontrollable async set/reset nets&lt;/li>
&lt;li>conflicting clock and async set/reset net&lt;/li>
&lt;li>shift register rules&lt;/li>
&lt;li>abstract segment rules&lt;/li>
&lt;li>&lt;code>check_dft_rules&lt;/code>&lt;/li>
&lt;li>&lt;code>fix_dft_violations&lt;/code> (only for muxed style)&lt;/li>
&lt;li>&lt;code>check_atpg_rules&lt;/code>
&lt;ul>
&lt;li>only generate script for Modus ATPG rule checker&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;code>check_design&lt;/code>&lt;/li>
&lt;li>&lt;code>analyze_atpg_testability&lt;/code>
&lt;ul>
&lt;li>run Modus&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="add-dft-logic">add DFT logic&lt;a class="anchor" href="#add-dft-logic">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>&lt;code>insert_dft *&lt;/code>&lt;/li>
&lt;li>identify shift register to save area (auto done)
&lt;ul>
&lt;li>cmd = &lt;code>identify_shift_register_scan_segments&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>mapping to scan in a already mapped netlist
&lt;ul>
&lt;li>&lt;code>set_scan_equivalent&lt;/code>: one-to-one correspondence between non-scan and scan flop lib cells&lt;/li>
&lt;li>&lt;code>replace_scan&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="connect-scan-chains">connect scan chains&lt;a class="anchor" href="#connect-scan-chains">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>&lt;code>connect_scan_chains&lt;/code>&lt;/li>
&lt;/ul>
&lt;h3 id="report-and-output">report and output&lt;a class="anchor" href="#report-and-output">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>&lt;code>report_scan_chains&lt;/code>&lt;/li>
&lt;li>&lt;code>report_scan_setup&lt;/code>&lt;/li>
&lt;li>&lt;code>write_scandef&lt;/code>&lt;/li>
&lt;li>&lt;code>write_dft_atpg*&lt;/code>: interface to ATPG tool&lt;/li>
&lt;li>&lt;code>write_dft_abstract_model&lt;/code>&lt;/li>
&lt;/ul>
&lt;h3 id="bottom-up-scan-flow">bottom-up scan flow&lt;a class="anchor" href="#bottom-up-scan-flow">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>block level
&lt;ul>
&lt;li>create block level chains&lt;/li>
&lt;li>&lt;code>write_hdl -abstract&lt;/code>&lt;/li>
&lt;li>&lt;code>write_dft_abstract_model&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>top level
&lt;ul>
&lt;li>&lt;code>read_dft_abstract_model&lt;/code>&lt;/li>
&lt;li>&lt;code>connect_scan_chains&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h2 id="module-09-lec">Module 09: LEC&lt;a class="anchor" href="#module-09-lec">#&lt;/a>&lt;/h2>
&lt;h3 id="guidance-to-address-formal-verification-challenge">guidance to address formal verification challenge&lt;a class="anchor" href="#guidance-to-address-formal-verification-challenge">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>challenges
&lt;ul>
&lt;li>datapath arch&lt;/li>
&lt;li>ungrouping: no manual random ungrouping&lt;/li>
&lt;li>boundary opt&lt;/li>
&lt;li>phase inversion&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>long run-time, werid mismatch&lt;/li>
&lt;/ul>
&lt;h3 id="recommended-2-step-verification">recommended 2-step verification&lt;a class="anchor" href="#recommended-2-step-verification">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>1st-step: synthesis with preserved datapath modules/hier, restrict certain opt, min ungrouping, and output intermediate gate netlist&lt;/li>
&lt;li>2nd-step: incremental synthesis with additional opt and ungrouping, and output final gate netlist&lt;/li>
&lt;li>compare: RTL vs intermediate netlist, then intermediate netlist vs final netlist&lt;/li>
&lt;/ul>
&lt;h3 id="cmd">cmd&lt;a class="anchor" href="#cmd">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>&lt;code>write_lec_script -revised_design inter.v&lt;/code>&lt;/li>
&lt;li>&lt;code>write_lec_script -revised_design final.v -golden_design inter.v&lt;/code>&lt;/li>
&lt;/ul>
&lt;h3 id="attr-affects-formal-verification">attr affects formal verification&lt;a class="anchor" href="#attr-affects-formal-verification">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>datapath: &lt;code>dp_*&lt;/code>&lt;/li>
&lt;li>boundary opt&lt;/li>
&lt;li>ungrouping&lt;/li>
&lt;li>retime&lt;/li>
&lt;li>&lt;code>wlec_*&lt;/code>&lt;/li>
&lt;/ul>
&lt;h3 id="in-lec">in LEC&lt;a class="anchor" href="#in-lec">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>&lt;code>analyze datapath&lt;/code>: to analyze datapath modules&lt;/li>
&lt;li>&lt;code>analyze abort -compare -thread 4&lt;/code>: multithreading abort resolving&lt;/li>
&lt;li>module-level datapath analysis (MDP)
&lt;ul>
&lt;li>improve quality&lt;/li>
&lt;li>&lt;code>analyze datapath -module xxx&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h2 id="module-10-interface">Module 10: interface&lt;a class="anchor" href="#module-10-interface">#&lt;/a>&lt;/h2>
&lt;h3 id="netlist">netlist&lt;a class="anchor" href="#netlist">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>possible modifications
&lt;ul>
&lt;li>bit blasted port/constants
&lt;ul>
&lt;li>&lt;code>set_attr write_vlog_bit_blast_mapped_ports true /&lt;/code> and &lt;code>set_attr bit_blasted_port_style %s_%d /&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>name changing: &lt;code>update_names&lt;/code> cmd&lt;/li>
&lt;li>loop breaker: break comb feedback loops&lt;/li>
&lt;li>remove assign statement (not needed in INVS)
&lt;ul>
&lt;li>&lt;code>set_attr remove_assigns true /&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h2 id="appendix">Appendix&lt;a class="anchor" href="#appendix">#&lt;/a>&lt;/h2>
&lt;h3 id="retiming">retiming&lt;a class="anchor" href="#retiming">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>&lt;code>set_attr retime true [find / -subd xxx]&lt;/code>&lt;/li>
&lt;li>&lt;code>retime -prepare -min_delay -effort high [find / -subd xxx]&lt;/code> before &lt;code>syn_gen&lt;/code>&lt;/li>
&lt;/ul>
&lt;h3 id="advanced-low-power-flow">advanced low-power flow&lt;a class="anchor" href="#advanced-low-power-flow">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>CPF&lt;/li>
&lt;li>MSMV&lt;/li>
&lt;/ul>
&lt;h3 id="common-ui">common ui&lt;a class="anchor" href="#common-ui">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>attr
&lt;ul>
&lt;li>set attr: &lt;code>set_db &amp;lt;attr_name&amp;gt; &amp;lt;value&amp;gt; &amp;lt;object&amp;gt;&lt;/code>&lt;/li>
&lt;li>query attr: &lt;code>get_db &amp;lt;attr_name&amp;gt; &amp;lt;object&amp;gt;&lt;/code>&lt;/li>
&lt;li>&lt;code>help *clock* -attribute&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>virtual directory structure
&lt;ul>
&lt;li>&lt;code>vcd&lt;/code>&lt;/li>
&lt;li>&lt;code>vls&lt;/code>&lt;/li>
&lt;li>&lt;code>rename_obj&lt;/code>&lt;/li>
&lt;li>&lt;code>vpopd&lt;/code>&lt;/li>
&lt;li>&lt;code>vpushd&lt;/code>&lt;/li>
&lt;li>&lt;code>delete_obj&lt;/code>&lt;/li>
&lt;li>&lt;code>vfind&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>examples
&lt;ul>
&lt;li>find all designs: &lt;code>get_db .designs&lt;/code>&lt;/li>
&lt;li>find all comb leaf inst under current directory: &lt;code>get_db . .insts -if .is_comb&lt;/code>&lt;/li>
&lt;li>find all inst of a certain cell type: &lt;code>get_db insts -if {.base_cell.name == DFFX1}&lt;/code>&lt;/li>
&lt;li>calc leakage power of a hier: &lt;code>expr [join [get_db hinst:CORE/ALU .insts.leakage_power] +]&lt;/code>&lt;/li>
&lt;li>fanout histogram&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;pre tabindex="0">&lt;code>set tot [llength [get_db nets]]
for {set i 0} {$i &amp;lt;= 100} {incr i 5} {
 set n [llength [get_db nets -if &amp;#34;.num_loads&amp;gt;$i &amp;amp;&amp;amp; .num_loads&amp;lt;[expr {$i+5}]&amp;#34;]]
 puts [string report &amp;#34;#&amp;#34; [expr $n * 100 / $tot]]
}&lt;/code>&lt;/pre>&lt;pre tabindex="0">&lt;code>- find all pins: `vls -la [vfind /designs/* -pin *]`&lt;/code>&lt;/pre>&lt;ul>
&lt;li>MMMC setup flow
&lt;ul>
&lt;li>&lt;code>read_mmmc&lt;/code>&lt;/li>
&lt;li>&lt;code>read_physical -lef&lt;/code>&lt;/li>
&lt;li>&lt;code>read_hdl&lt;/code>&lt;/li>
&lt;li>&lt;code>elab&lt;/code>&lt;/li>
&lt;li>&lt;code>read_def&lt;/code>&lt;/li>
&lt;li>&lt;code>read_power_intent&lt;/code>&lt;/li>
&lt;li>&lt;code>init_design -skip_sdc_read&lt;/code>&lt;/li>
&lt;li>&lt;code>syn_gen/map/opt&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="clipper-flow">clipper flow&lt;a class="anchor" href="#clipper-flow">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>block level physical synthesis &amp;lt;-&amp;gt; unit level physical synthesis
&lt;ul>
&lt;li>unit level cannot understand block level&amp;rsquo;s congestion and physical context issues&lt;/li>
&lt;li>so pass timing/physical context DEF and constraint from block level to unit level&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>CMD
&lt;ul>
&lt;li>&lt;code>create_clip&lt;/code> at higher level
&lt;ul>
&lt;li>block boundary must be preserved (remember, genus is very aggressive about optimizing)&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;code>read_clip&lt;/code> at lower level&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h2 id="advanced-synthesis">Advanced Synthesis&lt;a class="anchor" href="#advanced-synthesis">#&lt;/a>&lt;/h2></description></item><item><title>Case Study Clock Skew Control</title><link>https://jimwang99.github.io/posts/chip-design/case-study-clock-skew-control/</link><pubDate>Tue, 18 Apr 2017 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/chip-design/case-study-clock-skew-control/</guid><description>&lt;p>Question: how to control the clock skew between a group of clocks to be minimum, say less than 30ps, instead of utilizing useful skew? This case happens to our hard macros.&lt;/p>
&lt;p>A: in Innovus, use skew group&lt;/p>
&lt;pre tabindex="0">&lt;code>set min_skew_group {
 path/to/clock/NLVB_CKB
 path/to/clock/NLVA_CKB
 path/to/clock/NLVP_CKB
}

create_ccopt_skew_group \
 -name min_skew_group \
 -sources path/to/clock/source/CKB \
 -sinks $min_skew_group \
 -target_insertion_delay 0.500 \
 -rank 1
 -target_skew 0.000

set_ccopt_property constraints -skew_group min_skew_group ccopt&lt;/code>&lt;/pre></description></item><item><title>INNOVUS Training Notes</title><link>https://jimwang99.github.io/posts/chip-design/innovus-training-notes/</link><pubDate>Sat, 01 Apr 2017 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/chip-design/innovus-training-notes/</guid><description>&lt;blockquote class='book-hint '>
&lt;p>The following is my notes of INNOVUS training course on Cadence&amp;rsquo;s training module&lt;/p>&lt;/blockquote>&lt;h2 id="module-02-overview">Module 02: overview&lt;a class="anchor" href="#module-02-overview">#&lt;/a>&lt;/h2>
&lt;ul>
&lt;li>“gift” directory contains lots of useful scripts to help productivity&lt;/li>
&lt;li>Independent “viewlog” utility or “Tools-&amp;gt;Log Viewer” will start a GUI to help understand log files better.&lt;/li>
&lt;li>Batch mode: &lt;code>innovus -no_gui -init batch.tcl&lt;/code>
&lt;ul>
&lt;li>&lt;code>win&lt;/code> / &lt;code>win off&lt;/code> to show/hide GUI&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h2 id="module-03-import-design">Module 03: import design&lt;a class="anchor" href="#module-03-import-design">#&lt;/a>&lt;/h2>
&lt;h3 id="input">Input&lt;a class="anchor" href="#input">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>Netlist in Verilog&lt;/li>
&lt;li>Floorplan in DEF&lt;/li>
&lt;li>Clock tree spec auto gen from SDC&lt;/li>
&lt;li>Scan info in Tcl or DEF&lt;/li>
&lt;li>I/O info (pads or pins)&lt;/li>
&lt;li>GDS layer map (if want to dump GDS)&lt;/li>
&lt;li>Timing constraint in SDC&lt;/li>
&lt;li>Timing library in .lib&lt;/li>
&lt;li>LEF library of cells&lt;/li>
&lt;li>Tech file for extraction (cap table or qrc)&lt;/li>
&lt;/ul>
&lt;h3 id="import-design">import design&lt;a class="anchor" href="#import-design">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>Save all the input file paths and parameters in a .globals file, then next time just use &lt;code>source design.globals; init_design&lt;/code>.&lt;/li>
&lt;li>Q: what if there is errors, such as mismatch between netlist and libraries?
&lt;ul>
&lt;li>A: in early stage, mismatch is OK. For example, importing a new Verilog netlist but along with an old DEF containing floorplan. But in late stage, the mismatch is serious.&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Tips: save and load a workspace (window layout) use menu “Windows -&amp;gt; Save Workspace”.&lt;/li>
&lt;li>Tips: menu could be customized using terminal commands “ui*”&lt;/li>
&lt;/ul>
&lt;h3 id="check-design">check design&lt;a class="anchor" href="#check-design">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>&lt;code>checkDesign&lt;/code> to detect missing/inconsistency
&lt;ul>
&lt;li>ex: checkDesign -floorplan -outfile checkDesign.floorplan.rpt&lt;/li>
&lt;li>ex: checkDesign -timingLibrary&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="design-mode">design mode&lt;a class="anchor" href="#design-mode">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>&lt;code>setDesignMode -process 16 -flowEffort {express|standard|extreme}&lt;/code>&lt;/li>
&lt;/ul>
&lt;h2 id="module-04-select-and-highligh-obj">Module 04: select and highligh obj&lt;a class="anchor" href="#module-04-select-and-highligh-obj">#&lt;/a>&lt;/h2>
&lt;h3 id="select">select&lt;a class="anchor" href="#select">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>&lt;code>selectObjByProp &amp;lt;objType&amp;gt; &amp;lt;expression&amp;gt;&lt;/code>&lt;/li>
&lt;li>F12: dim the background&lt;/li>
&lt;li>“instance (right click) -&amp;gt; highlight instance nets”&lt;/li>
&lt;/ul>
&lt;h3 id="design-browser">design browser&lt;a class="anchor" href="#design-browser">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>“tools -&amp;gt; design brower”&lt;/li>
&lt;li>All design stuff in a tree; can use it to select obj and do placement&lt;/li>
&lt;li>Color the modules: right click on “modules”&lt;/li>
&lt;/ul>
&lt;h3 id="schematic-viewer">schematic viewer&lt;a class="anchor" href="#schematic-viewer">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>“tools -&amp;gt; schematic viewer”&lt;/li>
&lt;li>To explore design changes&lt;/li>
&lt;li>Can cross-probe to physical window&lt;/li>
&lt;/ul>
&lt;h3 id="three-views">three views&lt;a class="anchor" href="#three-views">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>floorplan view&lt;/li>
&lt;li>amoeba view : display outline of modules/sub-modules after placement to check locality of the module&lt;/li>
&lt;li>physical view: detailed placments and routing&lt;/li>
&lt;/ul>
&lt;h2 id="module-05-floorplan">Module 05: floorplan&lt;a class="anchor" href="#module-05-floorplan">#&lt;/a>&lt;/h2>
&lt;h3 id="sites-and-rows">sites and rows&lt;a class="anchor" href="#sites-and-rows">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>site: basic horizontal unit&lt;/li>
&lt;li>row: core rows / IO rows&lt;/li>
&lt;/ul>
&lt;h3 id="what-is-floorplanning">what is floorplanning?&lt;a class="anchor" href="#what-is-floorplanning">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>define die size&lt;/li>
&lt;li>place IO&lt;/li>
&lt;li>create soft blocks&lt;/li>
&lt;li>power planning&lt;/li>
&lt;li>macro placement&lt;/li>
&lt;li>early routing congestion/utilization check&lt;/li>
&lt;/ul>
&lt;h3 id="specify-floorplan">specify floorplan&lt;a class="anchor" href="#specify-floorplan">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>“floorplan -&amp;gt; specify floorplan” = &lt;code>floorPlan&lt;/code>&lt;/li>
&lt;li>Leave space between core and IO to place power rings&lt;/li>
&lt;li>Tips: evaluate routing resource (number of layers, routing tracks) to decide if the core should be high &amp;amp; thin or short &amp;amp; wide
&lt;ul>
&lt;li>core utilization = standard cell + macros / area&lt;/li>
&lt;li>cell utilization = standard cell / area&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>orientation
&lt;ul>
&lt;li>R0: no rotation&lt;/li>
&lt;li>R90: counter-clockwise 90 degree rotation&lt;/li>
&lt;li>MX: mirror through X axis&lt;/li>
&lt;li>MY: mirror through Y axis&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="assign-pads-and-pins">assign pads and pins&lt;a class="anchor" href="#assign-pads-and-pins">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>either read in DEF file or Innovus IO file format; or randomly create one then dump out and modify&lt;/li>
&lt;li>“edge 0” is the left-most edge at Y=0 which is the staring point for IO assignment&lt;/li>
&lt;li>“User Guide -&amp;gt; infrastructure -&amp;gt; data preparation -&amp;gt; generating the IO assignment file”&lt;/li>
&lt;/ul>
&lt;h3 id="automatic-floorplan">automatic floorplan&lt;a class="anchor" href="#automatic-floorplan">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>Seed: design blocks definition used to guide the auto floorplan (not a seed for randomization)&lt;/li>
&lt;li>Use &lt;code>setPlanDesignMode&lt;/code> to control some advanced placement options&lt;/li>
&lt;li>use &lt;code>planDesign&lt;/code> create init floorplan&lt;/li>
&lt;li>NOTE: most of the time, not useful at all&lt;/li>
&lt;/ul>
&lt;h3 id="floorplan-toolbox">floorplan toolbox&lt;a class="anchor" href="#floorplan-toolbox">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>placement blockage type
&lt;ul>
&lt;li>hard: restricted no&lt;/li>
&lt;li>soft: can be used during place opt, CTS, ECO, legalization&lt;/li>
&lt;li>partial: percentage of unavailability&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>routing blockage: layer / type specific&lt;/li>
&lt;li>rectilinear object use scissor tool&lt;/li>
&lt;li>rectilinear floorplan
&lt;ul>
&lt;li>view -&amp;gt; preference -&amp;gt; enable rectlinear design&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>“floorplan -&amp;gt; resize” = &lt;code>setResizeFPlanMode&lt;/code> / &lt;code>resizeFloorplan&lt;/code>&lt;/li>
&lt;li>stairway style floorplan edge&lt;/li>
&lt;li>relative floorplan: move in a group
&lt;ul>
&lt;li>&lt;code>create_relative_floorplan -horizontal_edge_separate {target_edge distance ref_edge} -vertical_edge_separate ...&lt;/code>&lt;/li>
&lt;li>&lt;code>delete_relative_floorplan&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>“floorplan -&amp;gt; edit floorplan” has the same tools with the toolbox to do the job&lt;/li>
&lt;/ul>
&lt;h3 id="create-row">create row&lt;a class="anchor" href="#create-row">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>“floorplan -&amp;gt; row -&amp;gt; create core row” = &lt;code>createRow&lt;/code>
&lt;ul>
&lt;li>need to choose the site&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>rows can be stretched using “floorplan -&amp;gt; row -&amp;gt; stretch core row”&lt;/li>
&lt;li>sometimes can create rows outside of core into the pad area&lt;/li>
&lt;/ul>
&lt;h3 id="rectilinear-blockage">rectilinear blockage&lt;a class="anchor" href="#rectilinear-blockage">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>also use the “cut rectilinear” tool&lt;/li>
&lt;/ul>
&lt;h3 id="module-constraint-types">module constraint types&lt;a class="anchor" href="#module-constraint-types">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>5 different levels of module constraint types
&lt;ol>
&lt;li>None&lt;/li>
&lt;li>soft guide (SoftGuide): weakly grouping instances under the same soft guide, but actually they can be placed through out the whole core area&lt;/li>
&lt;li>guide (Guide): preplacement guide for the module in the core design area&lt;/li>
&lt;li>region (Region): force instances in the regsion, but allow other modules in as well&lt;/li>
&lt;li>Fence (Fence): like region but don&amp;rsquo;t allow other modules. A module becomes a fence when the module is specified as partition. Usually used for hierarchical (bottom-up) design.&lt;/li>
&lt;/ol>
&lt;/li>
&lt;li>Tips: when first imported design, all std cell instances will be under the same top module, you have to ungroup it to break it down into pieces.&lt;/li>
&lt;li>Tips: TU = target utilization (std cell + macro); EU = effective utilization (std cell + macro + blockage)&lt;/li>
&lt;/ul>
&lt;h3 id="instance-placement-status">instance placement status&lt;a class="anchor" href="#instance-placement-status">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>change from “placed” to “fixed” after floorplan for macros&lt;/li>
&lt;li>softfixed: cannot be moved by global placement, but can be moved by legalization and upsize by optimization&lt;/li>
&lt;/ul>
&lt;h3 id="placement-halo">placement halo&lt;a class="anchor" href="#placement-halo">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>diff from blockage, halo move along with the target block&lt;/li>
&lt;li>“floorplan -&amp;gt; edit floorplan -&amp;gt; edit halo” = &lt;code>addHaloToBlock&lt;/code> / &lt;code>deleteHaloFromBlock&lt;/code>&lt;/li>
&lt;/ul>
&lt;h3 id="routing-halo">routing halo&lt;a class="anchor" href="#routing-halo">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>prevent signal integrity issues around macros
&lt;ul>
&lt;li>direct connection is ok&lt;/li>
&lt;li>no long wires&lt;/li>
&lt;li>no jogging&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>“addRoutingHalo” / “deleteRoutingHalo”&lt;/li>
&lt;li>Tips: snap objects (region/guide/macro/blockage to instance grid or routing track)&lt;/li>
&lt;li>Tips: “floorplan -&amp;gt; clear floorplan” = “deleteAllFPObjects” / “deleteSelectedFromFPlan”&lt;/li>
&lt;/ul>
&lt;h3 id="instance-group">instance group&lt;a class="anchor" href="#instance-group">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>create instance group to group instances without changing the logic hierarchy for physical implementation&lt;/li>
&lt;li>&lt;code>createInstGroup&lt;/code> / &lt;code>addInstToInstGroup&lt;/code>&lt;/li>
&lt;li>OR &lt;code>createLogicHierarchy&lt;/code> to change the netlist if preferred&lt;/li>
&lt;/ul>
&lt;h3 id="auto-finish-floorplan">auto finish floorplan&lt;a class="anchor" href="#auto-finish-floorplan">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>“floorplan -&amp;gt; automatic floorplan -&amp;gt; finish floorplan”&lt;/li>
&lt;/ul>
&lt;h3 id="save-floorplan">save floorplan&lt;a class="anchor" href="#save-floorplan">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>&lt;code>writeFPlanScript&lt;/code>&lt;/li>
&lt;/ul>
&lt;h3 id="how-to-reduce-die-size">How to reduce die size&lt;a class="anchor" href="#how-to-reduce-die-size">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>shape is important, considering routing layers are diff in horizontal and vertical directions&lt;/li>
&lt;li>start from 70% utilization, then iteratively evaluate the timing/congestion results&lt;/li>
&lt;/ul>
&lt;h2 id="module-06-power-plan">Module 06: power plan&lt;a class="anchor" href="#module-06-power-plan">#&lt;/a>&lt;/h2>
&lt;h3 id="what-is-power-planning">What is power planning?&lt;a class="anchor" href="#what-is-power-planning">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>create rings/strips&lt;/li>
&lt;li>define global power/ground nets, as well as power structures&lt;/li>
&lt;/ul>
&lt;h3 id="commands">commands&lt;a class="anchor" href="#commands">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>“power -&amp;gt; power plannning -&amp;gt; …” or &lt;code>addRing/addStripe/editPowerVia&lt;/code>
&lt;ul>
&lt;li>Full Geometry Checker (FGC) is enabled under 20nm process&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>“power -&amp;gt; connect global nets” or &lt;code>globalNetConnect&lt;/code> to connect PG nets logically
&lt;ul>
&lt;li>&lt;code>globalNetConnect VDD -type pgpin -pin VDD -all&lt;/code>&lt;/li>
&lt;li>&lt;code>globalNetConnect VSS -type tielow&lt;/code>&lt;/li>
&lt;li>?shouldn&amp;rsquo;t this be a part of UPF&amp;rsquo;s function?&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="add-rings">add rings&lt;a class="anchor" href="#add-rings">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>for whole core or blocks or IO&lt;/li>
&lt;li>advanced tab
&lt;ul>
&lt;li>extend to re-use core ring for block rings&lt;/li>
&lt;li>add ring around cluster of selected blocks
&lt;ul>
&lt;li>&lt;code>addRing -type block_rings -around cluster/shared_cluster&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>use wire groups, to avoid max width DRC&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="add-stripes">add stripes&lt;a class="anchor" href="#add-stripes">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>the concept of “set”
&lt;ul>
&lt;li>ex. physically the power/ground wires are (VDD–VSS——VDD–VSS), then the set distance is from first VDD to next VDD&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>also have the option to connect to wire groups of the ring&lt;/li>
&lt;li>advanced tab
&lt;ul>
&lt;li>break strips at block ring: “omit stripes inside block rings”&lt;/li>
&lt;li>“merge with ring” to save resource by defining a threshold&lt;/li>
&lt;li>&lt;code>setAddStipeMode -orthogonal offset&lt;/code> to control strips go beyond or within the edge of the block&lt;/li>
&lt;li>&lt;code>setAddStripMode -break_at blocks_without_same_net&lt;/code> to disjoint stripes in different power domain&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="via-gen">via gen&lt;a class="anchor" href="#via-gen">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>for overlap area of strip/ring&lt;/li>
&lt;li>shrink size of via to allow more routing resources for signal&lt;/li>
&lt;li>“target penetration”: how long the target wire goes into the fat wire&lt;/li>
&lt;/ul>
&lt;h3 id="tips">Tips&lt;a class="anchor" href="#tips">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>
&lt;p>ignore DRC during early stage, to save run-time&lt;/p></description></item><item><title>Register-based SRAM Read Circuit RTL Example using generate</title><link>https://jimwang99.github.io/posts/chip-design/register-based-sram-read-circuit-rtl-example-using-generate/</link><pubDate>Wed, 18 Jan 2017 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/chip-design/register-based-sram-read-circuit-rtl-example-using-generate/</guid><description>&lt;p>Some parameterized example RTL code for register-based SRAM read circuit using “generate” feature&lt;/p>
&lt;pre tabindex="0">&lt;code>parameter d = 32; // FIFO depth
parameter w = 64; // FIFO data bit-width

logic [w-1:0] mem [d-1:0]; // FIFO memory array
logic [d-1:0] rwl; // 1-hot read word line

// read circuit using &amp;#34;generate&amp;#34;
wire [w-1:0] word_or;
genvar width, depth;
generate
 for (width = 0; width &amp;lt; w; width++) begin: rbit
 wire [d-1:0] bit_or;
 for (depth = 0; depth &amp;lt; d; depth++) begin: rmux
 assign bit_or[depth] = mem[depth][width] &amp;amp; rwl[depth];
 end
 assign word_or[width] = |bit_or;
 end
endgenerate

reg [w-1:0] idout;
always @ (negedge CKB) begin
 idout &amp;lt;= word_or;
end&lt;/code>&lt;/pre></description></item><item><title>SystemVerilog for Design Note</title><link>https://jimwang99.github.io/posts/chip-design/systemverilog-for-design-note/</link><pubDate>Tue, 10 Jan 2017 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/chip-design/systemverilog-for-design-note/</guid><description>&lt;blockquote class='book-hint '>
&lt;p>This is my reading note of book “SystemVerilog for Design (2nd edition)&amp;quot;. As a non-full-time RTL designer, it has opened my mind. But still, I&amp;rsquo;m sad about the antient tool that we are using to design hardware.&lt;/p>&lt;/blockquote>&lt;h2 id="chapter-2-systemverilog-declaration-spaces">Chapter 2: SystemVerilog Declaration Spaces&lt;a class="anchor" href="#chapter-2-systemverilog-declaration-spaces">#&lt;/a>&lt;/h2>
&lt;h3 id="package">Package&lt;a class="anchor" href="#package">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>Verilog shortage: no global declaration&lt;/li>
&lt;li>&lt;code>package ... endpackage&lt;/code>
&lt;ul>
&lt;li>share user-defined type definitions across multiple modules&lt;/li>
&lt;li>independent of modules&lt;/li>
&lt;li>parameters cannot be redefined
&lt;ul>
&lt;li>in package, parameter is similar to localparam, cos in module localparam cannot be directly redefined while instantiation&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>referencing
&lt;ul>
&lt;li>&lt;code>::&lt;/code> the scope resolution operator
&lt;ul>
&lt;li>&lt;code>package_name::package_member&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>use &lt;code>import&lt;/code> to import package into current space
&lt;ul>
&lt;li>&lt;code>import package_name::package_member&lt;/code>
&lt;ul>
&lt;li>TIPS: importing an enumerated type definition will not import the labels automatically&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;code>import package_name::*&lt;/code>
&lt;ul>
&lt;li>what is used will be imported&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;code>$unit&lt;/code> declaration space&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>TIPS: synthesis guide
&lt;ul>
&lt;li>tasks and functions must be &lt;code>automatic&lt;/code>
&lt;ul>
&lt;li>storage for automatic task/function is allocated each time it&amp;rsquo;s called&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>cannot use &lt;code>static&lt;/code> variables, which are supposed to be shared by all instances&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="unit-compilation-unit-declarations">$unit: compilation-unit declarations&lt;a class="anchor" href="#unit-compilation-unit-declarations">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>declaration space &lt;strong>outside&lt;/strong> of package/module/interface/program
&lt;ul>
&lt;li>BUT it&amp;rsquo;s &lt;strong>not&lt;/strong> global&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>if put variables and nets in $unit
&lt;ul>
&lt;li>source code order can affect the usage of a declaration external to the module&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>each &lt;strong>compilation&lt;/strong> has one $unit
&lt;ul>
&lt;li>single-file compilation&lt;/li>
&lt;li>multiple-file compilation: source order is tricky&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>TIPS: coding guide
&lt;ul>
&lt;li>DONOT make any declarations in $unit space, only import packages into $unit&lt;/li>
&lt;li>ILLEGAL to import the same package more than once into the same $unit&lt;/li>
&lt;li>NOTE: donot work for global variables, static task/function&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;pre tabindex="0">&lt;code>// filename: def.pkg
`ifdef DEF_PKG
`define DEF_PKG

package def;
// ...
endpackage
`endif&lt;/code>&lt;/pre>&lt;pre tabindex="0">&lt;code>// in every design or testbench file that need package &amp;#34;def&amp;#34;
`include &amp;#34;def.pkg&amp;#34;&lt;/code>&lt;/pre>&lt;ul>
&lt;li>identifier search rules&lt;/li>
&lt;/ul>
&lt;ol>
&lt;li>local&lt;/li>
&lt;li>package
&lt;ul>
&lt;li>named first&lt;/li>
&lt;li>&lt;code>*&lt;/code> wildcard second&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;code>$unit&lt;/code>&lt;/li>
&lt;li>design hierarchy&lt;/li>
&lt;/ol>
&lt;ul>
&lt;li>TIPS: synthesis guide
&lt;ul>
&lt;li>use packages instead of $unit&lt;/li>
&lt;li>external task/function must be automatic&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="namedunnamed-statement-blocks">Named/unnamed statement blocks&lt;a class="anchor" href="#namedunnamed-statement-blocks">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>local variables in named blocks can be accessed hierarchically&lt;/li>
&lt;li>local variables in unnamed blocks (added in SV) has no hierarchical path
&lt;ul>
&lt;li>protecting from external, cross-module referencing&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="timing-units-and-precision">Timing units and precision&lt;a class="anchor" href="#timing-units-and-precision">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>problem with Verilog&amp;rsquo;s timescale directive: file order dependent&lt;/li>
&lt;li>SystemVerilog improvements
&lt;ul>
&lt;li>time value with time units: 5ns, 3.2ps
&lt;ul>
&lt;li>NOTE: there is no space between number and unit&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>scope-level time units and precision: timeunit &amp;amp; timeprecision keywords
&lt;ul>
&lt;li>must be immediately after module/interface/program declaration&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>search order&lt;/li>
&lt;/ul>
&lt;ol>
&lt;li>local&lt;/li>
&lt;li>parent module/interface&lt;/li>
&lt;li>&lt;code>timescale&lt;/code> in effect while compilation&lt;/li>
&lt;li>defined in $unit&lt;/li>
&lt;li>simulator default&lt;/li>
&lt;/ol>
&lt;h2 id="chapter-3-systemverilog-literal-values-and-built-in-data-types">Chapter 3: SystemVerilog Literal Values and Built-in Data Types&lt;a class="anchor" href="#chapter-3-systemverilog-literal-values-and-built-in-data-types">#&lt;/a>&lt;/h2>
&lt;h3 id="literal-value-enhancement">Literal value enhancement&lt;a class="anchor" href="#literal-value-enhancement">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>Verilog tricks to fill vector with all ones
&lt;ul>
&lt;li>&lt;code>data = ~0; // one's complement&lt;/code>&lt;/li>
&lt;li>&lt;code>data = -1; // two's complement&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>SystemVerilog: apostrophe(tick) ( &amp;rsquo; ) (Note: not back-tick ( ` ))
&lt;ul>
&lt;li>&lt;code>data = '1; // all 1's&lt;/code>&lt;/li>
&lt;li>&lt;code>data = 'z; // all z's&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="define-enhancement">DEFINE enhancement&lt;a class="anchor" href="#define-enhancement">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>String&lt;/li>
&lt;/ul>
&lt;pre tabindex="0">&lt;code>// Verilog
`define print(v) $display(&amp;#34;variable v = %h&amp;#34;, v)
`print(data); // = $display(&amp;#34;variable v = %h&amp;#34;, data);

// SystemVerilog
`define print(v) $display(`&amp;#34;varaible v = %h`&amp;#34;, v)
`print(data); // = $display(&amp;#34;variable data = %h&amp;#34;, data);

// SystemVerilog: escape with double `
`define print(v) $display(`&amp;#34;varaible `\`&amp;#34;v`\`&amp;#34; = %h`&amp;#34;, v)
`print(data); // = $display(&amp;#34;varaible \&amp;#34;data\&amp;#34; = %h&amp;#34;, data);&lt;/code>&lt;/pre>&lt;ul>
&lt;li>Construct identifier names: double back-tick w/o space will separate names that will allow 2 or more names to be replaced and form a new name.&lt;/li>
&lt;/ul>
&lt;pre tabindex="0">&lt;code>`define MY_NET(index) bit my_net``index``_bit;
`MY_NET(00) // = bit my_net00_bit;
`MY_NET(15) // = bit my_net15_bit;&lt;/code>&lt;/pre>&lt;h3 id="variables">Variables&lt;a class="anchor" href="#variables">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>Type
&lt;ul>
&lt;li>Net: “wire” keyword, only 4-state&lt;/li>
&lt;li>Variable: “var” keyword, most of the time it can be omitted&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Data type: value system
&lt;ul>
&lt;li>2-state: “bit” keyword&lt;/li>
&lt;li>4-state: “logic” keyword (to replace “reg” keyword)&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Explicit &amp;amp; implicit (shit-hole of SystemVerilog)&lt;/li>
&lt;/ul>
&lt;pre tabindex="0">&lt;code>// 4-state 8-bit varaible
logic [7:0] busA;
// to be explicitly
var logic [7:0] busA;

// 2-state 32-bit variable
bit [31:0] busB;
// to be explicitly
var bit[31:0] busB;

// 4-state 8-bit net
wire [7:0] busC;
// to be explicitly
wire logic [7:0] busC;

wire reg [31:0] busD; // ILLEGAL&lt;/code>&lt;/pre>&lt;ul>
&lt;li>Signed vs. Unsigned
&lt;ul>
&lt;li>Concatenation automatic create &lt;code>unsigned&lt;/code> result&lt;/li>
&lt;li>&lt;code>logic&lt;/code> are unsigned by default&lt;/li>
&lt;li>&lt;code>int&lt;/code> are signed by default&lt;/li>
&lt;li>syntax: &lt;code>&amp;lt;type&amp;gt; &amp;lt;signed/unsigned&amp;gt; &amp;lt;bit width&amp;gt; &amp;lt;name&amp;gt;;&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>TIPS: synthesis guide
&lt;ul>
&lt;li>Because 2-state data types begins simulation with default 0 instead of X, if they are used in RTL may cause RTL behavior mismatch gate-level netlist. So they are mostly used in verification&lt;/li>
&lt;li>Converting from 4-state to 2-state, X and Z are mapped to 0&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>High level data type:
&lt;ul>
&lt;li>2-state data type: used for abstract model or DPI (Direct Programming Interface) to work with C/C++ model
&lt;ul>
&lt;li>&lt;code>byte&lt;/code>: 8-bit&lt;/li>
&lt;li>&lt;code>shortint&lt;/code>: 16-bit&lt;/li>
&lt;li>&lt;code>int&lt;/code>: 32-bit&lt;/li>
&lt;li>&lt;code>longint&lt;/code>: 64-bit&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;code>void&lt;/code>: no storage&lt;/li>
&lt;li>&lt;code>shortreal&lt;/code>: 32-bit single-precision = float in C, while &lt;code>real&lt;/code> = double in C&lt;/li>
&lt;li>&lt;code>classes&lt;/code>: not covered in this book&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>NOTE: Most signals can be declared as &lt;code>logic&lt;/code> in RTL
&lt;ul>
&lt;li>&lt;code>logic&lt;/code> for single-driver&lt;/li>
&lt;li>&lt;code>wire&lt;/code> for multi-driver logic (wand/wor)&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Value drivers
&lt;ul>
&lt;li>Any number of &lt;code>initial&lt;/code> or &lt;code>always&lt;/code> blocks
&lt;ul>
&lt;li>NOTE: it&amp;rsquo;s only for back-compatable with Verilog which is not really circuit behavior&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Single &lt;code>always_comb/always_ff/always_latch&lt;/code> block&lt;/li>
&lt;li>Single &lt;code>assign&lt;/code> statement&lt;/li>
&lt;li>Single &lt;code>module/primitive output/inout&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Type casting
&lt;ul>
&lt;li>Static casting (synthesizable)
&lt;ul>
&lt;li>Size casting: &lt;code>&amp;lt;size&amp;gt;'(&amp;lt;expression&amp;gt;)&lt;/code>
&lt;ul>
&lt;li>ex. &lt;code>16'(2) // 16-bit wide&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Sign casting: &lt;code>&amp;lt;sign&amp;gt;'(&amp;lt;expression&amp;gt;)&lt;/code>
&lt;ul>
&lt;li>ex. &lt;code>signed'({a, b}) // unsigned concatenation result to signed&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Dynamic casting (has error check)
&lt;ul>
&lt;li>ex. &lt;code>cast(dest_var, source_exp);&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Varaible initialization
&lt;ul>
&lt;li>SystemVerilog in-line initialization is before time zero and does not cause a simulation event&lt;/li>
&lt;li>Testbench should initialize varaibles to their inactive state&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Static and automatic variables
&lt;ul>
&lt;li>&lt;code>static&lt;/code> vs &lt;code>automatic&lt;/code>
&lt;ul>
&lt;li>Storage&lt;/li>
&lt;li>Automatic variables can be used for re-entrant tasks and recursive functions.&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Module level, all varaibles are static&lt;/li>
&lt;li>&lt;code>begin … end&lt;/code> and &lt;code>fork … join&lt;/code> blocks, tasks and functions, all storage defaults to static&lt;/li>
&lt;li>Automatic tasks and functions have all automatic storages&lt;/li>
&lt;li>Initialization
&lt;ul>
&lt;li>Static variables are only initialized once&lt;/li>
&lt;li>Automatic variables are initialized each call&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h3 id="constants">Constants&lt;a class="anchor" href="#constants">#&lt;/a>&lt;/h3>
&lt;ul>
&lt;li>Verilog
&lt;ul>
&lt;li>&lt;code>parameter&lt;/code>: can be redefined when instantiation&lt;/li>
&lt;li>&lt;code>specparam&lt;/code>: can be redefined from SDF&lt;/li>
&lt;li>&lt;code>localparam&lt;/code>: elaboration-time constant, cannot be directly redefined&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>SystemVerilog: C-like const keyword
&lt;ul>
&lt;li>&lt;code>const int N = 5;&lt;/code>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h2 id="chapter-4-systemverilog-user-defined-and-enumerated">Chapter 4: SystemVerilog User-Defined and Enumerated&lt;a class="anchor" href="#chapter-4-systemverilog-user-defined-and-enumerated">#&lt;/a>&lt;/h2>
&lt;h3 id="typedef-keyword">&lt;code>typedef&lt;/code> keyword&lt;a class="anchor" href="#typedef-keyword">#&lt;/a>&lt;/h3>
&lt;p>Ex. &lt;code>typedef int unsigned uint;&lt;/code>&lt;/p></description></item><item><title>Scan Chain Problem of Clock Generator Flip-Flops</title><link>https://jimwang99.github.io/posts/chip-design/scan-chain-problem-of-clock-generator-flip-flops/</link><pubDate>Mon, 09 Jan 2017 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/chip-design/scan-chain-problem-of-clock-generator-flip-flops/</guid><description>&lt;p>&lt;img src="https://jimwang99.github.io/legacy-media/scan-chain-of-clock-gen.jpg" alt="img" />&lt;/p>
&lt;p>As shown in the schematic, we have some clock divider that divide root clock by half. While in scan mode, these flip-flops will be bypassed and treated as normal flip-flop that need to be inserted into the scan chain along with leaf flip-flops. But due to the nature of clock tree, clock divider will be in the upper stream and will have a much smaller clock insertion delay. Then it will cause large hold time violation from clock generator flip-flops to normal leaf flip-flops, and these violations cannot be fixed easily.&lt;/p></description></item><item><title>My experience with custom digital design</title><link>https://jimwang99.github.io/posts/chip-design/my-experience-with-custom-digital-design/</link><pubDate>Fri, 06 Jun 2014 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/chip-design/my-experience-with-custom-digital-design/</guid><description>&lt;h2 id="background">Background&lt;a class="anchor" href="#background">#&lt;/a>&lt;/h2>
&lt;p>This is the summary of my experience from project LBRAM in Marvell in the year of 2014.&lt;/p>
&lt;h3 id="the-first-thing-discuss-timingareapower-specs-in-details">The first thing: discuss timing/area/power SPEC’s in details&lt;a class="anchor" href="#the-first-thing-discuss-timingareapower-specs-in-details">#&lt;/a>&lt;/h3>
&lt;p>Most of the time, because custom design takes lots of time, it often starts ahead of chips. At that time, the design SPEC’s, such as timing/area/power, are not clear. So try to discuss it with your supervisor or the project leader or your customer to define these SPEC’s even if they are not accurate. And do remember to write them down in some documents, so if they want to change the SPEC’s later, you can show them how absurd the idea is. ;-)&lt;/p></description></item><item><title>Survey of Low Power Design</title><link>https://jimwang99.github.io/posts/chip-design/survey-of-low-power-design/</link><pubDate>Mon, 07 Sep 2009 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/chip-design/survey-of-low-power-design/</guid><description>&lt;blockquote class='book-hint '>
&lt;p>从2017年初的观点来看，这篇报告的部分内容过时了，但是整体结构还是比较适合的。希望今年有时间能够出一版更新的版本。&lt;/p>&lt;/blockquote>&lt;p>低功耗设计的最根本驱动力是集成电路芯片的功耗随着工艺的进步不仅没有下降反而不断上涨。因为晶体管速度和集成度的上升速度超过了电路单次翻转所消耗能量的下降速度，所以单位面积芯片的功耗在迅速上升。而根据ITRS的预测，固定电源供电设备和移动设备中芯片的功耗发展趋势如图表 1所示。从中我们不难看出，各类芯片的各种功耗都在不断飞速上升，已经成为芯片设计者不容小觑的问题。&lt;/p>
&lt;h2 id="简介">简介&lt;a class="anchor" href="#%e7%ae%80%e4%bb%8b">#&lt;/a>&lt;/h2>
&lt;p>低功耗设计的最根本驱动力是集成电路芯片的功耗随着工艺的进步不仅没有下降反而不断上涨。因为晶体管速度和集成度的上升速度超过了电路单次翻转所消耗能量的下降速度，所以单位面积芯片的功耗在迅速上升。而根据ITRS的预测，固定电源供电设备和移动设备中芯片的功耗发展趋势如图表 1所示。从中我们不难看出，各类芯片的各种功耗都在不断飞速上升，已经成为芯片设计者不容小觑的问题。&lt;/p>
&lt;p>图表 1：芯片功耗的发展趋势：固定电源供电设备（左）和移动设备（右）&lt;/p>
&lt;p>&lt;img src="https://jimwang99.github.io/legacy-media/survey-of-low-power-design-image001.png" alt="image001.png" />&lt;/p>
&lt;p>图表 2：不同工艺下芯片功耗发展趋势&lt;/p>
&lt;p>&lt;img src="https://jimwang99.github.io/legacy-media/survey-of-low-power-design-image003.png" alt="image003.png" />&lt;/p>
&lt;p>同时，随着工艺的进步，提升晶体管速度的难度在不断增加，导致晶体管延时的下降幅度不断减小，如图表 2所示。因此为了继续提升电路的整体性能，芯片设计者不断引入新技术来弥补晶体管速度的不足。例如，使用低介电常数（Low-K）的电介质和低电阻率的金属线（铜金属线）。除此之外，许多其他技术的引入会进一步增加功耗，例如使用SOI（Silicon-on-Insulator）衬底材料、增加载流子迁移率（Strained Silicon）和提高电磁场强度（Overdrive技术）。这些新技术的引入不仅增加了单位面积内功耗的总量，还增加了漏电功耗在整体功耗中的比重，从而使得一些移动应用迫切需要进行低功耗设计。&lt;/p>
&lt;p>时至今日，如何降低动态功耗是现在几乎所有IC设计者关注的焦点之一。对于使用电池供电的移动应用而言，降低芯片功耗能够延长产品的续航时间。这是一个非常具有诱惑力的特性。对于使用固定电源供电的应用而言，降低芯片功耗也能带来许多好处。例如，可以降低设备成本，因为能够使用更便宜的封装；能够达到更高的性能，因为芯片温度下降了。对于企业级数据存储和通信基站这样的系统而言，降低功耗更能够节约巨大的成本，因为可以使用更便宜的制冷系统。&lt;/p>
&lt;h3 id="功耗的基础概念">功耗的基础概念&lt;a class="anchor" href="#%e5%8a%9f%e8%80%97%e7%9a%84%e5%9f%ba%e7%a1%80%e6%a6%82%e5%bf%b5">#&lt;/a>&lt;/h3>
&lt;h4 id="功耗的分类">功耗的分类&lt;a class="anchor" href="#%e5%8a%9f%e8%80%97%e7%9a%84%e5%88%86%e7%b1%bb">#&lt;/a>&lt;/h4>
&lt;p>图表 3：功耗的分类&lt;/p>
&lt;p>&lt;img src="https://jimwang99.github.io/legacy-media/survey-of-low-power-design-image005.png" alt="image005.png" />&lt;/p>
&lt;p>如图表 3所示，我们将功耗分为动态功耗、短路功耗和漏电功耗。动态功耗为图中红色线条，当PMOS管开启时，为负载电容CL充电的电流引起的功耗。短路功耗为图中绿色线条，当输入信号发生翻转时，PMOS和NMOS会同时处于半开启状态，此时流经PMOS和NMOS的电流引起的功耗。漏电功耗为稳定状态下，MOS管的漏电涓电流（图中黄色线条）引起的功耗。&lt;/p>
&lt;p>&lt;img src="https://jimwang99.github.io/legacy-media/survey-of-low-power-design-image008.gif" alt="image008.gif" />&lt;/p>
&lt;p>从公式中我们可以看出，动态功耗与短路功耗均与翻转概率、频率和电压成正比。&lt;/p>
&lt;h4 id="eda工具功耗报告的分类">EDA工具功耗报告的分类&lt;a class="anchor" href="#eda%e5%b7%a5%e5%85%b7%e5%8a%9f%e8%80%97%e6%8a%a5%e5%91%8a%e7%9a%84%e5%88%86%e7%b1%bb">#&lt;/a>&lt;/h4>
&lt;p>在EDA工具的功耗报告中，功耗被分成三类：翻转功耗Switching，内部功耗Internal和漏电功耗Leakage。其中的漏电功耗很好理解，而翻转功耗与内部功耗与我们前面所提到的动态功耗和短路功耗稍有不同。&lt;/p>
&lt;p>由于EDA流程是基于标准单元的设计方法，所以翻转功耗与内部功耗是针对标准单元来说的。我们知道每个电路节点都有寄生电容，因此也会有动态功耗。EDA工具的功耗报告中所提及的内部功耗就是当输入发生变化时，标准单元内部逻辑门的功耗损失，包括内部节点的动态功耗和内部MOS管的短路功耗。而翻转功耗指的是当输出发生变化时，标准单元对连线负载节点充电所引起的动态功耗。需要注意的是，单元A对单元B的输入端口PORT_IN进行充电所引起的动态功耗计算在内部功耗之列。&lt;/p>
&lt;h3 id="能量限制-vs-功耗限制">能量限制 vs. 功耗限制&lt;a class="anchor" href="#%e8%83%bd%e9%87%8f%e9%99%90%e5%88%b6-vs-%e5%8a%9f%e8%80%97%e9%99%90%e5%88%b6">#&lt;/a>&lt;/h3>
&lt;p>在低功耗设计领域有两种不同的应用需要区分清楚：能量限制的应用和功耗限制的应用。能量限制的应用有手机、笔记本、MP3等由电池供电的设备；功耗限制的应用有RFID等由电磁场供电的设备。&lt;/p>
&lt;p>如图表 4所示，能量是由电流的面积积分决定的，所以能量限制的应用需要考虑的是一定时间范围内（通常是指工作和待机时间），其能量消耗（电流随时间的积分乘以电压）尽可能小。而峰值功耗则仅仅由电流的最大值决定，所以功耗限制的应用需要考虑的是瞬态电流（由于片上电容的存在，考虑一个微小时间片内的电流积分）小于电源能提供的最大电流。&lt;/p>
&lt;p>图表 4：能量 vs. 功耗&lt;/p>
&lt;p>&lt;img src="https://jimwang99.github.io/legacy-media/survey-of-low-power-design-image009.png" alt="image009.png" />&lt;/p>
&lt;p>这两种应用的功耗优化策略稍有不同，但是这样的不同之处是非常关键的。对于能量限制的应用而言，应该首先保证每次操作都是有效操作，并尽可能降低每次操作时消耗的能量。至于这些有效操作何时发生并不是非常重要。而对于功耗限制的应用而言，应该在时间上尽可能分散操作，不让芯片的峰值功耗超过供电源能提供的最大功耗值，但是刻意使得某一时刻功耗非常低也是没有意义的。&lt;/p>
&lt;h3 id="本文的结构">本文的结构&lt;a class="anchor" href="#%e6%9c%ac%e6%96%87%e7%9a%84%e7%bb%93%e6%9e%84">#&lt;/a>&lt;/h3>
&lt;p>在本文中，我们将首先讨论RTL级和门级（Gate-Level）下的一些优化策略。&lt;/p>
&lt;p>能够归类于RTL级的低功耗技术比较少，因为在这一级大部分的优化都与具体的设计有直接的联系，所以能够提取出通用的方法并不多，主要是一些通用的在RTL编码中应该注意的问题。&lt;/p>
&lt;p>现今的数字电路设计通常都是基于标准单元的设计方法，所以在门级我们主要关注一些可以借助EDA工具的优化策略和门单元的低功耗设计。&lt;/p>
&lt;p>从前面的介绍中我们可以看到，动态功耗和短路功耗都与电压有直接的联系。所以降低电源电压是最直接也是最有效的低功耗设计手段。所以在接下来的一节我们将对这一类的设计方法进行讨论。有一些技术虽然还处于学术界的研究范畴，还没有在工业界大展拳脚，我们仍然需要对它们进行适当关注。&lt;/p>
&lt;p>在现今的芯片设计中，时钟网络消耗了大量的功耗，必要对这一部分进行独立思考和优化。有不少低功耗技术的优化目标就是时钟网络。这部分内容我们将在之后进行介绍。&lt;/p>
&lt;p>最后，我们将关心漏电功耗的优化策略，了解学术界和工业界对这部分功耗的优化成果。虽然在我们的应用中并不需要特别关注漏电功耗，但是这也是芯片功耗的一个重要组成部分，特别是对于手机这一类的手持移动设备而言更是如此。&lt;/p>
&lt;h2 id="rtl级优化策略">RTL级优化策略&lt;a class="anchor" href="#rtl%e7%ba%a7%e4%bc%98%e5%8c%96%e7%ad%96%e7%95%a5">#&lt;/a>&lt;/h2>
&lt;h3 id="rtl代码优化">RTL代码优化&lt;a class="anchor" href="#rtl%e4%bb%a3%e7%a0%81%e4%bc%98%e5%8c%96">#&lt;/a>&lt;/h3>
&lt;p>在RTL级，大部分对于低功耗优化策略都与实际的设计有密切的关系。需要RTL工程师对于代码综合以后的结果有清晰的认识。在这里我们将首先提出一些针对RTL代码的优化策略。虽然说某些设计精良的综合工具能够帮助我们实施某些优化，但是为了以防万一，也为了缩短综合工具的优化时间，我们应该在设计RTL代码时就将这些优化策略考虑进去。&lt;/p>
&lt;h4 id="提取公因子">提取公因子&lt;a class="anchor" href="#%e6%8f%90%e5%8f%96%e5%85%ac%e5%9b%a0%e5%ad%90">#&lt;/a>&lt;/h4>
&lt;p>尽可能找出计算中重复使用的公因子进行重用，有的时候我们需要根据实际情况进行一些功能相等的转换。需要注意的是，综合工具对某些类型的公因子不敏感，所以我们应该在RTL设计中进行预处理。&lt;/p>
&lt;p>提取公因子举例：能减少一个加法器&lt;/p>
&lt;ul>
&lt;li>优化前&lt;/li>
&lt;/ul>
&lt;pre tabindex="0">&lt;code>if (test)
 y0 = a + b;
else
 y1 = c - a - b;&lt;/code>&lt;/pre>&lt;ul>
&lt;li>优化后&lt;/li>
&lt;/ul>
&lt;pre tabindex="0">&lt;code>assign t = a + b;
if (test)
 y0 = t;
else
 y1 = c - t;&lt;/code>&lt;/pre>&lt;h4 id="资源重用">资源重用&lt;a class="anchor" href="#%e8%b5%84%e6%ba%90%e9%87%8d%e7%94%a8">#&lt;/a>&lt;/h4>
&lt;p>通过代码优化将运算部件（例如关系运算、加减乘除等）进行重用。有的时候综合工具并不能很好的识别出需要做改进的地方。&lt;/p></description></item><item><title>Chisel3 Systolic Array Generator</title><link>https://jimwang99.github.io/posts/chip-design/chisel3-systolic-array-generator/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://jimwang99.github.io/posts/chip-design/chisel3-systolic-array-generator/</guid><description>&lt;p>#hardware #accelerator #chisel #open-source&lt;/p>
&lt;p>I&amp;rsquo;ve tried to learn Chisel 5 years ago, but gave up and went back to SystemVerilog to design our RISC-V CPU + AI custom instructions from scratch. After 5 years, both Chisel and myself grow a lot, so I&amp;rsquo;m giving it another shot. This time, I found Chisel3&amp;rsquo;s documentation and tutorial is way better than 5 years ago. With numerous answers on Stack Overflow, LLVM&amp;rsquo;s CIRCT, Chisel simulator Treadle, and unit-test framework ChiselTest, it&amp;rsquo;s much easier to write and test Chisel/Scala.&lt;/p></description></item></channel></rss>