205 D-FIR

205 : D-FIR

Design render

General Operation

  • Load 12-bit samples or coefficients over SPI interface
    • Use FIR_MODE ui_in[0] pin to switch between coefficient/sample loading
      • FIR_MODE = 1 (SAMPLE)
      • FIR_MODE = 0 (COEFFICIENT)
    • Coefficients are in format Q1.11, 11 decimal bits
      • When generating coefficient vector, ensure you multiply each floating point coefficient by $2^{11}$ and round to the nearest whole number to match Q1.11 input range of [-2048, 2047] for 12-bit signed integers
      • Coefficients are not required to be symmetric, allowing both linear-phase and non-linear-phase FIR filters to be implemented
    • Coefficients must be loaded in reverse order (last tap first) due to the coefficient shift register architecture
      • Your coefficient vector [c35, c34, ..., c1, c0] should be sent over SPI as c0, c1, ..., c34, c35
    • There is an expected amount of error due to the fixed point representation and truncated multiplier precision
      • Truncated multiplier accumulates error across taps due to serial MAC architecture (one multiplier is used across all taps)
      • See the Area–Error Tradeoff section below for a quantified analysis with Sky130A synthesis data
    • Recommended first test:
      • Generate own array of filter coefficients using Python or other online tools (ex. Python SciPy firwin function)
      • Load coefficients over SPI
      • Send impulse response (impulse is equal to 2047 which is max positive value at 12 bit signed)
      • Verify outputs match within an acceptable range to loaded filter coefficients (acceptable range is around 1-2 integer steps, you may load a coefficient of -5 but receive back an output of -6)

Coefficient Reloading

  • Coefficients are runtime programmable through the use of the FIR_MODE pin with some additional notes
    • On fresh startup of filter, best practice is to load your coefficients for all taps and then continuously stream samples
    • If a coefficient reload is desired to change filter behavior during runtime, the recommended setup is to flush the entire sample line with 0x0000/NOPs equal to tap count (36 taps) to not corrupt the filter math with stale samples and then load new coefficients
    • Outputs during this NOP/reloading phase should be ignored; they are stale calculations from the previous filter settings

Validation

  • Design was functionally validated on a Gowin GW5A FPGA using an STM32 Nucleo-F446RE SPI master
  • Validation methodology and measured results are documented in bringup/README.md

SPI Overview

  • System clock is 40 MHz, recommended SCLK $\leq$ ~4-5 MHz
  • 16 bit frames; MSB leading, remaining lower bits padded with 0's
    • With 16 bit frames @ SCLK = 5 MHz, the sampling rate is about:
      • Total SPI frame time = 200ns per bit * 16 bits + cs_n high between frames = 3400ns -> 294 kSps / 2 (Nyquist) = 147 kHz maximum theoretical recoverable signal frequency; real-world maximum will land slightly below
  • SPI Mode 0 only
  • MISO is driven LOW during idle, not tri-stated; do not share MISO line with other SPI slaves unless externally isolated or muxed
  • CS_N high time between frames must be one cycle of SCLK, (For 5 MHz, CS_N high between frames should be 200ns)
  • MISO is valid 2 core clock cycles after CS_N falls (3-stage synchronizer + edge detection). At 40 MHz this is 50 ns. The SPI master must not start SCLK sooner than this after asserting CS_N low, or the first MISO bit may be stale.
  • Set ui_in[0] pin (FIR_MODE) at the beginning of each SPI transaction to your desired transmission type
    • FIR_MODE = 1 (SAMPLE)
    • FIR_MODE = 0 (COEFFICIENT)
    • The mode pin is sampled in the middle of the SPI frame; recommend setting pin at the beginning of the frame
  • Safe operation of filter assumes SPI clock within documented ranges, there is no backpressure/overrun control if you operate above recommended ranges

SPI Output Timing

  • Full duplex SPI interface
    • MISO returns previous samples' results while you send new samples on MOSI
    • This creates a output pipeline with latency dependent on the chosen SCLK frequency
SCLK speed Pipeline lag (k) Leading garbage frames Trailing NOPs needed
2–5 MHz 2 First 2 MISO outputs Send 2x "0x0000" at end
≤1 MHz 1 First 1 MISO output Send 1x "0x0000" at end

Example (k=2): sending [A, B, C, D, E] over MOSI:

Frame MOSI in MISO out
1 A garbage
2 B garbage
3 C result of A
4 D result of B
5 E result of C
6 0x0000 result of D
7 0x0000 result of E
  • The pipeline drain only is only needed with finite sample batches, continous sample streaming does not require trailing NOPs; only required in the case you stop continous streaming

SPI Timing

  • CS must go high between each 16-bit frame as noted in prior sections

Loading Samples:

MODE _/¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯\_
CS_N  \__________________________________________________________________________________/
SCLK  _/¯\_/¯\_/¯\_/¯\_/¯\_/¯\_/¯\_/¯\_/¯\_/¯\_/¯\_/¯\_/¯\_/¯\_/¯\_/¯\_/¯\_/¯\_/¯\_/¯\_/¯\_
MOSI  X  b15  b14  b13  b12  b11  b10  b9   b8   b7   b6   b5   b4   b3   b2   b1   b0  X
MISO  X  d15  d14  d13  d12  d11  d10  d9   d8   d7   d6   d5   d4   d3   d2   d1   d0  X

Loading Coefficients:

MODE \___________________________________________________________________________________/
CS_N  \__________________________________________________________________________________/
SCLK  _/¯\_/¯\_/¯\_/¯\_/¯\_/¯\_/¯\_/¯\_/¯\_/¯\_/¯\_/¯\_/¯\_/¯\_/¯\_/¯\/¯\_/¯\_/¯\_/¯\_/¯\_
MOSI  X  b15  b14  b13  b12  b11  b10  b9   b8   b7   b6   b5   b4   b3   b2   b1   b0  X
MISO  X  d15  d14  d13  d12  d11  d10  d9   d8   d7   d6   d5   d4   d3   d2   d1   d0  X

Truncated Baugh-Wooley Multiplier Design

  • 12x12 bit multiplier, truncating (dropping) the bottom 8 LSP (Least Significant Product) bits [7:0]
  • Uses the Baugh-Wooley algorithm to handle signed multiplication
    • "A Two's Complement Parallel Array Multiplication Algorithm" Baugh and Wooley
  • Uses error correction scheme with IC terms to handle error stemming from truncation
    • "Low Error Truncated Multipliers for DSP Applications" Garofalo et al.
    • Optimized for lowest mean square error along with area

Overview

  • Truncated multiplier to minimize area
    • We use truncation specifically due to the fixed point of the coefficient terms in the FIR filter math
    • In standard binary multiplication, we would compute all 24 product columns and their partial products stemming from the 12x12 multiplication
    • The idea behind the truncated multiplier is that because our coefficients are represented in Q1.11 (11 fractional bits), we can drop a number of the fractional bits and maintain a good level of precision while saving hardware area
    • In our case, we drop 8 of the 11 fractional bits. This means that in our multiplication of 12x12 bit numbers, we never compute the lower 8 partial product columns, saving area that would be spent on adders, AND gates, etc.
  • Partial Product Calculation
    • This section will give an overview to the general steps taken to perform binary multiplication in our case
    • We first create a bit vector of all partial products per column, largest possible width is 12 partial products in a column in our case (width of multiplicand). Each bit position in this vector will store one partial product which is the result of an AND operation between two of the multiplier operand bits ($a_i$ & $b_j$ = $pp_{ij}$)
    • With bit vectors for each partial product column computed, we go down the column and sum them how you would in normal multiplication; ex. partial product column 2 may have a result that is [1,0,1] = sum across that vector and the final result is binary 010 (2 in decimal).
    • Finally, we accumulate across each summed product columns into a final result accumulator vector, making sure to account for the weight of each column (ex. we shift left 12 bits for partial product column 12 to align with the correct bit position and preserve weighting)
  • Baugh-Wooley Algorithm ("A Two's Complement Parallel Array Multiplication Algorithm" Baugh and Wooley)
    • An issue with the truncation method is that it does not natively handle signed (2's complement) multiplication
    • This is fixed by using the Baugh-Wooley algorithm which inverts certain terms in partial products along with the addition of new terms in partial product columns
      • Every y and x term less than ($y_i$ < $y_{m-1}$) and ($x_i$ < $x_{n-1}$) is inverted in their respective partial product pairs
      • Special case for $y_i$ = $y_{m-1}$ and $x_i$ = $x_{n-1}$ where this partial product pair is left unchanged (no inversion)
      • Extra terms added
        • Inverted single terms of $y_{m-1}$ and $x_{n-1}$ are added to the OutputWidth - 2 (12x12 = 24 bit -> Column/bit 22) partial product column
        • Single bit "1" added to the final partial product column, OutputWidth - 1 (Column/bit 23)
        • Two single $y_{m-1}$ and $x_{n-1}$ terms are added at $P_{m-1}$ and $P_{n-1}$ respectively, special case when m = n (both operands are equal bit width) where terms are added to the same partial product column
  • IC Error Correction ("Low Error Truncated Multipliers for DSP Applications" Garofalo et al.)
    • Truncation involves some intrinsic error due to lost partial product terms and carry out from dropped partial product columns
    • We can recover some accuracy by implementing an error correction scheme
    • Since we drop the bottom 8 bits/columns, and our bit width for our inputs equals to 12, we compute our "h" term to be input bit width - dropped bits = 4.
      • "h" represents the number of extra bits kept beyond the minimum necessary of n = 12 bits (matching our input bit width)
    • In order to implement the error compensation, we must calculate the partial products in the one column below our first non-dropped column
      • In our case, drop bits = 8, in our truncated multiplier we compute partial products for columns [23:8]. We must now compute one column lower at column 7 (8 - 1 = column 7)
      • Once this extra lower column has their partial products computed, we must apply the correct f(IC) function on specific bits (Eq. 20, pg. 31)
      • For partial products i = 1, 2, n-h-1 (12 - 4 - 1 = 7), n-h (12 - 4 = 8) we will sum them together and shift them all by $2^{-n-h-1}$ where -n-h-1 = -12-4-1 = -17
        • These are our "edge IC terms"
      • For partial products 2 < i < n-h-1 (12 - 4 - 1 = 7), we apply the shift of $2^{-n-h}$ or $2*2^{-n-h-1}$ (-n-h = -12-4 = -16 -> $2^{-16}$) to them, with our K term coming to be 0
        • These are our "middle IC terms"
      • The term shift can be confusing in this context especially with using negative power of 2 terms
        • The paper uses reverse style for MSB and LSB where the MSB (or most significant product) is at $2^{-1}$ place and the LSB (least significant product) is at the $2^{-2n}$ place
        • Since our output with 12x12 multiplication is 24 bits and we truncate the bottom 8 bits for our output, we must align these at 17 and 16 bits from the left respectively
        • This comes out to 24 - 17 = 7 and 24 - 16 = 8. These fit into partial product column 7 and 8 respectively
      • We must align the edge IC sum partial products and the middle IC sum partial products to their respective bit positions and add them to the main accumulation term
      • One thing to notice is that we do not have a bit 7 in our output of [24:8]. We must expand our accumulator term by one bit width on the LSB side, align the edge IC term with that lower bit and then add normally
      • The middle IC terms align to bit 8, so since we add that lower bit to account for the edge IC addition we shift left by 1 to align the middle IC terms with the correct bit and add normally
      • Note: this extra bit alters the original truncated multiplier shifting in the accumulate stage; we must shift one extra bit each time to account for the extra guard bit added for the IC term alignment to preserve prior weighting for each column

Area–Error Tradeoff

Truncated Multiplier Tradeoff

  • The truncated multiplier was synthesized across all DropBits values (0–11) against the Sky130A HD standard cell library. The plot shows max and mean absolute error vs cell count and chip area.
  • Absolute error in full precision LSBs grows at the rate of ~$2^{DropBits}$ because each further dropped bit carries increasing weight which can contribute to large error at high DropBits
DropBits Max |error| (full LSB) Cells Area (µm²)
0 ≤2 657 5241
4 ≤32 634 4937
8 ≤512 522 4150
11 ≤4096 398 3133

Design point: DropBits = 8 (vertical dashed line). At this level:

  • Error bound: ≤±2 truncated-output LSBs (≤512 full-precision LSBs)
  • ~27% cell savings vs a raw * operator synthesized through the same flow (717 cells baseline)
  • FIR output plot exactly overlaid on from the ideal floating-point reference (see Noisy Sinusoid Filtering below)

Noisy Sinusoid Filtering

Noisy Sine Filtering

A 2kHz sinusoid with added Gaussian noise filtered by the DUT with low-pass coefficients (10 kHz) vs the Python lfilter model.

Impulse Response

Impulse Response

The DUT output (red) overlaid on the ideal fixed-point coefficients (blue). Generated with 50–60 kHz band-pass coefficients.

Step Response

Step Response

The DUT step response (red) vs Python lfilter (blue). The FIR fills in over 36 taps and then converges to the expected steady state step response.

Frequency Response

Frequency Response

50–60 kHz band-pass frequency response, reconstructed from measured impulse response test

IO

#InputOutputBidirectional
0FIR_MODECS_N
1MOSI
2MISO
3SCLK
4
5
6
7

Chip location

Controller Mux Mux Mux Mux Mux Mux Mux Mux Mux Mux Analog Mux Mux Mux Mux Mux Mux Mux Mux Mux Mux Analog Mux Mux Mux Mux Mux Mux Mux Mux Mux Mux tt_um_chip_rom (Chip ROM) tt_um_factory_test (Tiny Tapeout Factory Test) tt_um_teuscher_eml_fabric (EML Fabric — analog exp/ln compute cells) tt_um_wokwi_465656663515438081 (Convert binary to hex on 7 segments) tt_um_nikita_face_detect (FPGA Face Detection) tt_um_obstacle_avoider (Obstacle Avoider State Machine) tt_um_poket_animal (Poket Animal) tt_um_drewbabel_uart (Configurable FIFO-buffered UART with APB CSR) tt_um_wokwi_469163916296039425 (TT Workshop Test) tt_um_jonahsaunders_slsvga (tt_um_jonahsaunders_slsvga) tt_um_fatigue_monitor (Fatigue Monitor (PPG Pulse-Interval Variability)) tt_um_vedam_dual_port_ram (Dual Port RAM) tt_um_wokwi_469739097665887233 (Tiny Tapeout Template Copy) tt_um_spdif_to_i2s_kilpelaj (S/PDIF to I2S receiver) tt_um_morse_converter (ASCII to Morse Code Converter) tt_um_wokwi_469806914724000769 (Spin, Text and VGA) tt_um_wokwi_469701770572338177 (TinyTapeout) tt_um_garnetkoebel_communotron (Communotron) tt_um_wokwi_469449970323169281 (full adder) tt_um_duzabf_2026_ow (A WIP Online Workshop 2026 project) tt_um_wokwi_469807513638180865 (Tiny Tapeout NAK) tt_um_ttsky26c_oguz (ttsky26c-202607-mehmetoguzderin by Oguz) tt_um_kashif_fp4_sparse_tpu (FP4 Sparse Mini-TPU) tt_um_moein_maleki_arm16 (arm16) tt_um_wokwi_469453454643027969 (ON Check System) tt_um_felixcheng_neural_core (Neural Compute Core (V0.15)) tt_um_wokwi_469788774011248641 (Spin Display - select-reset-reverse) tt_um_wokwi_469449443070765057 (Samuel's first chip) tt_um_wokwi_469449007236383745 (testinttrsv01) tt_um_vga_ca (VGA cellular Automaton) tt_um_dosci_500hz (Digital Oscillator 500 Hz) tt_um_wokwi_469747443569078273 (XOR test project - Tiny Tapeout workshop) tt_um_wokwi_469585758593419265 (spinner) tt_um_fp16_mac (FP32 Math Unit) tt_um_1DC_vga_dyoa (VGA Design Your Own ASIC) tt_um_haydenevans_top (Systolic Processing Element) tt_um_wokwi_469806252715961345 (TT_Proj_SA) tt_um_wokwi_469448996577604609 (Tiny Tapeout - Reto) tt_um_ehofmannbr_pmodvga_06 (VGA Color Tiles) tt_um_lfglabs_lsc1u (leanSilicon LSC-1 Micro arithmetic kernel) tt_um_wokwi_469804280240495617 (Zetterling SRAM) tt_um_wokwi_469806066852696065 (TileTestchase) tt_um_wokwi_469449118072978433 (binary_add_v1) tt_um_voltage_amplifier_neuron (Voltage Amplfier Neuron) tt_um_wokwi_469449686545956865 (Tiny Tapeout Template Copy_JinoShiono) tt_um_wokwi_469448887171240961 (Tiny Tapeout - Mini CORDIC) tt_um_wokwi_469809033878555649 (Tiny Tapeout Yummy Chip - bgianfo) tt_um_sirajmuhammad_bpsk_mod (BPSK Baseband Modulator) tt_um_K_coder_9 (TENs device frequency controller) tt_um_wokwi_469758119198926849 (LL_6BitShiftRegister_ToggleEnabledFeedback) tt_um_Asaadkhex_6x6u (6x6 UART Bussbar Switch) tt_um_wokwi_469809198944364545 (tt8-8bit-cpu Copy) tt_um_wokwi_469710279607305217 (Tiny Tapeout Submission KL - SiliDize) tt_um_wokwi_469629799092815873 (2:1 Mux with differential outputs) tt_um_poundbrad_reciprocal_counter (Two-Channel Reciprocal Counter) tt_um_joonatanalanampa_cordic (CORDIC-1) tt_um_x4ntha_nova (Data General Nova 1200 CPU) tt_um_quick_bus (quick_bus) tt_um_wokwi_470058539448408065 (Nigel's Tiny Tapeout Project) tt_um_wokwi_470058244557293569 (Tiny Tapeout Kabisan) tt_um_wokwi_470058241869790209 (Abdi's desgin) tt_um_wokwi_470060107756808193 (Sukhraj Deol's Chip) tt_um_wokwi_470058578588614657 (The Chip of Master George Stead) tt_um_wokwi_470069286344622081 (Tiny Tapeout ISHA) tt_um_ucl_display (Flashing... lights) tt_um_wokwi_470058746279043073 (Arihant's first Wokwi design) tt_um_wokwi_470060103260512257 (Tiny Tapeout Jabriel Copy) tt_um_wokwi_470069460157662209 (haadi's tiny tapeout) tt_um_wokwi_470058418706939905 (Kitty) tt_um_wokwi_470058490118136833 (Iris) tt_um_wokwi_470060098828179457 (Temz_ tiny tapeout) tt_um_wokwi_470058023187099649 (Osman WOKWI project 1) tt_um_wokwi_470057988621827073 (Viraj Tiny Template Full Adder TEST) tt_um_wokwi_470069802034377729 (Tiny Tapeout Template Copy) tt_um_wokwi_470070136685362177 (full adder) tt_um_wokwi_470070449402211329 (Anastasia Copy (2)) tt_um_wokwi_470059864883484673 (Keyaan’s first Wokwi design) tt_um_wokwi_470071200164912129 (full adder tiny tapeout Copy) tt_um_wokwi_470060671178857473 (SBUSixth First Chip Design Mentored by Tiny Tapeout) tt_um_wokwi_470099562753182721 (Isaac Tiny Tapeout) tt_um_wokwi_470120538476737537 (efwz8voices) tt_um_lelo_gr01_analogicus (LELO-GR01) tt_um_lelo_gr04_analogicus (LELO-GR04) tt_um_lelo_gr02_analogicus (LELO-GR02) tt_um_pump_out (60 Hz RMS Pump-Out Controller) tt_um_urish_simon (Simon Says memory game) tt_um_lelo_gr03_analogicus (LELO-GR03) tt_um_wokwi_470299374901578753 (Shrimp) tt_um_vga_clock (VGA clock) tt_um_frequency_counter (Frequency counter) tt_um_z2a_rgb_mixer (RGB Mixer demo) tt_um_mattvenn_r2r_dac_3v3 (Analog 8 bit 3.3v R2R DAC) tt_um_rebeccargb_universal_decoder (Universal Binary to Segment Decoder) tt_um_rebeccargb_hardware_utf8 (Hardware UTF Encoder/Decoder) tt_um_rebeccargb_intercal_alu (INTERCAL ALU) tt_um_rebeccargb_vga_pride (VGA Pride) tt_um_ogggggish_ota_ldo (SSF Capless LDO) tt_um_hariri4534_audioplayback (audioplayback) tt_um_wokwi_470637150792846337 (Joni - Tiny Tapeout Teardown2026 Workshop) tt_um_wokwi_470635013242210305 (Tom's first Wokwi design) tt_um_wokwi_470635780983408641 (Tiny Tapeout-AyeshaTeardown26) tt_um_wokwi_470639152626282497 (KeKoaM Tiny Tapeout) tt_um_wokwi_470637073520124929 (Tiny Tapeout workshop) tt_um_toby43479_iox (IO Expander with PWM) tt_um_wokwi_470635764113915905 (Divider Demo) tt_um_wokwi_470635580461052929 (Mann-teardown-project) tt_um_wokwi_470639672984256513 (KCs 001 TinyTapeout Design) tt_um_wokwi_470635507665754113 (Tiny Tapeout Template Copy) tt_um_wokwi_470637047364443137 (Pixel-Curio-Chip) tt_um_terihear_tinytearout (TinyTearout) tt_um_wokwi_470643025042834433 (TT 2026) tt_um_wokwi_470637360757626881 (Tiny Tapeout Template Copy) tt_um_wokwi_470635627278929921 (Tiny Tapeout Workshop) tt_um_wokwi_474471160110403585 (Cylon-Scanner) tt_um_wokwi_470646659230201857 (bloopbloop) tt_um_pthomas_sigma_delta (Continuous-Time Sigma-Delta ADC (1st order)) tt_um_sky_tpu_3x3 (Sky TPU 3x3) tt_um_tpcannon7_fir (tinyfir) tt_um_bruniliomuy_top (Fir_Filter) tt_um_semiqa_diff_opamp (Diff-In-Diff-Out-OpAmp) tt_um_TinyProcessor_naiyar_ (TinyProcessor) tt_um_CCDmos3D (ADC for CCDmos3D pixel) tt_um_snn_lif_neuron (snn_lif_neurons) tt_um_galaguna_NanoSys_fit (Nano-120_CPU@ler.uam.mx) tt_um_rowles_regime (Single-Bit Macro Regime Classifier) tt_um_rowles_fedmodel (The Fed Model (F1/F2)) tt_um_sky26c (tt_sky26c) tt_um_aka_regfile_ecc (regfile_ecc) tt_um_fwilson12_mac (int8 MAC) tt_um_davidbroughsmyth_ecg_sar12 (heart_monitor_adc_art) tt_um_foxworks_picorv32 (TCD Foxworks PicoRV32) tt_um_saltworks_ndf_c32 (Neural dataflow fabric — bit-serial MAC cells on a self-routing switch) tt_um_yjeum11 (DTMF (Touch-Tone) decoder) tt_um_vedic_mult (4-bit Vedic Multiplier) tt_um_atx_phased_interferometer (Acoustic Interferometer) tt_um_tilesos_dual_adc (Dual-Path Noise-Shaping ADC) tt_um_darga_cirom (Darga CiROM digital read + ternary MAC) tt_um_azara_cirom (Azara CiROM ternary read) tt_um_spi_reg_bank (8-bit Modified RISC-V) tt_um_aialaqili_updown_counter (4-bit Up/Down Counter) tt_um_noahzperez29_riscv_core (Noah RISC-V Core) tt_um_fp8_fpu (FP8 (E4M3) Floating-Point Unit) tt_um_costinemanuelv_gps_daily_trigger (GPS Daily Trigger) tt_um_ja_achtung_1x1 (JA Achtung Compact) tt_um_ja_achtung_1x2 (JA Achtung Full) tt_um_pwm_spice (spice-pwm-tapeout) tt_um_wecallemjazzyfact_bgr_ldo (BGR + LDO 3.3V/1.8V Integrated IP) tt_um_lelo_temp_wulffern (LELO-TEMP) tt_um_wokwi_472389622799861761 (3-Bit 101 Pattern Detector) tt_um_LnL_SoC (Lab and Lectures SoC) tt_um_dash_lucas_risc (risc_processor) tt_um_serdes_ephotonics (UCIe-style SERDES with analog TX driver & RX slicer) tt_um_joram200 (Kalman Filter Hardware Accelerator) tt_um_colbywonn_poly_synth (Poly Synth v1.0) tt_um_nobleg30_uart_vga_scroller (UART VGA Text Scroller) tt_um_multi_precision_mult (Multi-Precision Multiplier) tt_um_pratibha_munnangi_qkt_mac (QKT MAC Accelerator) tt_um_akankaan_bf16_fma (BF16 Fused Multiply-Add (FMA)) tt_um_rtfce (RTFCE - Reconfigurable Temporal Fault/Constraint Engine) tt_um_hdc_classifier (HDC Classifier) tt_um_preethi8a_adaptive_lfsr_prng (Self-Seeding Adaptive 16-bit Galois LFSR PRNG) tt_um_dilip951_cpu_systolic_array (Reconfigurable mixed-precision 2x2 systolic MAC array) tt_um_pqc_ntt_bfly (Crypto-Agile NTT Butterfly (ML-KEM / ML-DSA / FN-DSA)) tt_um_mlkem_coefficient_integrity (Fault-Aware Constant-Time FO Backend for ML-KEM) tt_um_vital_ap (VITAL-AP: Adaptive Pixel Register) tt_um_olaf8 (OLAF-8: Bounded-Memory Online Adaptive Fuzzy Inference) tt_um_Median_MAD (Streaming Median-MAD Estimator) tt_um_tnt_mosbius (tnt's variant of SKY130 mini-MOSbius) tt_um_undip_ann_q610 (UNDIP ANN Accelerator (SPI + bring-up self-test)) tt_um_cpu8 (CPU8) tt_um_vaishnavipatil5_configurable_cam (Configurable CAM with Masked Pattern Matching and Priority Resolution) tt_um_gina_env_monitor (Environmental Mapping Processor) tt_um_manasvibhat_bloom_filter (Bloom Filter Membership Tester) tt_um_amazing_sage_snn (LIF Neuron SNN) tt_um_nkanderson_lut_snn (LUT Spiking Network Classifier) tt_um_bigmanraffa_clm (Clementine: 4-lane int8 SIMT GPU) tt_um_adityarprasad_fft (Adaptive-Precision FFT) tt_um_oscillating_bones (Oscillating Bones) tt_um_silicon_edge_ns_sar_adc (NS SAR ADC) tt_um_sishi888_tinymind (TinyMind SoC) tt_um_afra_123_ecc_memory (Runtime-Reconfigurable ECC Memory) tt_um_kenchangh_mnist (MNIST Digit Recognition) tt_um_ece298a_8_bit_cpu_top (8-Bit CPU) tt_um_libormiller_SIMON_V2 (SIMON V2) tt_um_WaiMingLee888_nanov_1tile (NanoV RV32E one-tile RISC-V processor) tt_um_four_bit_nn_accel (4-bit Neural Network Accelerator) tt_um_rsa_simple (RSA Simple Encryptor) tt_um_synapticrw_lif_neuron (LIF Neuron (SynapticRW Teardown 2026)) tt_um_smunigan_ipv4_filter (IPv4 Header Filter) tt_um_jjy_spi_watchdog (SPI-Configurable Watchdog Timer) tt_um_osian_beam_controller (Programmable Metasurface Beam Controller) tt_um_namramazhar_popcnt_shiftreg (17-bit Wallace-tree POPCNT with shift-register input) tt_um_obookstay_puf (An arbiter PUF) tt_um_arminkardovic_montenegro_securekey (Montenegro SecureKey) tt_um_rcyaon_droop (All-Digital Supply Droop Detector) tt_um_ctw_spms (CTW-SPMS — Programmable Smart Power Management & Supervisor) tt_um_taiwoopesade_tempo_detector_sky26c (Hardware Audio Tempo Detector) tt_um_wokwi_470059878406973441 (Ehan's first TinyTapeout Project) tt_um_wokwi_470637170309995521 (My First Wokwi Thing!) tt_um_wokwi_470637401137246209 (Teardown Tiny Tapeout) tt_um_wokwi_469443433165025281 (Tiny Tapeout First Design Beth Plummer) tt_um_wokwi_472423526521678849 (4-bit to 5x7 Matrix Decoder for Tiny Tapeout) tt_um_wokwi_470057961258181633 (Tiny Tapeout Template Kavana) tt_um_wokwi_470057993933917185 (ivane- Tiny Tapeout (full adder)) tt_um_wokwi_470088776251343873 (training_project_kaylem) tt_um_neuropong (NeuroPong) tt_um_tamagotchi (TamaGotThis) tt_um_group02_seethebeat (SeeTheBeat) tt_um_kul_chromechain (Chrome Chain) tt_um_baked_weights (Baked-Weights Shakespeare GPT) tt_um_gilangfajrul_sar_adc (sar-adc) tt_um_Logy_FMAC (FMAC) tt_um_porkfreezer_rrio_opamp (RRIO Op-amp) tt_um_diff_engine (DSLX finite_difference) tt_um_dragonochi (WISH) tt_um_siliconsonics (ultrasonic sonar: range and bearing) tt_um_kul_conway (Interactive Conway's Game of Life) tt_um_algofoogle_ttsky26c_analog (Assorted analog in 1 tile) tt_um_mariavictoriaalm_qubit_sim ( tt-2qubit-sim) tt_um_andre_dpe (Dot product engine) tt_um_rmranjitkarNULL_pong_top (last_minute_Pong) tt_um_SAR_ADC (CTW LDO and Dynamic Comparator) tt_um_fabulous_sky_26c (Tiny FABulous FPGA) tt_um_tomvdsch_tiny32_soc (Tiny32 RV32IMA Zephyr-target SoC) tt_um_np523_pong (Pong) tt_um_usfq_adc_procmon (USFQ 8-bit Tracking ADC and Process Variation Monitor) tt_um_rangfuu_alu (Tiny ALU PD) tt_um_wokwi_473800139156677633 (Tiny Snake with PRISM 8) tt_um_mini_nn (Four-MAC Core Neural Network Inference Engine) tt_um_kianv_rv32_regfile (KianV uLinux RISC-V regfile edition) tt_um_2048_vga_game (2048 sliding tile puzzle game (VGA)) tt_um_urish_rings (VGA Rings) tt_um_silicon_art_vga_screensaver (VGA Screensaver with Silicon Art ROM) tt_um_rom_vga_screensaver (VGA Screensaver with embedded bitmap ROM) tt_um_krisjdev_manchester_baby (Manchester Baby) tt_um_urish_sic1 (SIC-1 8-bit SUBLEQ Single Instruction Computer) tt_um_ThomasCowieEngineering_LMC (Little Man Computer CPU) tt_um_pranavUl_ascon_aead128 (Ascon bit-serial permutation engine) tt_um_orca (ORCA — Online Reconfigurable Circuit with Adaptation) tt_um_krisjdev_artwork (Silicon Artwork) tt_um_htfab_caterpillar (Simon's Caterpillar) tt_um_htfab_vga_tester (Video mode tester) Available Available Available Available Available Available Available Available Available Available