Bumble BytesFPGA and ASIC IP
Bumble Bytes · FPGA and ASIC IP

RankFilter: verified median and rank-filter IP for video.

Streaming 3×3 to 7×7 median, percentile and Bayer defect-pixel-correction cores for FPGA and ASIC. A 5×5 median needs 52 comparators per pixel, against 85 for the best published design. Every comparator network is proved correct in Lean 4, and the cores run on Intel Cyclone V silicon.

AXI4-Stream or video timing 1, 2 or 4 pixels per clock 8 to 16-bit pixels
Running live on an Analogue Pocket's Intel Cyclone V at 148.5 MHz: noisy input on the left of the yellow line, the filtered output on the right.
52 vs 85
Comparators per pixel for a 5×5 median, against the best known single-window design
4K60
At 4 pixels per clock on the slowest ECP5, Artix-7 and Cyclone V speed grades
−22%
Less logic for a 7×7 median on Cyclone V than the strongest baseline we built, signed off in Quartus
0 errors
In about 22 million self-test runs on Cyclone V silicon at up to 170 MHz
Overview

Less logic for the same filter, with proof that it is right.

A streaming filter computes a new window every clock, and neighbouring windows overlap almost completely. RankFilter's comparator networks reuse that overlap instead of recomputing it. We found them by computer search and proved each one correct, for every input, in the Lean 4 proof assistant.

AXI4-Streamin Line buffersK−1 lines, BRAMborders Column sortonce percolumn Shared networkreused acrosswindows Window networkper output pixel,Lean-proved OutputDPC select(Bayer), FIFO AXI4-Streamout Credit-based flow control: formally proved never to drop or overwrite data, under any back-pressure

The honey-coloured stages are the comparator networks: found by search, proved in Lean 4.

Features

  • Median, 25th and 75th percentile; other ranks generated on request
  • 3×3, 5×5 and 7×7 windows
  • Defect-pixel correction for Bayer and monochrome sensors with a run-time threshold; the monochrome core also takes a static bad-pixel map
  • 1, 2 or 4 pixels per clock
  • AXI4-Stream with full back-pressure, or a smaller video-timing version (1 pixel per clock) for sensor streams that cannot stall
  • Frame size set at run time; edge pixels replicated or replaced by a constant
  • Pixel width is a parameter, verified at 8, 10, 12 and 16 bits
  • Plain synthesisable Verilog-2005, with no vendor primitives required

Applications

  • Hot and dead pixel correction on Bayer colour and monochrome sensors, including thermal and SWIR, from a threshold or a factory defect map
  • Impulse (salt-and-pepper) noise removal
  • Pre-processing for machine vision and industrial inspection
  • Rank-order and grey-scale morphological filtering
  • Scientific, medical and broadcast imaging

Supported devices

Any FPGA or ASIC flow that accepts Verilog. Measured on Lattice ECP5, AMD Artix-7 and Intel Cyclone V, and in SkyWater 130 nm and ASAP7 7 nm ASIC flows.

Comparators per output pixel

Fewer comparators than published designs

Comparator count drives logic and power. RankFilter is lower than Adams' 2021 separable networks and the best single-window designs at every shipped size, apart from a tie at 3×3 and one pixel per clock.

WindowRankFilterAdams 2021Single window
3×3, 1 px/clk131319
5×5, 1 px/clk526385
5×5, 4 px/clk3742.7585
7×7, 1 px/clk140187231
7×7, 2 / 4 px/clk92117 / 93.25231

Single window: Dobbelaere's published records. Adams 2021 (ACM Transactions on Graphics) is a software method; its 4-pixel figures are 2×2 output tiles, and 117 is our count of his method. The networks that HLS vision libraries typically unroll need 234 for a 5×5 median.

Performance

Measured size and speed

Complete filters with AXI4-Stream interface, line buffers and border handling, 8-bit pixels. Fmax in MHz.

VariantECP5-85F −6
LUT4 / FF / EBR
FmaxArtix-7 35T −1
LUT / FF
FmaxCyclone V C8
ALM / M10K
Fmax
3×3 median928 / 691 / 2154861 / 773203371 / 7180
3×3 median, 4 px/clk2233 / 1783 / 81641968 / 1752183939 / 8180
5×5 median2056 / 1905 / 41642105 / 1687199896 / 10180
5×5 median, 4 px/clk5184 / 6600 / 151615354 / 37611852216 / 28164
5×5 25th percentile1783 / 1622 / 41701838 / 1476194769 / 10180
7×7 median4455 / 4612 / 61594613 / 35221961941 / 20171
7×7 median, 4 px/clk11452 / 14490 / 2314312112 / 74381565276 / 46160
Bayer DPC (same-colour 3×3)1092 / 1019 / 41661060 / 986181435 / 9164
Bayer DPC, 4 px/clk2673 / 2458 / 151542593 / 21201741115 / 11180

ECP5 and Artix-7: median of 5 placement seeds with open-source tools (yosys, nextpnr), not vendor sign-off. Cyclone V: Quartus Prime Lite timing analysis, worst slow corner, median of 3 seeds; 180 MHz is the device's clock-network limit here. 1080p60 needs 148.5 MHz at 1 pixel per clock and 4K60 needs 124.4 MHz at 4 pixels per clock; every 4-pixel variant meets 4K60 on every seed at 8 and 12 bits. 12-bit pixels add 22 to 47 % logic. The product brief and datasheet list every variant.

Against other cores5×57×7
Strongest baseline we built, same wrapper and tools (ECP5 LUT4)−11 %−21 %
Same, Cyclone V ALMs (Quartus)−18 %−22 %
Lattice Median Filter IP, published LatticeECP3 LUTs2908 → 205611,536 → 4455
ASIC area / power, SkyWater 130 nm at 200 MHz (cores)−16 / −18 %−25 / −39 %

The strongest baseline is shared column sorting with the best selection stage we found, built through the same wrapper and tools. Lattice's figures come from its product page (IP v1.0, Diamond and Synplify): a different family and toolchain, shown for scale. At 3×3 Lattice's published core is smaller (680 LUTs against 928) because our AXI4-Stream wrapper dominates at that size.

For sensor streams that cannot stall, the video-timing version of the 3×3 filter uses 567 ECP5 LUT4s (470 at a fixed frame size), against 689 for the most widely copied open-source 3×3 median, which has no border handling or run-time frame size. ASIC results use open PDKs.

Verification

Checked at every level, from proof to silicon.

A filter that is wrong only on rare inputs can pass every demo and still fail in the field. Every claim below is backed by a repeatable script.

  • Every comparator network is proved in Lean 4 to select the right rank for every input and any pixel type.
  • The generated RTL is simulated on every binary window (225 for 5×5; every sorted-column state for 7×7), which by the 0-1 principle covers all pixel widths.
  • SymbiYosys proves the flow control never overflows under any back-pressure and follows the AXI4-Stream rules.
  • Output on photographs is byte-identical to an independent Python reference model, under random back-pressure.
  • On Intel Cyclone V silicon, a continuous self-test of five variants found zero mismatches in about 22 million runs at 74 to 170 MHz. Each run repeats the same two test frames under changing random back-pressure, so this is a soak test of timing and flow control.
Evaluation and licensing

Try it in your own simulator first

A simulation evaluation package is available on request. It drops into your testbench with the same module names and ports as the licensed source.

Evaluation package

  • Gate-level simulation netlists of the FPGA cores
  • Self-checking SystemVerilog testbench and golden vectors
  • Bit-exact Python reference model
  • One-command test script for Verilator or Icarus

Licensed deliverables

  • Verilog source for each variant: an FPGA set and an ASIC-tuned set
  • Datasheet with interface, timing and integration notes
  • Board self-test with expected checksum
  • The Lean proof files

Custom variants

Licences per project or per company. Other window sizes, ranks, pixel widths and pixels per clock can be generated and verified to order, and integration support is available on contract.

Contract engineering

Talk to us about RankFilter

Tell us your device, resolution, frame rate and pixel width, and we'll reply with the matching variant and an evaluation package.

Bumble Bytes also takes on cloud, data and AI contracts