IP CORE FAMILY

LDPC Decoders

Layered and folded quasi-cyclic LDPC decoders for 3GPP 5G NR, IEEE 802.11n, and CCSDS AR4JA, generated from one parameterized architecture, verified bit-exact at every configuration, and timing-closed against a commercial reference IP on the same silicon.

FLAGSHIP

5G NR decoder, measured against a commercial IP

Our flagship configuration (base graph 2, lifting size 384, K = 3840) was placed and routed on the same AMD Kintex UltraScale+ device class the commercial reference IP publishes its numbers on. All figures below are from Vivado Design Timing Summary and utilization reports.

463 MHz
Place-and-route closed clock, xcku13p MEASURED
459 MHz
Commercial reference IP's published clock on the same device class
>1 Gbps
Information throughput with syndrome early termination at operating SNR MEASURED
0 DSP
Zero DSP48 blocks: the entire decoder lives in logic and BRAM MEASURED

One generator, twenty configurations

5G NR specifies two base graphs and 51 lifting sizes. Because our decoder is emitted by a parameterized generator, we regenerated and re-implemented 20 distinct configurations end to end:

Result across 20 generated configurationsValueProvenance
Configurations matching or beating the commercial IP's 459 MHz clock19 / 20Vivado P&R, per config MEASURED
Median clock ratio vs commercial IP1.03×Vivado Design Timing Summary MEASURED
Fastest configuration566 MHzVivado Design Timing Summary MEASURED
Slowest configuration (densest base matrix, routing congestion)408 MHzReported as-is. We publish misses too MEASURED

Fair-comparison note: the commercial IP is a runtime-flexible multi-standard core; our cores are per-configuration specialized netlists, which is precisely what the generator makes economical. Throughput depends on iteration count: figures with early termination are stated at the operating SNR; fixed-iteration apples-to-apples figures are available in the datasheet on request.

SILICON

CCSDS AR4JA: the evaluation twin, proven on a deployed board

Deep-space links get one chance at a frame: there is no retransmission across an interplanetary distance, so the decoder has to be right the first time and it has to keep running. The build below is our evaluation twin: it mirrors the MathWorks CCSDS LDPC Decoder block module by module in scalar mode, which makes it directly comparable against a known reference and makes its correctness auditable line by line. It trades throughput for that comparability, so it is slower than the folded configuration we ship for rate-critical links. It is also the one we took all the way to a board, so it is the one whose numbers we publish.

9 / 9
Code points of CCSDS 131.0-B-3 covered by the evaluation twin MEASURED
675
Frames decoded on the board with 0 bit mismatches vs the golden model MEASURED
81 / 81
Reproducibility gate: 3 fresh boots × 9 configs × 3 trials, zero failures MEASURED
775 / 775
Frames bit-exact in simulation before silicon, all 9 configs MEASURED

Latency, per code point

Measured by interface lockstep at a fixed 8 iterations: the first output bit lands at an exactly predicted cycle for every configuration. These are the evaluation twin's figures, in scalar mode. The folded product configuration processes several lanes per cycle and is correspondingly faster; ask us for figures on your target part rather than scaling these.

ConfigurationTransmitted bitsCycles / iterationFrame latency (cycles)Cycles / info bit
k = 1024, rate 1/220482525,0974.98
k = 1024, rate 2/315363165,0974.98
k = 1024, rate 4/512804445,8655.73
k = 4096, rate 1/28192124022,2175.42
k = 4096, rate 2/3614484417,0014.15
k = 4096, rate 4/5512044412,7773.12
k = 16384, rate 1/232768496088,8415.42
k = 16384, rate 2/324576337667,9774.15
k = 16384, rate 4/520480258457,5453.51

Fixed 8 iterations, no early termination, so these are worst-case figures rather than best-case ones. Early termination on the syndrome check is available and reduces the average iteration count at operating SNR. This twin closed timing at its 208.3 MHz constraint on a ZCU102 (xczu9eg) with WNS +0.091 ns post-route; we are not publishing a clock for the RFSoC part, because the builds we have on that part do not meet that constraint. Device clock and resource figures for the folded product configuration on your target part are produced per configuration on request. We publish only numbers we have measured, on the part we measured them on.

Replay the silicon capture in your browser

Real recorded data from the board: every frame, every configuration, the bit-error comparison and the BER waterfall. No hardware needed.

CATALOG

Family members

5G NR (3GPP TS 38.212)

BG1 & BG2 · all lifting sizes

Layered normalized-min-sum decoding with syndrome early termination. Golden model cross-validated bit-exact against MATLAB 5G Toolbox encode/decode across 40 trials and 7 configurations.

Clock (Z=384 flagship)463 MHz MEASURED
BRAM36 (flagship)90 MEASURED

CCSDS AR4JA

Deep space · satellite

All nine AR4JA code points of CCSDS 131.0-B-3: block lengths 1024, 4096 and 16384 information bits, each at rate 1/2, 2/3 and 4/5. The folded configuration trades parallelism against area for rate-critical or power-constrained platforms; a scalar evaluation twin of the MathWorks reference block is the build we verified bit-exact on real silicon, across all nine.

Frames decoded on silicon675, 0 bit errors MEASURED
Reproducibility gate81 / 81 runs MEASURED

Read the silicon case study

IEEE 802.11n

Wi-Fi · streaming

Continuously streaming decoder accepting one codeword after another with no inter-frame gap. The architectural template behind the whole family.

Clock (Virtex-7)312 MHz MEASURED
Initiation intervalII = 1

Need a different standard, rate, or block size? New configurations are generated and re-verified in days: see IP customization.

Interactive case study

Walk through the full 5G NR decoder design (architecture, verification chain, and the timing-closure campaign) in our interactive engineering report.

Open the report
Request datasheet & evaluation