Skip to content

Source: ParallelSlots.h

ParallelSlots

WS2812 encode for parallel buses: the contract between a parallel driver and its peripheral.

It is named for the wire unit it builds, one pixel-clock slot being one bus word. The peripheral sets that word's width, and both the i80 and Parlio buses use it. A pure data transform with no platform include, so the host test pins it without an ESP32.

Prior art: the technique is hpwit's, Adafruit's and FastLED's, studied rather than copied.

Functions

Return Name Description
uint64_t MM_RAMFUNC transposeBits8x8 inline
void MM_RAMFUNC transposeBits8x8Pair inline The same butterfly on a register pair, which is what the 32-bit device wants.
void MM_RAMFUNC transposeLanes8x8 inline Transpose 8 lane bytes into 8 bit-plane bytes; inactive lanes must be passed as 0.
void MM_RAMFUNC transposeLanes16x8 inline Transpose 16 lane bytes into 8 bit-plane halfwords, for a 16-bit bus.
void MM_RAMFUNC encodeWs2812ParallelSlots inline Encode one row, one light across every lane, in direct mode with one lane per pin.
void MM_RAMFUNC encodeWs2812ShiftLatchPad inline Close a shift-register frame with one latch-only word, without which the register keeps presenting the last data byte.
void encodeWs2812ShiftSlots inline Encode one row through the shift-register expander, indexed by expanded lane.
void MM_RAMFUNC shiftActivePins inline Which pins carry an active strand on each shift cycle, one bus word per cycle.
void MM_RAMFUNC prefillWs2812ShiftConstants inline Write the frame's constant shift-mode words once, so the per-light encoder skips them.
void MM_RAMFUNC encodeWs2812ShiftData inline The per-light data encode, writing ONLY the data word of each slot.

transposeBits8x8

inline

inline uint64_t MM_RAMFUNC transposeBits8x8(uint64_t x)

transposeBits8x8Pair

inline

inline void MM_RAMFUNC transposeBits8x8Pair(uint32_t & lo, uint32_t & hi)

The same butterfly on a register pair, which is what the 32-bit device wants.


transposeLanes8x8

inline

inline void MM_RAMFUNC transposeLanes8x8(const uint8_t * in, uint8_t * out)

Transpose 8 lane bytes into 8 bit-plane bytes; inactive lanes must be passed as 0.


transposeLanes16x8

inline

inline void MM_RAMFUNC transposeLanes16x8(const uint8_t * in, uint16_t * out)

Transpose 16 lane bytes into 8 bit-plane halfwords, for a 16-bit bus.


encodeWs2812ParallelSlots

inline

template<class Slot> inline void MM_RAMFUNC encodeWs2812ParallelSlots(const uint8_t * wire, Slot activeMask, uint8_t channels, Slot * out)

Encode one row, one light across every lane, in direct mode with one lane per pin.


encodeWs2812ShiftLatchPad

inline

template<class Slot> inline void MM_RAMFUNC encodeWs2812ShiftLatchPad(uint8_t latchBit, Slot * out)

Close a shift-register frame with one latch-only word, without which the register keeps presenting the last data byte.


encodeWs2812ShiftSlots

inline

template<class Slot> inline void encodeWs2812ShiftSlots(const uint8_t * wire, uint64_t activeMask, uint8_t physPins, uint8_t latchBit, uint8_t outPerPin, uint8_t channels, Slot * out)

Encode one row through the shift-register expander, indexed by expanded lane.


shiftActivePins

inline

template<class Slot> inline void MM_RAMFUNC shiftActivePins(uint64_t activeMask, uint8_t physPins, uint8_t outPerPin, Slot(&) out)

Which pins carry an active strand on each shift cycle, one bus word per cycle.


prefillWs2812ShiftConstants

inline

template<class Slot> inline void MM_RAMFUNC prefillWs2812ShiftConstants(uint64_t activeMask, uint8_t physPins, uint8_t latchBit, uint8_t outPerPin, uint8_t channels, uint32_t rows, Slot * out)

Write the frame's constant shift-mode words once, so the per-light encoder skips them.


encodeWs2812ShiftData

inline

template<class Slot> inline void MM_RAMFUNC encodeWs2812ShiftData(const uint8_t * wire, uint64_t activeMask, uint8_t physPins, uint8_t latchBit, uint8_t outPerPin, uint8_t channels, Slot * out)

The per-light data encode, writing ONLY the data word of each slot.

Variables

Return Name Description
constexpr uint8_t kPinExpanderOutputs constexpr Outputs per physical data pin when a 74HCT595 expander is fitted; one register is eight.

kPinExpanderOutputs

constexpr

constexpr uint8_t kPinExpanderOutputs = 8

Outputs per physical data pin when a 74HCT595 expander is fitted; one register is eight.

More info

The three slots

Every WS2812 data bit becomes three bus slots. A 1 bit is high for two of them, a 0 bit for one:

slot 0: activeMask       every active lane high, the pulse start
slot 1: data bits & mask lane L's current bit at bus bit L
slot 2: 0x00             every lane low, the pulse tail
A short lane leaves both slot 0 and slot 1 once its lights run out. It then idles low rather than flashing white, its mask bit being clear.

What the encoders take

Slot is deduced from the call: a uint8_t for an 8-lane bus and a uint16_t for a 16-lane one. One bus word is one slot, bit L being data line L, so the word width is the lane count and the 8-bit call sites need no change.

wire holds the corrected wire bytes lane-major, indexed as wire[lane * channels + channel]. Only lanes set in the active mask are read, so an inactive lane may hold garbage. The lane stride is channels rather than a fixed four, so a light of any channel count works.

The shift-register path

The same wire contract is fanned out through \'595 expanders, so each data pin drives several strands. The fan-out multiplies the slot count rather than the pin count.

The part is a 74HCT and not a plain 74HC. At 5 V the plain part's input threshold sits above 3.3 V, so a high is not guaranteed to read as a 1. The symptom is flaky strands rather than dead ones.

The fan-out grows on pins rather than cascade depth, because depth doubles the required pixel clock and there is no exact divide in the band that would need. The board details are on the drivers page.

Two shift-mode optimizations worth knowing

The prefill writes the frame's constant words once, because only one of a slot's three words carries pixel data. Rewriting the other two per light would burn two thirds of the encoder's stores.

The active-pin helper is deliberately not called by the whole-slot encoder. There the mask test is fused with the lane gather, so routing it through the helper would walk the pins twice.

Why the 64-bit mask is split into halves

Splitting once makes every strand test a 32-bit variable shift, which is a single Xtensa instruction. The direct form, masking against a 1ULL shifted by a runtime value, compiles on the 32-bit Xtensa to an __ashldi3 library call of roughly 50 cycles. At 144 tests a row that call cost about 7,000 cycles a row, the encoder's dominant cost, attributed on the bench. It is the sibling of the 32-bit SWAR lesson: 64-bit operations synthesize on this target.

The SWAR pack follows the same shape. Lane p is byte p of the 8x8 matrix: byte p of the first register when p is under four, byte p - 4 of the second otherwise. A 16-lane bus needs a second pair for pins 8 to 15, which the 8-bit path never touches and the compiler drops.

What a prefill run covers

The rows argument is how many rows share one active set. The caller re-prefills per run of rows with the same mask, because an exhausted strand changes the mask at the row where it runs out. Uniform-length strands, the common case, are a single run.

Why the transpose is the emit

Each shift cycle's eight bit-planes are stored the moment they are computed, while they are still in registers. Staging them in an array first cannot work on this target. Eight cycles of two words each exceeds the register file, so every plane spills to the stack and is reloaded once per bit. Measured on an S3, that staging cost 97 word loads and stores a light against the 17 byte-stores of actual output.

This is the lesson the lane array taught one level down, where removing it took a light from 8.85 to 6.19 microseconds, and staging the planes was the identical pattern. hpwit's driver stages nothing either, his transpose storing straight into the DMA buffer at its pulse offsets; that was studied rather than copied.

The price is a strided store, one cycle's eight planes landing a bit-stride apart rather than contiguously. That is one address add per store, far cheaper than a spill and a reload.

Why the transpose is SWAR

The data slot is an 8 by 8 bit-matrix transpose, and the measured render-loop hot spot. It is around 85% of the driver frame at scale. So it uses the branch-free delta-swap from Hacker's Delight rather than a per-bit gather loop. Same result, no table, far fewer operations.

The 8 by 8 bit-transpose on the packed representation, which keeps it in registers.