Skip to content

Source: moonlive_asm_xtensa.h

moonlive_asm_xtensa

The classic ESP32 and S3 backend, written against the same named-instruction interface as the host one.

Only the encodings and the windowed ABI differ: the emitted routine opens with entry and returns with retw.n, the host arguments arriving in a2 upward. Branch displacements are back-patched, so no offset is hand-computed.

Enumerations

Name Description
[Reg] Ten registers, mapping to a2 upward; the top of the window is the return path, not a register.
[Cond] A branch condition, holding the ones the IR needs.

Reg

enum Reg

Ten registers, mapping to a2 upward; the top of the window is the return path, not a register.

Value Description
R0
R1
R2
R3
R4
R5
R6
R7
R8
R9
R10
R11
R12
R13
kRegCount
R0
R1
R2
R3
R4
R5
R6
R7
R8
R9
R10
R11
R12
R13
kRegCount
R0
R1
R2
R3
R4
R5
R6
R7
R8
R9
kRegCount

Cond

enum Cond

A branch condition, holding the ones the IR needs.

Value Description
Lo
Hs
Ne
Ge
Lo
Hs
Lo
Hs

Functions

Return Name Description
const uint8_t * xtRegMap The register map, for the device-codegen test.

xtRegMap

const uint8_t * xtRegMap(uint8_t & count)

The register map, for the device-codegen test.

XtensaAssembler

class XtensaAssembler
src/platform/esp32/moonlive_asm_xtensa.h:46

Public Methods

inline ~XtensaAssembler() : Frees the buffer only when this emitter owns it; a copy would double-free, so copying is deleted.

inline explicit XtensaAssembler(size_t cap = kCodeCap) : Allocate a cap-byte code buffer, sized per script because backends differ by up to 1.9x.

inline XtensaAssembler(uint8_t * out, size_t cap) : Emit straight into the caller's staging buffer, which halves a compile's transient heap.

XtensaAssembler(const XtensaAssembler &) = delete : Not copied: a copied owner would free the same buffer twice.

XtensaAssembler & operator=(const XtensaAssembler &) = delete : Not copy-assigned, for the same reason.

inline void finalize() : Resolve every fixup against its bound label; call it once, after the last instruction.

inline void alignForEntry() : Pad to the four-byte boundary a function entry needs, which this architecture requires twice over.

inline const uint8_t * bytes() const : The finished bytes, valid only after finalize.

inline size_t size() const : How many bytes were emitted.

inline bool overflowed() const : Whether any write was dropped for want of room.

void prologue(uint8_t slots = 0) : Open the routine's frame, widened to carry slots spilled values; must be the first instruction.

void spillStore(Reg r, uint8_t slot) : Write a register into a spill slot.

void spillLoad(Reg r, uint8_t slot) : Read a spill slot back into a register.

void slotAddr(Reg d, uint8_t slot) : Address one slot, which is how a call builds its argument block.

Label newLabel() : A fresh label, to be bound once and branched to any number of times.

void bind(Label l) : Fix this label's position at the current offset.

void movPtr(Reg d, const void * p) : A full-width address into a register (ConstPtr).

void movImm(Reg d, int32_t imm) : An immediate of any width into a register, through the narrow move where it fits.

void movReg(Reg d, Reg a) : Mov.n aD, aA.

void addImm(Reg d, Reg a, int32_t imm) : Addi.n aD, aA, #imm (1..15).

void addReg(Reg d, Reg a, Reg b) : add.n aD, aA, aB.

void mulReg(Reg d, Reg a, Reg b) : Mull aD, aA, aB.

void mulhi(Reg d, Reg a, Reg b) : Mulsh aD, aA, aB: the SIGNED high 32 bits.

void shlImm(Reg d, Reg a, uint8_t n) : slli aD, aA, #n (1..31).

void sarImm(Reg d, Reg a, uint8_t n) : srai aD, aA, #n (0..31), arithmetic.

void shrImm(Reg d, Reg a, uint8_t n) : LOGICAL right shift (srli / extui).

void store8(Reg base, Reg off, Reg val) : S8i via computed address (add then s8i,0).

void load8(Reg d, Reg base, int32_t imm) : L8ui aDst, aBase, #imm: a control read.

void load32(Reg d, Reg base, int32_t imm) : L32i.n aDst, aBase, #imm: a whole 4-byte slot.

void store32(Reg base, int32_t imm, Reg val) : s32i.n aVal, aBase, #imm (offset IMMEDIATE).

void load32Idx(Reg d, Reg base, Reg off) : add.n tmp,base,off ; l32i.n d,tmp,0.

void store32Idx(Reg base, Reg off, Reg val) : add.n tmp,base,off ; s32i.n val,tmp,0.

void load8Idx(Reg d, Reg base, Reg off) : add.n tmp,base,off ; l8ui d,tmp,0.

void branchIfZero(Reg a, Label l) : Beqz aA, l (nLights==0 guard).

void branchGeU(Reg a, Reg b, Label l) : Bgeu aA, aB, l (Bounds: skip if a>=b).

void branchGeS(Reg a, Reg b, Label l) : Bge aA, aB, l (a script's own comparison).

void branchNe(Reg a, Reg b, Label l) : Bne aA, aB, l (loop test).

void call(Reg d, Reg a, Reg b, Reg c, const void * fn) : Windowed call8 to a host built-in.

void callLabel(Label l, Reg d = R0, bool take = false) : Call a function in this block by label: the script-to-script call.

void epilogue() : Retw.n.

void retValue(Reg a) : Park a where the ABI returns a value, before the epilogue tears the frame down.

Public Static Attributes

constexpr uint8_t kMaxSpillSlots = kTotalSlots : The allocator's slot range plus the parked host arguments.

Public Types

using RegType = Reg : The register type the shared lowering works in, each backend's Reg being its own enum.

More info

The window is the return path

Ten registers map to a2 upward, and the top of the window is not a general register. A routine opened with entry returns through retw.n, which reads the caller's linkage from the window's top. Using those as registers corrupted that linkage and returned to a garbage address the moment a scripted layout ran.

Entry alignment, twice over

The entry instruction must sit on a four-byte boundary, and a call encodes its target in four-byte units, so an unaligned callee cannot be expressed at all. Instructions here are two or three bytes, so a function following another lands anywhere and needs the pad.

Why the buffer is heap

The assembler is a stack local, so a cap-sized member put 2 KB on the compile chain's stack and overflowed the task on a classic ESP32.