Source:
moonlive_asm_xtensa.h
moonlive_asm_xtensa¶
The classic ESP32 and S3 backend, written against the same named-instruction interface as the host one.
Only the encodings and the windowed ABI differ: the emitted routine opens with entry and returns with retw.n, the host arguments arriving in a2 upward. Branch displacements are back-patched, so no offset is hand-computed.
Enumerations¶
| Name | Description |
|---|---|
[Reg] |
Ten registers, mapping to a2 upward; the top of the window is the return path, not a register. |
[Cond] |
A branch condition, holding the ones the IR needs. |
Reg¶
Ten registers, mapping to a2 upward; the top of the window is the return path, not a register.
| Value | Description |
|---|---|
R0 |
|
R1 |
|
R2 |
|
R3 |
|
R4 |
|
R5 |
|
R6 |
|
R7 |
|
R8 |
|
R9 |
|
R10 |
|
R11 |
|
R12 |
|
R13 |
|
kRegCount |
|
R0 |
|
R1 |
|
R2 |
|
R3 |
|
R4 |
|
R5 |
|
R6 |
|
R7 |
|
R8 |
|
R9 |
|
R10 |
|
R11 |
|
R12 |
|
R13 |
|
kRegCount |
|
R0 |
|
R1 |
|
R2 |
|
R3 |
|
R4 |
|
R5 |
|
R6 |
|
R7 |
|
R8 |
|
R9 |
|
kRegCount |
Cond¶
A branch condition, holding the ones the IR needs.
| Value | Description |
|---|---|
Lo |
|
Hs |
|
Ne |
|
Ge |
|
Lo |
|
Hs |
|
Lo |
|
Hs |
Functions¶
| Return | Name | Description |
|---|---|---|
const uint8_t * |
xtRegMap |
The register map, for the device-codegen test. |
xtRegMap¶
The register map, for the device-codegen test.
XtensaAssembler¶
src/platform/esp32/moonlive_asm_xtensa.h:46Public Methods¶
inline ~XtensaAssembler()
: Frees the buffer only when this emitter owns it; a copy would double-free, so copying is deleted.
inline explicit XtensaAssembler(size_t cap = kCodeCap)
: Allocate a cap-byte code buffer, sized per script because backends differ by up to 1.9x.
inline XtensaAssembler(uint8_t * out, size_t cap)
: Emit straight into the caller's staging buffer, which halves a compile's transient heap.
XtensaAssembler(const XtensaAssembler &) = delete
: Not copied: a copied owner would free the same buffer twice.
XtensaAssembler & operator=(const XtensaAssembler &) = delete
: Not copy-assigned, for the same reason.
inline void finalize()
: Resolve every fixup against its bound label; call it once, after the last instruction.
inline void alignForEntry()
: Pad to the four-byte boundary a function entry needs, which this architecture requires twice over.
inline const uint8_t * bytes() const
: The finished bytes, valid only after finalize.
inline size_t size() const
: How many bytes were emitted.
inline bool overflowed() const
: Whether any write was dropped for want of room.
void prologue(uint8_t slots = 0)
: Open the routine's frame, widened to carry slots spilled values; must be the first instruction.
void spillStore(Reg r, uint8_t slot)
: Write a register into a spill slot.
void spillLoad(Reg r, uint8_t slot)
: Read a spill slot back into a register.
void slotAddr(Reg d, uint8_t slot)
: Address one slot, which is how a call builds its argument block.
Label newLabel()
: A fresh label, to be bound once and branched to any number of times.
void bind(Label l)
: Fix this label's position at the current offset.
void movPtr(Reg d, const void * p)
: A full-width address into a register (ConstPtr).
void movImm(Reg d, int32_t imm)
: An immediate of any width into a register, through the narrow move where it fits.
void movReg(Reg d, Reg a)
: Mov.n aD, aA.
void addImm(Reg d, Reg a, int32_t imm)
: Addi.n aD, aA, #imm (1..15).
void addReg(Reg d, Reg a, Reg b)
: add.n aD, aA, aB.
void mulReg(Reg d, Reg a, Reg b)
: Mull aD, aA, aB.
void mulhi(Reg d, Reg a, Reg b)
: Mulsh aD, aA, aB: the SIGNED high 32 bits.
void shlImm(Reg d, Reg a, uint8_t n)
: slli aD, aA, #n (1..31).
void sarImm(Reg d, Reg a, uint8_t n)
: srai aD, aA, #n (0..31), arithmetic.
void shrImm(Reg d, Reg a, uint8_t n)
: LOGICAL right shift (srli / extui).
void store8(Reg base, Reg off, Reg val)
: S8i via computed address (add then s8i,0).
void load8(Reg d, Reg base, int32_t imm)
: L8ui aDst, aBase, #imm: a control read.
void load32(Reg d, Reg base, int32_t imm)
: L32i.n aDst, aBase, #imm: a whole 4-byte slot.
void store32(Reg base, int32_t imm, Reg val)
: s32i.n aVal, aBase, #imm (offset IMMEDIATE).
void load32Idx(Reg d, Reg base, Reg off)
: add.n tmp,base,off ; l32i.n d,tmp,0.
void store32Idx(Reg base, Reg off, Reg val)
: add.n tmp,base,off ; s32i.n val,tmp,0.
void load8Idx(Reg d, Reg base, Reg off)
: add.n tmp,base,off ; l8ui d,tmp,0.
void branchIfZero(Reg a, Label l)
: Beqz aA, l (nLights==0 guard).
void branchGeU(Reg a, Reg b, Label l)
: Bgeu aA, aB, l (Bounds: skip if a>=b).
void branchGeS(Reg a, Reg b, Label l)
: Bge aA, aB, l (a script's own comparison).
void branchNe(Reg a, Reg b, Label l)
: Bne aA, aB, l (loop test).
void call(Reg d, Reg a, Reg b, Reg c, const void * fn)
: Windowed call8 to a host built-in.
void callLabel(Label l, Reg d = R0, bool take = false)
: Call a function in this block by label: the script-to-script call.
void epilogue()
: Retw.n.
void retValue(Reg a)
: Park a where the ABI returns a value, before the epilogue tears the frame down.
Public Static Attributes¶
constexpr uint8_t kMaxSpillSlots = kTotalSlots
: The allocator's slot range plus the parked host arguments.
Public Types¶
using RegType = Reg
: The register type the shared lowering works in, each backend's Reg being its own enum.
More info¶
The window is the return path¶
Ten registers map to a2 upward, and the top of the window is not a general register. A routine opened with entry returns through retw.n, which reads the caller's linkage from the window's top. Using those as registers corrupted that linkage and returned to a garbage address the moment a scripted layout ran.
Entry alignment, twice over¶
The entry instruction must sit on a four-byte boundary, and a call encodes its target in four-byte units, so an unaligned callee cannot be expressed at all. Instructions here are two or three bytes, so a function following another lands anywhere and needs the pad.
Why the buffer is heap¶
The assembler is a stack local, so a cap-sized member put 2 KB on the compile chain's stack and overflowed the task on a classic ESP32.