ZEngine¶
Engine Architecture & Memory Management¶
Archive presentation — concepts and historical TLSF integration plan
Point-in-time material: implementation status, profile sizes, and roadmap dates in the remaining slides are not current. Use
docs/design/memory-budget.md,docs/reference/memory-management.md, anddocs/design/completed/tlsf-allocator-integration.mdfor the maintained contracts.TLSFSlabis implemented and spinlock-protectsAlloc,Realloc, andFree.POSIX correction: slides below describe a proposed
PROT_NONE/mprotectreserve-and-commit model. That model is now implemented for independently reserved configured owners; the 8 GiB value is a capacity-validation limit rather than one writable mapping. The maintained production contract is indocs/design/future-plan/memory-budget.md.
Agenda¶
Part I — Context 1. Why Custom Memory Management 2. Engine Memory Architecture · Budget
Part II — The Allocators 3. Lifetime Model · Allocation Flow · Decision Framework 4. ArenaAllocator — Design, API, Performance, Pros & Cons 5. PoolAllocator — Design, API, Performance, Pros & Cons
Part III — The Gap & TLSF 6. The Gap — What Neither Allocator Handles · Fragmentation 7. TLSF — Guarantees, Performance, Internals, Engine Wrapper
Part IV — Integration Plan 8. Upload Pipeline Before/After · Sequence · TextureDeferral · Sizing
Part V — Future Scope & Roadmap 9. AssetManager · ECS Storage · Roadmap
Part VI — Platform Implications 10. Windows · macOS Apple Silicon · Linux x86-64
Why Custom Memory Management?¶
The goal is not to be clever. It is to know exactly what lifetime each allocation has, and pick the cheapest allocator that matches it.
Engine Memory Architecture¶
Current implementation note: this presentation's historical root-arena diagrams and provisional capacity figures predate #835. Configured application runs now reserve named owners independently; the 8 GiB value is a validation limit, not one mapped root arena. See
memory-budget.mdfor the current profile and reservation contract.
graph TD
OS["OS — mmap / VirtualAlloc"]
ROOT["Root Arena 8 GB virtual reservation"]
OS --> ROOT
ROOT --> VD["VulkanDevice 1 GB\nVMA · command pools · descriptor pools"]
ROOT --> IP["ImportPipeline 1 GB\nGltf · Assimp · EnvMap importers"]
ROOT --> AM["AssetManager 512 MB\nMeshes · Materials · UUID maps"]
ROOT --> ECS["ECSScene 512 MB\nComponentStorage · EntityRegistry"]
ROOT --> SZ["Serializer 256 MB\nScene file scratch"]
ROOT --> AN["AnimationMgr 256 MB\nAnimation clips · blend trees · state machines"]
ROOT --> UI["UIContext 128 MB\nImGui · editor components"]
ROOT --> VFS["VirtualFS 64 MB\nMount table · scanner · filewatcher"]
ROOT --> SC["ShaderCache 64 MB\nSPIR-V · reflection data"]
ROOT --> HR["Headroom ~4.2 GB\nStreaming · Physics · Navigation"]
style ROOT fill:#1a1a2e,color:#fff
style OS fill:#0f3460,color:#fff
style HR fill:#555,color:#ccc
Total committed: ~3.8 GB. Virtual: 8 GB. Physical RSS much lower — pages committed via mprotect on demand, then physically backed by OS on first write.
Memory Budget — 8 GB Root Arena¶
pie title Committed Budget (~3.8 GB of 8 GB virtual)
"VulkanDevice" : 1024
"ImportPipeline" : 1024
"AssetManager" : 512
"ECSScene" : 512
"Serializer" : 256
"AnimationMgr" : 256
"UIContext" : 128
"VirtualFS" : 64
"ShaderCache" : 64
"Other (logging, input, swapchain)" : 20
Memory Budget — Key Numbers¶
| Subsystem | Budget | What Lives There |
|---|---|---|
| VulkanDevice | 1 GB | VMA, descriptor pools, command buffers, swapchain |
| ImportPipeline | 1 GB | GltfImporter (64 MB) + AssimpImporter (128 MB) × 2 instances |
| AssetManager | 512 MB | Meshes[], Materials[], Textures[], UUID hash maps (5000-entry cap each) |
| ECSScene | 512 MB | ComponentStorage dense arrays, EntityRegistry, ActorManager |
| Serializer | 256 MB | EditorSceneSerializer scratch (150 MB sub-arena) |
| AnimationMgr | 256 MB | Animation clips, blend tree nodes, state machine transition data |
| UIContext (editor) | 128 MB | ImguiLayer (64 MB) → DockspaceUI (32 MB) → AssetImporter (8 MB) |
| ShaderCache | 64 MB | SPIR-V, reflection data |
| VirtualFS | 64 MB | Mount table, scanner cache, file watcher events |
Headroom reserved for future systems:
| Planned System | Budget | Notes |
|---|---|---|
| StreamingManager | 2 GB | Open-world chunk streaming, cleared per region transition |
| PhysicsEngine | 512 MB | Rigid bodies, broad-phase |
| NavigationEngine | 256 MB | NavMesh, pathfinding, agent state |
Part II¶
The Allocators¶
Lifetime Model · Arena · Pool · Decision Framework
Memory Model — Lifetime Hierarchy¶
graph LR
subgraph Engine["Engine Lifetime — shutdown only"]
direction TB
VD2["VulkanDevice Arena\nShader cache, VMA, descriptor pools"]
ECS2["ECSScene Arena\nArchetype tables, entity registry"]
AM2["AssetManager Arena\nMesh / material / texture arrays"]
end
subgraph Scene["Scene Lifetime — scene load/unload"]
direction TB
ES["EditorScene::LocalArena 200 MB\nInstance arrays, scene graph, strings"]
end
subgraph Task["Per-Task — cleared after task"]
direction TB
IP["ImportPipeline 1 GB\nGltf + Assimp decode scratch\nCleared after each import session"]
end
subgraph Frame["Per-Frame — reset end of frame"]
direction TB
AT["ArenaTemp (ZGetScratch)\nBarrier batch, draw list\nCamera UBO staging"]
end
subgraph Object["Per-Object — individual free"]
direction TB
PA["PoolAllocator\nEntity slots, command buffer handles\nMesh instance slots"]
end
Engine --> Scene --> Task --> Frame
Engine --> Object
style Engine fill:#1a1a2e,color:#fff
style Scene fill:#16213e,color:#fff
style Task fill:#0f3460,color:#fff
style Frame fill:#2ecc71,color:#000
style Object fill:#3498db,color:#fff
Before writing new or std::vector, identify the lifetime. Pick the cheapest allocator that matches it. If nothing fits, the lifetime is unclear — clarify it first.
Memory Model — Allocation Flow¶
graph TD
ROOT["Root Arena 8 GB"]
ROOT --> VD["VulkanDevice 1 GB\n(ArenaAllocator)"]
ROOT --> AM["AssetManager 512 MB\n(ArenaAllocator)"]
ROOT --> ECS["ECSScene 512 MB\n(ArenaAllocator)"]
VD --> RP["AppRenderPipeline\n30 MB sub-arena"]
VD --> TLS["TLSFSlab × N workers\n64–128 MB each\n(Phase 1 target — not yet in codebase)"]
RP --> SC["ZGetScratch\nArenaTemp — reset each frame"]
AM --> PM["Array of AssetMesh · Array of AssetMaterial\nbacked by ArenaAllocator (grows → dead blocks)"]
AM --> PH["UnorderedHashMap of uuid to Handle\nbacked by ArenaAllocator (grows → dead blocks)"]
ECS --> PE["PoolAllocator\nEntitySlots fixed-cap"]
ECS --> PC["PoolAllocator\nComponentSlots fixed-cap"]
style ROOT fill:#1a1a2e,color:#fff
style SC fill:#2ecc71,color:#000
style TLS fill:#e94560,color:#fff
No allocation escapes its designated arena. The root cursor is monotonic — the only reclamation is arena reset or pool destruction.
Allocation Decision Tree¶
flowchart TD
START(["New allocation needed"])
Q1{"Group lifetime?\ne.g. reset each frame\nor after import"}
Q1b{"Stable after init?\nno grows once setup is done"}
Q2{"Fixed size?\nsame N bytes\nevery time"}
Q3{"Variable size +\nindividual lifetime?"}
WARN(["Re-examine the lifetime.\nDo NOT use std::vector / new."])
A1["ArenaAllocator or ArenaTemp\nPick arena whose lifetime matches.\nAlloc: ~3 cycles. No free."]
A1g["TLSFSlab ← Part III\nGroup lifetime but grows unpredictably.\nPre-size OR use slab for realloc.\nAlloc/free: ~30 cycles. O(1)."]
A2["PoolAllocator\nCarve from parent arena at init.\nAlloc: ~5 cyc + memset(chunk). Free: ~5 cyc. O(1)."]
A3["TLSFSlab ← Part III\nOne slab carved from arena at init.\nAlloc/free: ~30 cycles. O(1). Variable size."]
START --> Q1
Q1 -- YES --> Q1b
Q1b -- YES --> A1
Q1b -- NO, grows --> A1g
Q1 -- NO --> Q2
Q2 -- YES --> A2
Q2 -- NO --> Q3
Q3 -- YES --> A3
Q3 -- NO --> WARN
style A1 fill:#2ecc71,color:#000
style A1g fill:#e94560,color:#fff
style A2 fill:#3498db,color:#fff
style A3 fill:#e94560,color:#fff
style WARN fill:#e74c3c,color:#fff
ArenaAllocator — Design¶
The fundamental primitive. Everything else is built on top of it.
mmap(NULL, ZGiga(8), PROT_NONE, MAP_PRIVATE|MAP_ANONYMOUS, ...)
│
└── PROT_NONE: address space reserved, NO access yet
write to PROT_NONE → SIGSEGV (not a silent page commit)
On each Allocate() call, if cursor crosses a page boundary:
mprotect(base + committed, new_pages, PROT_READ|PROT_WRITE)
│
└── pages now accessible; physical RAM committed on first write
→ RSS << virtual reservation
m_memory (base pointer)
│
▼
┌──────────────────────────────────────────┐
│ alloc A │ alloc B │ alloc C │ ... │ committed
└───────────┴───────────┴───────────┴─────┤
▲ │
m_current_offset │
├───────────────►│ m_size
uncommitted
// precondition (asserted inside memory_align()):
// align > 0 && (align & (align-1)) == 0 — must be power of two
aligned = (current + align - 1) & ~(align - 1)
next = aligned + size
// guard: size > m_total_size - aligned (avoids wrap)
m_current_offset = next
return m_memory + aligned
// postcondition: returned memory is zeroed (secure_memset)
ArenaAllocator — API Surface¶
struct ArenaAllocator {
void Initialize(size_t size); // mmap / VirtualAlloc reservation
void Shutdown(); // munmap / VirtualFree
void* Allocate(size_t size, // bump forward, align, return pointer
size_t alignment = DEFAULT_ALIGNMENT);
void* Resize(void* ptr, size_t old_size, // extend in-place if last alloc,
size_t new_size, size_t alignment); // else new alloc + copy
void Clear(); // reset cursor to initial offset
void CreateSubArena(size_t size, ArenaAllocator* out);
};
// Scratch (temp) arena — LIFO pairing
ArenaTemp scratch = ZGetScratch(arena); // saves current cursor
// ... transient allocations ...
ZReleaseScratch(scratch); // restores cursor to saved position
Key invariant¶
Resize on the last allocation in the arena extends it in-place (cursor rewind + rebump). On any other allocation it allocates a new block forward and copies — the old block becomes permanently dead weight.
This is safe and correct for scratch arenas. It is a slow memory leak for long-lived growing containers.
Detection caveat: "last allocation" is detected via
m_memory + m_previous_offset == old_ptr(usingm_previous_offset, not a computed size). The caller does not need to pass the aligned size — the implementation uses the stored previous-offset directly.Resize copy semantics: copies
min(old_size, new_size)bytes. Shrink zeroes the freed tail. Old block abandoned in place on slow-path (not zeroed). OOM on slow-path asserts viaZENGINE_VALIDATE_ASSERT(always-on crash handler) — never returns null.
Scratch scope constraint¶
// CORRECT — flat scratch, or nested with LIFO release
ArenaTemp outer = ZGetScratch(arena); // saves cursor A
// ... outer allocations (cursor → B) ...
ArenaTemp inner = ZGetScratch(arena); // saves cursor B
// ... inner allocations (cursor → C) ...
ZReleaseScratch(inner); // cursor → B ✓
ZReleaseScratch(outer); // cursor → A ✓
// INCORRECT — releasing outer before inner
ArenaTemp outer = ZGetScratch(arena); // saves cursor A
ArenaTemp inner = ZGetScratch(arena); // saves cursor B > A
ZReleaseScratch(outer); // cursor → A — inner's save point (B) now stale!
// arena cursor is at A; inner.CurrentOffset = B > A
// any new allocation overwrites B..C that inner "thinks" it owns
ZReleaseScratch(inner); // cursor → B — unwinds forward past A — corrupt
ZGetScratch / ZReleaseScratch must be released in strict LIFO order. Releasing an outer scratch while an inner scratch is still live leaves the inner's save point above the current cursor — its subsequent release unwinds in the wrong direction. The common footgun is passing the same arena to two independent functions that each call ZGetScratch without knowing the other holds a live scratch.
Thread safety¶
ArenaAllocator is not thread-safe. All arenas are carved on the main thread at engine startup before any worker starts. Workers receive pre-carved TLSFSlab handles — they never call ArenaAllocator::Allocate after init. CreateSubArena must not be called concurrently.
ArenaAllocator — Performance¶
Allocate(n, align):
aligned = memory_align(cur, align) // 2 ops + is_power_of_two assert
if n > m_total_size - aligned: return null // overflow-safe OOM check
if aligned+n > committed: mprotect(...) // OS call only on page boundary cross
cur = aligned + n // 1 store
memset(base+aligned, 0, n) // O(n) zeroing — dominates for large n
return base + aligned // 1 add
ArenaAllocator — Pros¶
ArenaAllocator — Cons & Constraints¶
### Grow leaks dead blocks
Array<T> initial:
┌────────────────────────┐ free
│ 8 entries (64 B) │──►
└────────────────────────┘
▲ cursor
Array<T> after push (grows to 16):
┌────────────────────────┬─────────────────────────────┐ free
│ 8 entries (DEAD) │ 16 entries (live) │──►
└────────────────────────┴─────────────────────────────┘
▲ cursor
▲ 64 bytes permanently lost
PoolAllocator — Design¶
Fixed-size free list. Carved from an ArenaAllocator at initialization time.
Parent arena gives PoolAllocator a single block:
┌──────────────────────────────────────────┐
│ chunk 0 │ chunk 1 │ chunk 2 │ ... │
└──┬────────┴──┬────────┴──┬────────┴──────┘
│ │ │
│ free list (links stored INSIDE chunks):
│
head ──► [chunk 2] ──► [chunk 0] ──► nullptr
▲ PoolFreeNode* stored in first bytes of each free chunk
void PoolAllocator::Initialize(
ArenaAllocator* arena,
size_t total_size,
size_t chunk_size,
size_t alignment)
{
// One arena carve for all chunks at once
memory = static_cast<uint8_t*>(arena->Allocate(total_size, alignment));
total_size = size;
chunk_size = chk_size; // already aligned to power-of-two boundary
head = nullptr;
Clear(); // zeroes all chunks and builds the free list
}
// Clear() — O(capacity): zero every chunk, then rebuild free list
void PoolAllocator::Clear()
{
auto count = total_size / chunk_size;
for (size_t i = 0; i < count; i++) {
void* ptr = &memory[i * chunk_size];
secure_memset(ptr, 0, chunk_size);
// C-style cast: technically UB (no placement new), but PoolFreeNode is
// trivially copyable — all compilers generate correct code in practice.
// Conformant fix: ::new (ptr) PoolFreeNode{head}
auto* node = (PoolFreeNode*) ptr;
node->Next = head;
head = node;
}
// Result: head → chunk[count-1] → chunk[count-2] → ... → chunk[0] → nullptr
}
PoolAllocator — API Surface¶
struct PoolAllocator {
using Arena = ArenaAllocator;
void Initialize(Arena* arena, size_t total_size,
size_t chunk_size, size_t alignment = DEFAULT_ALIGNMENT);
void* Allocate(); // O(1) — pop free-list head, zero chunk, return; nullptr if exhausted
void Free(void* ptr); // O(1) — push free-list head (bounds-checked, always-on assert)
void Clear(); // O(capacity) — zero all chunks, rebuild free list
};
Usage in engine¶
// Entity slot pool — 4096 entities, each 128 bytes
PoolAllocator entity_pool;
entity_pool.Initialize(&ecs_arena, 4096 * 128, 128);
// Allocate one entity slot
void* slot = entity_pool.Allocate(); // pops free list + zeroes chunk — ~5 cyc + memset(128)
// Return the slot
entity_pool.Free(slot); // pushes free list, asserts ownership — ~8 cycles
Safety invariants enforced¶
Freeassertsptr >= m_memory && ptr < m_memory + total_sizeFreeasserts(ptr - m_memory) % chunk_size == 0(chunk-aligned)- Use-after-free: subsequent
Allocate()returns the freed chunk — the pool zeroes it automatically (secure_memsetinsideAllocate) before returning. Caller receives a pre-zeroed slot; no manual clear needed. - Double-free: second
Freeof same ptr silently corrupts the free list in all builds — the range and alignment assertions do NOT catch it (the pointer is still valid by both checks). Detection requires a per-chunk allocated/free bitset, which this allocator does not maintain. Guard against it at the call site.
PoolAllocator — Performance¶
PoolAllocator backing for 4096 entities × 128 B:
Total reserved: 524 288 bytes (512 KB)
Metadata: 1 pointer in each free chunk
= 0 bytes overhead on allocated chunks
Header cost: sizeof(PoolAllocator) = ~48 bytes
vs jemalloc for same workload:
Each 128-byte alloc gets a ~32-byte header → 25% overhead
Total overhead: ~131 KB just for metadata
PoolAllocator — Pros & Cons¶
Part III¶
The Gap & TLSF¶
Where Arena and Pool fall short · TLSF solution
The Gap — What Neither Allocator Handles¶
stbi_load / DeserializeEnvMap
│
▼
std::vector<float> output_buf ← malloc(w*h*4*4)
│
Bitmap vertical_cross ← std::vector inside → malloc
Bitmap cubemap ← std::vector inside → malloc
│
std::vector<uint8_t> buffer ← malloc(final_bytes)
│
TextureDeferral { Buffer = std::vector }
│
ThreadSafeQueue
│
CompleteDeferrals()
│
~TextureDeferral() ← free()
Fragmentation: Arena vs Pool vs TLSF¶
Texture upload session — 10 textures, mixed sizes:
[16MB] [256KB] [8MB] [16MB] [256KB] [256KB] [8MB] [16MB] [256KB] [8MB]
ArenaAllocator (if it had free — it does not, showing the leak pattern):
┌──────┬───┬──────┬──────┬───┬───┬──────┬──────┬───┬──────┐
│ 16MB │256│ 8MB │ 16MB │256│256│ 8MB │ 16MB │256│ 8MB │ total: ~73MB
└──────┴───┴──────┴──────┴───┴───┴──────┴──────┴───┴──────┘
All freed? Cursor reset = all 73MB back. But no INDIVIDUAL free.
PoolAllocator (chunk = 16MB, 8 slots = 128MB reserved):
┌──────┬──────┬──────┬──────┬──────┬──────┬──────┬──────┐
│ 16MB │ 16MB │ 16MB │ 16MB │ 16MB │ 16MB │ 16MB │ 16MB │ 128MB reserved
│ used │wasted│ used │ used │wasted│wasted│ used │ used │ 64MB wasted
└──────┴──────┴──────┴──────┴──────┴──────┴──────┴──────┘
TLSF (one 64MB slab):
Alloc 16MB: [16MB used] [48MB free]
Alloc 256KB: [16MB] [256KB] [~47.7MB free]
Free 16MB: [16MB FREE] [256KB] [~47.7MB free]
Coalesce: [16MB FREE] [256KB] [~47.7MB free] ← left merge skipped (256KB between)
Alloc 8MB: [8MB used] [8MB FREE] [256KB] [~47.7MB free] ← fits in 16MB freed slot
After all 10 textures freed: single contiguous 64MB free block.
Fragmentation: zero. All adjacent frees coalesced.
TLSF — What Is It?¶
Two-Level Segregated Fit (mattconte/tlsf) is a dynamic allocator designed for systems where alloc/free latency must be deterministic.
// You provide the memory
uint8_t backing[64 * 1024 * 1024]; // 64 MB slab
// TLSF manages it
tlsf_t pool = tlsf_create_with_pool(backing, sizeof(backing));
// Allocate from it
void* a = tlsf_malloc(pool, 16 * 1024 * 1024); // 16 MB
void* b = tlsf_malloc(pool, 256 * 1024); // 256 KB
// Free returns to the slab — coalesces with neighbors
tlsf_free(pool, a); // 16 MB block merges back into free space
// The slab is immediately reusable for the next allocation
void* c = tlsf_malloc(pool, 14 * 1024 * 1024); // 14 MB — succeeds
TLSF — Performance Profile¶
TLSF — Internal Design: Two-Level Bitmap¶
TLSF organizes free blocks into a two-tier segregated free list indexed by a bitmap. Finding the right free list is a single bit-scan instruction.
First-Level Index (FL): floor(log2(block_size_in_bytes))
Second-Level Index (SL): subdivides each FL band into 16–32 sub-classes
(For MB-range blocks the actual FL values are 22–26; the diagram uses
abbreviated band indices for readability — the structure is identical.)
FL bitmap (one bit per size class)
┌──┬──┬──┬──┬──┬──┬──┬──┐
│0 │0 │1 │0 │1 │1 │0 │1 │ 1 = at least one free block in this class
└──┴──┴──┴──┴──┴──┴──┴──┘
│ │ │ │
│ │ └─── band 7 → SL bitmap for 64MB–128MB range
│ └────── band 5 → SL bitmap for 16MB–32MB range
└─────────────── band 3 → SL bitmap for 4MB–8MB range
Second-Level bitmap for band 5 (16 MB – 32 MB):
┌──┬──┬──┬──┬──┬──┬──┬──┬──┬──┬──┬──┬──┬──┬──┬──┐
│0 │0 │1 │0 │0 │0 │1 │0 │0 │0 │0 │1 │0 │0 │0 │0 │
└──┴──┴──┴──┴──┴──┴──┴──┴──┴──┴──┴──┴──┴──┴──┴──┘
│ │ │
17.5MB 19MB 22MB ← free block size classes
Free list head for each SL:
[17.5MB] → block → block → nullptr
[19MB ] → block → nullptr
[22MB ] → block → block → block → nullptr
tlsf_malloc(n):
1. Compute FL = floor(log2(n)), SL = sub-class of n within FL band — 2 ops
2. Find next set bit in the bitmap at or above (FL, SL) — 1 BSF instruction
3. Pop the head of that free list — 1 load
4. Split remainder back into a smaller free block — 2 list ops
Total: ~8–12 operations — O(1) by construction.
TLSF — Block Structure¶
Each block in the TLSF heap has a small header. Headers form a doubly-linked physical list for coalescing.
Memory layout of a TLSF slab:
┌───────────────────────────────────────────────────────────────────┐
│ TLSF slab (e.g. 64 MB) │
├──────────┬──────────────────────┬──────────────┬─────────────────┤
│ TLSF hdr │ Allocated Block │ Free Block │ Allocated Block │
│ ~3 KB │ │ │ │
└──────────┴─────────────┬────────┴──────────────┴─────────────────┘
│
Block header (16–32 bytes):
┌────────────────────────────────────┐
│ prev_phys_block* │ size | flags │ ← all blocks (16 B)
│ (free only: │ │
│ next_free*, │ F = free bit │ ← free only (+16 B)
│ prev_free*) │ P = prev free │
└────────────────────────────────────┘
│ │
│◄── header overhead: 16 B alloc ►│
│ 32 B free │
Coalesce on free:
┌───────────┬──────────┬───────────┐ ┌──────────────────────┐
│ free 8MB │ free 4MB │ free 6MB │ ──► │ merged free 18MB │
└───────────┴──────────┴───────────┘ └──────────────────────┘
All three blocks are adjacent in physical memory →
tlsf_free() merges them in O(1) using prev_phys_block pointer.
Coalescing is what prevents the "fragmentation cliff" seen with multiple pools: TLSF can always recombine adjacent free space regardless of what size class it was in.
TLSFSlab — The Engine Wrapper¶
// ZEngine/ZEngine/Core/Memory/TLSFSlab.h
struct TLSFSlab {
void* Backing = nullptr;
tlsf_t Pool = nullptr;
void Init(ArenaAllocator* arena, size_t bytes);
void* Alloc(size_t n);
void* Realloc(void* ptr, size_t n);
void Free(void* ptr);
void Shutdown();
size_t Overhead() const; // diagnostic: TLSF metadata bytes
};
Device->Arena (ArenaAllocator, 1 GB)
│
└── ZPushSize(&Device->Arena, 64MB, 64)
│
▼
┌──────────────────────────────────┐
│ TLSFSlab::Backing (64 MB) │
│ │
│ [tlsf metadata ~3KB] │
│ [free block: 64MB - 3KB] │
│ │
│ tlsf_malloc ──► alloc from here │
│ tlsf_free ──► merge here │
└──────────────────────────────────┘
Arena cursor moves once at Init.
Arena is NOT touched by subsequent Alloc/Free calls.
void TLSFSlab::Init(ArenaAllocator* arena, size_t bytes)
{
Backing = ZPushSize(arena, bytes, 64);
ZENGINE_VALIDATE_ASSERT(Backing, "TLSFSlab: arena OOM");
Pool = tlsf_create_with_pool(Backing, bytes);
ZENGINE_VALIDATE_ASSERT(Pool, "TLSFSlab: tlsf_create failed");
}
Thread Safety — Per-Worker Slabs¶
graph TD
SUB["Submitter Thread\nThreadPoolHelper::Submit(lambda)"]
SUB -- "round-robin\ncursor % WorkerCount" --> W0
SUB -- "round-robin" --> W1
SUB -- "round-robin" --> W2
SUB -- "round-robin" --> W3
subgraph TP["ThreadPool (hardware_concurrency - 1, max 16)"]
W0["Worker 0\nWorkerRun(0)\nt_worker_slab = &slab[0]"]
W1["Worker 1\nWorkerRun(1)\nt_worker_slab = &slab[1]"]
W2["Worker 2\nWorkerRun(2)\nt_worker_slab = &slab[2]"]
W3["Worker 3\nWorkerRun(3)\nt_worker_slab = &slab[3]"]
end
W0 -- "exclusive\nno lock" --> S0["TLSFSlab[0]\n64 MB"]
W1 -- "exclusive\nno lock" --> S1["TLSFSlab[1]\n64 MB"]
W2 -- "exclusive\nno lock" --> S2["TLSFSlab[2]\n64 MB"]
W3 -- "exclusive\nno lock" --> S3["TLSFSlab[3]\n64 MB"]
S0 & S1 & S2 & S3 -- "one arena carve at init" --> VA["VulkanDevice Arena\n1 GB"]
style VA fill:#1a1a2e,color:#fff
style S0 fill:#e94560,color:#fff
style S1 fill:#e94560,color:#fff
style S2 fill:#e94560,color:#fff
style S3 fill:#e94560,color:#fff
Thread-local t_worker_slab set at WorkerRun(idx) start. Lambdas call GetWorkerSlab(). Zero contention between workers — each slab is touched by exactly one worker at a time.
Open problem — cross-thread free:
CompleteDeferrals()runs on the render thread and callsd.Slab->Free(d.Pixels)on a worker's slab. If that worker is concurrently executing the next decode task, both threads touch the same TLSF pool simultaneously — undefined behavior. TLSF has no internal lock. This must be resolved before Phase 1 ships. Options: (a) per-slab spinlock on Free only, (b) worker drains a lock-free deferred-free queue at task start, (c) render thread posts the free back to the worker's queue.
Part IV¶
Integration Plan¶
Phase 1 · Upload Pipeline · TextureDeferral
Upload Pipeline — Before vs After¶
Worker thread
├── std::vector<float> output_buf
│ └── malloc(w × h × 4 × 4) ← system heap
│
├── Bitmap vertical_cross
│ └── std::vector<uint8_t> ← malloc
│
├── Bitmap cubemap
│ └── std::vector<uint8_t> ← malloc
│
├── std::vector<uint8_t> buffer
│ └── malloc(final_bytes) ← malloc
│
└── TextureDeferral {
Buffer = std::vector<uint8_t> ← ownership move into queue
}
Render thread (CompleteDeferrals)
└── ~TextureDeferral()
└── ~vector<uint8_t>() ← free() on system heap
Worker thread (t_worker_slab = &slab[idx])
├── float* raw = slab->Alloc(w×h×4×4) ← TLSF, no lock
│ stbi_load writes into raw ← requires STBI_MALLOC override
│
├── Bitmap vertical_cross (uses slab) ← TLSF
│ └── slab->Free(raw) ← input decode buffer freed
│
├── Bitmap cubemap (uses slab) ← TLSF
│ └── freed immediately
│
├── uint8_t* pixels = slab->Alloc(n) ← TLSF, only live alloc
│ └── memmove(pixels, cubemap, n)
│
└── TextureDeferral {
Pixels = pixels
ByteSize = n
Slab = &slab[idx] ← back-ref for free
}
Render thread (CompleteDeferrals) ← NOTE: cross-thread free (see thread safety)
└── UploadTextureBuffer(pixels)
d.Slab->Free(d.Pixels) ← O(1) TLSF free, block merges
Upload Pipeline — Sequence View¶
sequenceDiagram
participant S as Submitter
participant TP as ThreadPool Worker N
participant SL as TLSFSlab[N]
participant Q as DeferralQueue
participant RT as Render Thread
participant GPU as GPU Upload
S ->> TP : Submit(decode task)
TP ->> SL : Alloc(w×h×4×4) ── decode buffer
Note over TP: stbi_load writes here only if STBI_MALLOC overridden
TP ->> TP : stbi_load / DeserializeEnvMap → raw pixels
TP ->> SL : Alloc(cubemap_bytes) ── final pixel buffer
TP ->> SL : Free(decode buffer) ── intermediates freed immediately
TP ->> Q : Enqueue TextureDeferral { Pixels*, ByteSize, Slab* }
RT ->> Q : Pop TextureDeferral
RT ->> GPU : UploadTextureBuffer(Pixels)
Note over RT,SL: Cross-thread free — requires lock or deferred-free queue (see Thread Safety slide)
RT ->> SL : Free(Pixels) ── O(1), block merges back
The slab block is live only from decode completion to GPU upload confirmation — the narrowest possible window. Every free immediately coalesces with adjacent free space, keeping the slab in a clean, fully-mergeable state for the next texture.
TextureDeferral — Struct Change¶
struct TextureDeferral {
// std::variant<> occupies max(sizeof(T)) bytes
// + discriminant + alignment padding
// = ~40 bytes for the variant alone
std::variant<
unsigned char*, // borrowed (IsLarge = false)
std::vector<uint8_t> // owned (IsLarge = true)
> Buffer;
Textures::TextureHandle TexHandle = {};
uint8_t FrameIdx = 0;
uint8_t ThreadIdx = 0;
bool IsLarge = false;
};
struct TextureDeferral {
unsigned char* Pixels = nullptr;
size_t ByteSize = 0;
Textures::TextureHandle TexHandle = {};
uint8_t FrameIdx = 0;
uint8_t ThreadIdx = 0;
bool IsLarge = false;
// null when IsLarge=false (borrowed pointer, slab not involved)
Core::Memory::TLSFSlab* Slab = nullptr;
};
// Before:
auto& buf = std::get<std::vector<uint8_t>>(d.Buffer);
UploadTextureBuffer(d.FrameIdx, d.ThreadIdx, d.TexHandle,
buf.data());
// implicit: vector destructs → free()
// After:
UploadTextureBuffer(d.FrameIdx, d.ThreadIdx, d.TexHandle,
d.Pixels);
if (d.IsLarge && d.Slab)
d.Slab->Free(d.Pixels); // O(1), merges back into slab
Sizing and Memory Impact¶
Part V¶
Future Scope & Roadmap¶
AssetManager · ECS · Timeline
Future Scope — AssetManager¶
The AssetManager arena backs long-lived Array<T> and UnorderedHashMap<K,V> containers. Every grow() wastes a dead block.
AssetManager::Arena (512 MB)
│
├── Meshes[] Array<AssetMesh> — grows with every import
├── NodeHierarchies[] Array<AssetNode> — grows with every import
├── Materials[] Array<AssetMaterial> — grows with every import
├── UUIDToTextureHandle UnorderedHashMap<uuid,TextureHandle>
└── UUIDToMaterialSlot UnorderedHashMap<uuid,uint32_t>
After a session importing 500 assets with 3 resizes each:
Dead blocks per container = C₀ + 2C₀ + 4C₀ (3 doublings from initial capacity C₀)
= 7 × C₀
AssetMesh: C₀ ≈ 64 entries × ~128 B = 8 KB → 7 × 8 KB = ~56 KB dead
AssetNode: C₀ ≈ 64 entries × ~64 B = 4 KB → 7 × 4 KB = ~28 KB dead
AssetMaterial:C₀ ≈ 64 entries × ~128 B = 8 KB → 7 × 8 KB = ~56 KB dead
UUIDToTex: C₀ ≈ 256 buckets × ~24 B = 6 KB → 7 × 6 KB = ~42 KB dead
UUIDToMat: C₀ ≈ 256 buckets × ~20 B = 5 KB → 7 × 5 KB = ~35 KB dead
Total per session: ~217 KB of dead arena blocks
This is a modest but permanent leak — across 100 import sessions in a long editor run it accumulates to ~20 MB. It does not disappear until the AssetManager arena is torn down at shutdown.
A TLSFSlab backing these containers would give realloc() behavior: if the block immediately following the container's current allocation is free, TLSF extends in-place — zero copy. Dead block accumulation drops to near zero.
Requires: Array<T> and UnorderedHashMap<K,V> to accept a typed allocator parameter. This is a separate refactor — the container interface currently only accepts ArenaAllocator*.
Future Scope — ECS Component Storage¶
ComponentStorage today (planned — issue #642):
ArchetypeTable[Position] PoolAllocator — 3× float, fixed size, fine
ArchetypeTable[Velocity] PoolAllocator — 3× float, fixed size, fine
ArchetypeTable[MeshRenderer] PoolAllocator — uuid + handle, fixed size, fine
ArchetypeTable[Physics] ??? — variable: shape data, constraints
ArchetypeTable[Script] ??? — variable: script bytecode + state
When component types carry variable-size payloads (script state, animation clips, physics shapes), a PoolAllocator per archetype requires knowing the maximum instance size upfront. A TLSFSlab per archetype table handles heterogeneous payloads cleanly:
ArchetypeTable[Physics]
├── backing: TLSFSlab (16 MB, carved from ECSScene arena)
├── entity 1: RigidBody { mass, inertia, 6 constraint params } = 72 bytes
├── entity 2: Cloth { mesh_ref, wind_params[32], pin[16] } = 320 bytes
└── entity 3: Softbody { particle[128] } = 512 bytes
No pre-declared fixed size. TLSF handles the mix in O(1).
Roadmap¶
gantt
title TLSF Integration Roadmap
dateFormat YYYY-MM-DD
axisFormat %b %Y
section Phase 1 — Upload Pipeline (v0.5.0)
TLSFSlab wrapper (h + cpp) :p1a, 2026-09-01, 3d
ThreadPool TLS slab pointer :p1b, after p1a, 2d
Cross-thread free resolution :p1c, after p1b, 3d
RRM slab init + sizing :p1d, after p1c, 2d
TextureDeferral struct change :p1e, after p1d, 2d
Replace upload std::vector :p1f, after p1e, 3d
STBI_MALLOC override :p1g, after p1f, 2d
Profile + validate :p1h, after p1g, 3d
section Phase 2 — AssetManager (v0.6.0)
Array/HashMap typed allocator :p2a, 2026-10-01, 7d
AssetManagerSlab creation :p2b, after p2a, 3d
Container migration :p2c, after p2b, 5d
Dead-block waste measurement :p2d, after p2c, 2d
section Phase 3 — ECS Storage (v1.0.0)
Per-archetype TLSFSlab :p3a, 2026-12-01, 5d
Variable-payload component types :p3b, after p3a, 7d
Roadmap — Implementation Phases¶
Part VI¶
Platform Implications¶
Windows · macOS Apple Silicon · Linux x86-64
Platform Implications — Windows / macOS arm64 / Linux¶
| Property | Windows | macOS arm64 | Linux x86-64 |
|---|---|---|---|
| Reserve | VirtualAlloc(MEM_RESERVE, PAGE_NOACCESS) |
mmap(PROT_NONE) |
mmap(PROT_NONE) |
| Commit | VirtualAlloc(MEM_COMMIT, PAGE_READWRITE) |
mprotect(PROT_READ\|PROT_WRITE) |
mprotect(PROT_READ\|PROT_WRITE) |
| Release | VirtualFree(MEM_RELEASE) |
munmap |
munmap |
| Physical page size | 4 KB | 16 KB | 4 KB (default) |
| Commit granularity | 4 KB | 16 KB | 4 KB |
| GPU memory model | Discrete — PCIe staging buffer | UMA — CPU/GPU share pool | Discrete — PCIe staging buffer |
| Huge pages | MEM_LARGE_PAGES (2 MB, privileged) |
Not exposed | madvise(MADV_HUGEPAGE) or MAP_HUGETLB |
| Virtual address space | 128 TB (user mode) | 256 TB (arm64) | 128 TB (x86-64) |
| Overcommit | Pagefile-backed; no OOM killer | Not available | Configurable; OOM killer if exhausted |
Platform Implications — Engine Impact¶
Summary¶
| Allocator | Cost | Free? | Size | Best For |
|---|---|---|---|---|
| ArenaAllocator | ~3 cyc + memset(n) | No | Any | Scratch, import, per-frame |
| PoolAllocator | ~5 cyc + memset(chunk) | Yes | Fixed | Entity slots, handles, same-type objects |
| TLSFSlab | ~30 cycles | Yes | Variable | Upload buffers, asset metadata, ECS payloads |
| System heap | ~100–500 cycles | Yes | Variable | Nothing on the hot path |
Arena and Pool cover 95% of engine allocations today. TLSF fills the remaining 5% — variable size, individual lifetimes — currently leaking through to the system heap. The upload pipeline is Phase 1: measurable, isolated, zero risk to other systems.