Memory Management¶
Status: Current allocator primitives and GPU-domain reference, with historical planning
notes called out below. The authoritative configured CPU profile is
design/memory-budget.md.
ZEngine primarily uses explicit arena, pool, and TLSF-slab ownership. This page documents the live primitives, allocation patterns, GPU memory domains, and teardown rules for Vulkan handles. It does not claim that every allocation in the repository avoids the system heap: third-party libraries and a few non-hot-path facilities still use their own allocation policies.
See also: Engine Architecture · Asset Manager
Table of Contents¶
- Philosophy
- Lifetime Model
- Allocation Decision Framework
- CPU Memory — ArenaAllocator
- CPU Memory — PoolAllocator
- CPU Memory — TLSFSlab
- Allocation Macros
- Scratch Arenas
- Thread Safety
- Memory Budget
- Performance Comparison
- Container Ownership Rules
- GPU Memory — VMA Allocator
- GPU Memory Domains
- Arena-Allocated Vulkan Objects
- Platform Notes
- Memory Profiler
Philosophy¶
- Named up-front reservations. A configured
MemoryManagerindependently reserves each named CPU owner. The 8 GiB value passed byObeliskis a profile-capacity validation limit, not oneMainArenamapping. Individual objects never callmalloc/newoutside of third-party libraries. - Owners and child arenas have fixed budgets. Each subsystem gets a dedicated owner sized to its worst-case working set; short-lived consumers carve child arenas from that owner. Running out of either is a budgeting error to fix at design time, not a runtime failure to handle.
- Lifetime = scope. Objects allocated from an arena are freed by
ArenaAllocator::Clear()(cursor reset). There is no per-object free. Pick the allocator whose lifetime matches the object's lifetime. - No destructor guarantee.
ZPushStructCtorplaces objects via placement-new, but arena release does not call destructors. Any object that owns an OS or GPU resource must have its destructor called explicitly before the arena is cleared. - Zero hot-path touches. Alloc/free on the render thread or in inner simulation loops is off the table.
Before writing
neworstd::vector, identify the lifetime. Pick the cheapest allocator that matches it. If nothing fits, the lifetime is unclear — clarify it first.
Lifetime Model¶
Each allocation belongs to exactly one lifetime tier:
| Tier | When freed | Allocator | Examples |
|---|---|---|---|
| Engine | Shutdown only | ArenaAllocator |
VulkanDevice, ECSScene, AssetManager arenas |
| Scene | Scene load/unload | ArenaAllocator |
target editor/session data; it is not yet a separately budgeted production editor arena |
| Per-task | Task-specific boundary | ArenaAllocator or TLSFSlab |
importer scratch and variable-lifetime upload/decode data |
| Per-frame | End of frame / GPU completion | ArenaTemp or PerFrameUploadHeap |
CPU scratch; mapped GPU uniform/storage/indirect uploads |
| Per-object | Individual free needed | PoolAllocator |
Entity slots, command buffer handles, mesh instance slots |
| Variable | Individual free, variable size | TLSFSlab |
Texture decode buffers and growable engine containers |
Allocation Decision Framework¶
New allocation needed
│
▼
Group lifetime? (reset all at once after frame / import / scene)
YES ──► Stable after init? (no grows once setup is done)
YES ──► ArenaAllocator (~3 cyc + memset)
NO ──► TLSFSlab (~30 cyc, O(1) — roadmap v0.5.0)
NO
│
▼
Fixed size? (same N bytes every time)
YES ──► PoolAllocator (~5 cyc + memset(chunk), O(1))
NO
│
▼
Variable size + individual lifetime?
YES ──► TLSFSlab (~30 cyc, O(1) — roadmap v0.5.0)
NO ──► Re-examine the lifetime. Do NOT use std::vector / new.
CPU Memory — ArenaAllocator¶
File: ZEngine/ZEngine/Core/Memory/Allocator.h
struct ArenaAllocator
{
void Initialize(size_t size);
void* Allocate(size_t size, size_t alignment = DEFAULT_ALIGNMENT);
void* Resize(void* ptr, size_t old_size, size_t new_size, size_t alignment);
void CreateSubArena(size_t size, ArenaAllocator* out);
void Clear(); // reset cursor to 0, keep pages
void Shutdown(); // unmap pages
};
Allocate bumps a cursor — O(1), no locks. Virtual address space is reserved up-front; physical pages are committed on first write (RSS much lower than virtual reservation). Every Allocate calls secure_memset(ptr, 0, n) — negligible for small objects, dominant for large buffers (e.g. ~0.4 ms for a 16 MB decode buffer at 40 GB/s).
Resize extends in-place if the pointer is the most recent allocation; otherwise allocates a new block forward and copies, leaving the old block permanently dead. Safe for scratch arenas; a slow memory leak for long-lived growing containers.
CreateSubArena advances the parent cursor by size. The child manages its own cursor independently.
Key invariants¶
ArenaAllocatoris not thread-safe — all arenas are carved on the main thread before any worker starts.- Arena release does not call destructors. Call
ptr->~T()explicitly on objects owning OS/GPU handles. ZReleaseScratchpairs must be released in strict LIFO order — see Scratch Arenas.
CPU Memory — PoolAllocator¶
File: ZEngine/ZEngine/Core/Memory/Allocator.h
Fixed-size free list backed by a single arena carve at init. Suited for objects of a single known size (entity slots, component handles) with individual lifetimes.
struct PoolAllocator
{
void Initialize(ArenaAllocator* arena, size_t total_size,
size_t chunk_size, size_t alignment = DEFAULT_ALIGNMENT);
void* Allocate(); // O(1) — pop free-list head, zero chunk
void Free(void*); // O(1) — push free-list head; asserts range + alignment
void Clear(); // O(capacity) — zero all chunks, rebuild free list
};
Free list links are stored inside free chunks — zero separate metadata. After Clear() or across alloc/free cycles, allocation order is LIFO-scrambled; sequential layout is only guaranteed at init.
Safety invariants¶
| Check | Enforced? |
|---|---|
Free ptr in range |
Always-on assert |
Free ptr chunk-aligned |
Always-on assert |
| Double-free | Not detected — second Free corrupts the free list silently (issue #697) |
| Exhaustion | Allocate returns nullptr — caller must check (issue #681) |
When NOT to use PoolAllocator¶
- Multiple object sizes — requires multiple pools or wasteful over-sizing to largest.
- Capacity unknown at init — no in-place growth; growing requires a new arena carve.
- Per-frame
Clear()— O(capacity) traversal is too expensive.
CPU Memory — TLSFSlab¶
Status: Implemented. TLSFSlab is backed by mattconte/tlsf, protected by an internal
atomic spinlock, and used by the asset manager, render-resource manager, bitmap helpers, and
TLSF-aware Array/UnorderedHashMap variants.
TLSFSlab wraps mattconte/tlsf (already vendored via FetchContent) with a backing buffer carved from a parent ArenaAllocator. Fills the gap for variable-size, individually-freed allocations that neither Arena nor Pool can handle: texture decode buffers, asset metadata containers that grow unpredictably.
struct TLSFSlab {
void Init(ArenaAllocator* arena, size_t bytes);
void* Alloc(size_t n); // O(1) worst-case — asserts on exhaustion
void* Realloc(void* ptr, size_t n); // O(1) if in-place, O(n) copy otherwise
void Free(void* ptr); // O(1) — coalesces with adjacent free blocks
void Shutdown(); // tlsf_destroy; does NOT free backing
size_t Overhead() const;
};
The backing buffer is carved from the parent arena once at Init. Subsequent Alloc/Free never touch the arena. Internal fragmentation is bounded at ≤ 1.0625× requested size. Adjacent frees always coalesce — no fragmentation cliff over time.
Current texture/upload use¶
Texture/upload work can use worker slabs and the render-resource manager's upload slabs. The
slab's lock makes a cross-thread Free safe; it is not a lock-free per-worker-only allocator.
Individual importers and third-party decoders may still have their own allocations, so this is
not a promise of zero process-heap activity.
Historical rollout record¶
| Phase | Target | Scope |
|---|---|---|
| 1 | v0.5.0 target | Per-worker upload slabs, TextureDeferral refactor, STBI_MALLOC override |
| 2 | v0.6.0 target | AssetManager containers (typed allocator for Array<T> / UnorderedHashMap) |
| 3 | v1.0.0 | Per-archetype ECS slab for variable-payload component types |
Allocation Macros¶
File: ZEngine/ZEngine/ZEngineDef.h
| Macro | Equivalent | Notes |
|---|---|---|
ZKilo(n) |
uint64_t(n) * 1024 |
Always 64-bit — no overflow |
ZMega(n) |
uint64_t(n) * 1024² |
Always 64-bit |
ZGiga(n) |
uint64_t(n) * 1024³ |
Always 64-bit |
ZPushArray(arena, T, count) |
arena->Allocate(count * sizeof(T), alignof(T)) |
Returns T*, no constructor |
ZPushStruct(arena, T) |
ZPushArray(arena, T, 1) |
Returns T*, no constructor |
ZPushStructCtor(arena, T) |
new (ZPushStruct(arena, T)) T() |
Placement-new, default constructor |
ZPushStructCtorArgs(arena, T, ...) |
new (ZPushStruct(arena, T)) T(...) |
Placement-new with args |
When to use each:
- ZPushStruct / ZPushArray — POD structs, trivial types.
- ZPushStructCtor — objects with non-trivial default constructor (CommandPool, Semaphore, …).
- ZPushStructCtorArgs — objects requiring constructor arguments (GameWindow, VulkanDevice, …).
Calling delete on an arena-allocated pointer is undefined behavior. Call ptr->~T() explicitly, then set the pointer to nullptr.
Scratch Arenas¶
Short-lived per-call temporaries use a scratch arena to avoid polluting long-lived arenas.
sequenceDiagram
participant Code as Caller
participant SA as ZGetScratch / ZReleaseScratch
participant TA as Thread-local arena pair [A, B]
Code->>SA: ZGetScratch(&my_arena)
SA->>TA: pick arena that is NOT &my_arena
SA-->>Code: ScratchArena { .Arena = chosen, .checkpoint }
Code->>Code: allocate temporaries from scratch.Arena
Code->>SA: ZReleaseScratch(scratch)
SA->>TA: reset chosen arena cursor to checkpoint
Rules:
- Never store a pointer into a scratch arena past ZReleaseScratch.
- Always pair ZGetScratch / ZReleaseScratch — no early returns between them.
- ZGetScratch / ZReleaseScratch must be released in strict LIFO order. Releasing an outer scratch while an inner scratch is still live leaves the inner's save point stale — a subtle corruption that manifests later.
- Each thread has its own arena pair — scratch arenas are not shared across threads.
Thread Safety¶
| Allocator | Thread-safe? | Notes |
|---|---|---|
ArenaAllocator |
No | All arenas carved on main thread before workers start. Workers never call ArenaAllocator::Allocate after init. |
PoolAllocator |
No | All current pools are single-threaded (main thread or one render thread). CAS / spinlock needed if shared. |
TLSFSlab |
Yes | An internal atomic spinlock protects Alloc, Realloc, and Free, including cross-thread release. |
GpuAllocator (VMA) |
Yes | VMA handles its own synchronization internally. |
Memory Budget¶
MemoryBudgetConfig in ZEngine/ZEngine/Core/Memory/MemoryManager.h defines the profile.
Obelisk validates that profile against an 8 GiB configured-capacity limit before initialization.
The actual configured totals are 7,604 MiB (Default), 7,868 MiB (Editor), and 6,324 MiB
(Server); the largest slots are ImportPipeline (4,096 MiB), AssetManager (1,024 MiB by
default and 1,280 MiB for the calibrated editor profile), VulkanDevice (1,024 MiB), and
ECSScene (512 MiB). UIContext is 64 MiB by default and 128 MiB for the editor. Not every
declared slot is materialized by every runtime mode.
Configured owners are separate virtual reservations: Windows starts them PAGE_NOACCESS, while
macOS/Linux start them PROT_NONE. The allocator promotes only the required page range on each
allocation (VirtualAlloc(MEM_COMMIT) or mprotect), with owner-local page tracking for nested
child arenas. Consequently, configured Linux startup does not need one writable 8 GiB mapping;
the test suite exercises it below a constrained RLIMIT_AS limit. Do not use MAP_NORESERVE
alone as a replacement for explicit commit admission, because it can defer failure to a later
uncontrolled write.
The GPU allocator's domains and the 384 MiB persistent environment-lighting gate are separate
GPU policies, not deductions from this CPU profile. See
design/memory-budget.md for the full live table.
Performance Comparison¶
Approximate cycle counts on a cache-warm allocation path (bookkeeping only — does not include memset(n) zeroing which scales linearly with size):
| Allocator | Alloc cost | Free cost | Fragmentation | Best for |
|---|---|---|---|---|
| ArenaAllocator | ~3–5 cyc + memset(n) |
N/A | Zero | Scratch, import, per-frame |
| PoolAllocator | ~5–8 cyc + memset(chunk) |
~5–8 cyc | Zero | Entity slots, fixed-size objects |
| TLSFSlab | ~20–40 cyc | ~20–40 cyc | ≤ 1.0625× | Upload buffers, growing containers |
| System heap (jemalloc) | ~50–300 cyc | ~50–300 cyc | Accumulates | Nothing on the hot path |
| System heap (ptmalloc) | ~100–500 cyc | ~100–500 cyc | Accumulates | Nothing on the hot path |
The cycle counts are planning estimates, not a current benchmark suite. TLSFSlab is available for variable-size individual lifetimes; it is not a claim that all such allocations have been migrated.
Container Ownership Rules¶
File: ZEngine/ZEngine/Core/Containers/Array.h
Array<T> is move-only¶
Array<T> copy constructor and copy assignment are deleted. The arena owns the backing memory; a shallow copy would alias the same buffer. Moving transfers the pointer and nulls the source.
Array<T>(const Array&) = delete;
Array<T>& operator=(const Array&) = delete;
Array<T>(Array&& other) noexcept;
Array<T>& operator=(Array&& other) noexcept;
Passing conventions¶
void Inspect(const Array<uint32_t>& arr); // read-only
void Mutate(Array<uint32_t>& arr); // in-place mutation
void Consume(Array<uint32_t> arr); // ownership transfer — caller std::move()
ArrayView<T> for non-owning slices¶
ArrayView<T> is a plain {T*, size_t} — freely copyable, no ownership semantics.
Growing containers and allocator choice¶
Every Array<T>::grow() that reallocates on an arena abandons the old block for that arena's
lifetime. Pre-size stable containers. For a genuinely variable, individually released workload,
use the implemented TLSF-aware container initialization rather than treating arena growth as a
general allocator.
HashMap / UnorderedHashMap with move-only values¶
map.insert(key, std::move(my_array)); // rvalue overload for move-only values
for (auto& [k, v] : my_map) { v.push(42); } // reference — no copy
The insert(const K&, const V&) overload is gated with requires std::is_copy_assignable_v<V> — using it with a move-only value is a compile error.
GPU Memory — VMA Allocator¶
File: ZEngine/ZEngine/Core/Memory/GpuAllocator.h
GPU memory is managed by Vulkan Memory Allocator (VMA). GpuAllocator wraps VmaAllocator and exposes typed helpers:
BufferView AllocateBuffer(VkDeviceSize, VkBufferUsageFlags, GpuMemoryDomain, const char* debug_name);
void FreeBuffer(BufferView&);
BufferImage AllocateImage(VkImageCreateInfo&, GpuMemoryDomain, VkDevice,
VkImageAspectFlagBits, VkImageViewType, uint32_t layers, const char*);
void FreeImage(BufferImage&, VkDevice);
BufferView and BufferImage hold raw VkHandles + VmaAllocation. They are not arena-allocated and must be freed explicitly before the device is destroyed.
GPU Memory Domains¶
graph LR
DG["DeviceGeometry\nVMA_MEMORY_USAGE_AUTO\ndevice-local preferred → VRAM\nGlobal vertex/index/storage buffers"]
DT["DeviceTexture\nVMA_MEMORY_USAGE_AUTO\ndevice-local preferred → VRAM\nTexture images"]
HU["HostUniform\nVMA_MEMORY_USAGE_AUTO\nhost-visible required → BAR / shared\nTransformSB, DrawDataSB"]
HS["HostStaging\nPersistent staging ring plus one-shot fallback\nSequential CPU writes"]
HR["HostReadback\nMapped read-mostly allocations\nGPU-to-CPU readback"]
Rule: HostUniform includes the mapped per-frame heaps and is flushed when non-coherent.
DeviceGeometry and DeviceTexture receive GPU copies through the staging ring or a one-shot
fallback. RenderTarget intentionally follows the default allocator path rather than a custom
domain pool.
Arena-Allocated Vulkan Objects¶
Arena Clear() does not call destructors. Objects holding VkCommandPool, VkSemaphore, etc. must have their destructor called explicitly before the device is destroyed.
flowchart TD
A["Arena-allocated object owns VkHandle"]
B["Subsystem Shutdown() / Deinitialize()"]
C{"GPU-idle\nguaranteed?"}
D["Direct: ptr→~T() → vkDestroy*\nat QueueWaitAll point"]
E["Deferred: Device→DeferFree(entry)\ndrained when timeline value ≥ stamp"]
F["ptr = nullptr"]
A --> B --> C
C -->|Yes| D --> F
C -->|No| E --> F
| Class | Strategy | Reason |
|---|---|---|
CommandPool |
Direct | Always freed at GPU-idle |
FramebufferVNext |
Direct | Called after QueueWaitAll |
GraphicPipeline |
Direct | Same |
Semaphore |
Deferred | Can be signalled; deferred prevents in-flight use |
Fence |
Deferred | Same |
DeferredFreeQueue is a 2048-slot circular buffer, drained in Deinitialize() and Dispose().
Checklist for a new arena-allocated class holding a Vulkan handle:
1. Add an explicit destroy call in Shutdown() or Deinitialize().
2. Decide: direct (GPU-idle guaranteed) or deferred.
3. Set the pointer to nullptr after destruction.
4. Never call delete on an arena-allocated pointer.
Platform Notes¶
macOS Apple Silicon — 16 KB pages¶
mprotect rounds to 16 KB boundaries. The arena uses sysconf(_SC_PAGE_SIZE) → 16384 on arm64. A single 1-byte first-allocation commits 16 KB of physical RAM (vs 4 KB on Linux/Windows). Creating many small arenas at startup is 4× more expensive in physical pages than on Linux.
macOS Apple Silicon — Unified Memory Architecture¶
CPU and GPU share the same physical memory pool. The renderer still uses Vulkan/VMA allocations
and staging copies through MoltenVK; it does not expose a direct TLSF-to-MTLBuffer path.
ARM64 weak memory ordering¶
ARM64 (Apple Silicon, Linux ARM) uses a weakly ordered memory model. Concurrent ownership paths
must use the C++ atomic/lock synchronization already present in TLSFSlab and queue handoffs;
the old cross-thread slab-free data-race warning was resolved by that lock.
Linux — Transparent Huge Pages¶
Configured owners use PROT_NONE reservations and promote allocation page ranges with
mprotect, so startup does not depend on permissive overcommit for one whole-root mapping.
madvise for already committed hot ranges remains a separate optimization opportunity.
On Linux with THP = madvise, calling madvise(ptr, size, MADV_HUGEPAGE) on hot arenas promotes pages to 2 MB huge pages. TLB coverage improves from 4 KB × 512 entries = 2 MB to 2 MB × 512 = 1 GB per miss. Measurable win for dense ECS archetype iteration. No code change required beyond one madvise call in ArenaAllocator::Initialize for arenas larger than 2 MB.
Windows — Commit semantics¶
On Windows, VirtualAlloc(MEM_COMMIT) reserves pagefile space immediately (not demand-paged like POSIX mprotect). The engine only commits as the cursor advances — correct. However, auditing memory pressure on Windows requires checking pagefile reservation, not RSS, because the semantics differ from Linux.
Memory Profiler¶
File: ZEngine/ZEngine/Profiling/MemoryProfiler.h
MemoryProfiler records tracked-arena current and peak offsets and warns above 80% with a
60-second cooldown. A budgeted arena is registered only if profiling is enabled. The main-thread
frame boundary calls MemoryProfiler::Update(), so the editor profiler shows a live current and
peak feed for named CPU owners. These values do not represent VMA/driver memory, persistent
environment resources, or render-graph transients; those are separate accounting domains.