Skip to content

Memory Management

Status: Current allocator primitives and GPU-domain reference, with historical planning notes called out below. The authoritative configured CPU profile is design/memory-budget.md.

ZEngine primarily uses explicit arena, pool, and TLSF-slab ownership. This page documents the live primitives, allocation patterns, GPU memory domains, and teardown rules for Vulkan handles. It does not claim that every allocation in the repository avoids the system heap: third-party libraries and a few non-hot-path facilities still use their own allocation policies.

See also: Engine Architecture · Asset Manager


Table of Contents


Philosophy

  1. Named up-front reservations. A configured MemoryManager independently reserves each named CPU owner. The 8 GiB value passed by Obelisk is a profile-capacity validation limit, not one MainArena mapping. Individual objects never call malloc/new outside of third-party libraries.
  2. Owners and child arenas have fixed budgets. Each subsystem gets a dedicated owner sized to its worst-case working set; short-lived consumers carve child arenas from that owner. Running out of either is a budgeting error to fix at design time, not a runtime failure to handle.
  3. Lifetime = scope. Objects allocated from an arena are freed by ArenaAllocator::Clear() (cursor reset). There is no per-object free. Pick the allocator whose lifetime matches the object's lifetime.
  4. No destructor guarantee. ZPushStructCtor places objects via placement-new, but arena release does not call destructors. Any object that owns an OS or GPU resource must have its destructor called explicitly before the arena is cleared.
  5. Zero hot-path touches. Alloc/free on the render thread or in inner simulation loops is off the table.

Before writing new or std::vector, identify the lifetime. Pick the cheapest allocator that matches it. If nothing fits, the lifetime is unclear — clarify it first.


Lifetime Model

Each allocation belongs to exactly one lifetime tier:

Tier When freed Allocator Examples
Engine Shutdown only ArenaAllocator VulkanDevice, ECSScene, AssetManager arenas
Scene Scene load/unload ArenaAllocator target editor/session data; it is not yet a separately budgeted production editor arena
Per-task Task-specific boundary ArenaAllocator or TLSFSlab importer scratch and variable-lifetime upload/decode data
Per-frame End of frame / GPU completion ArenaTemp or PerFrameUploadHeap CPU scratch; mapped GPU uniform/storage/indirect uploads
Per-object Individual free needed PoolAllocator Entity slots, command buffer handles, mesh instance slots
Variable Individual free, variable size TLSFSlab Texture decode buffers and growable engine containers

Allocation Decision Framework

New allocation needed
        │
        ▼
Group lifetime? (reset all at once after frame / import / scene)
    YES ──► Stable after init? (no grows once setup is done)
                YES ──► ArenaAllocator  (~3 cyc + memset)
                NO  ──► TLSFSlab        (~30 cyc, O(1) — roadmap v0.5.0)
    NO
        │
        ▼
    Fixed size? (same N bytes every time)
        YES ──► PoolAllocator  (~5 cyc + memset(chunk), O(1))
        NO
            │
            ▼
        Variable size + individual lifetime?
            YES ──► TLSFSlab  (~30 cyc, O(1) — roadmap v0.5.0)
            NO  ──► Re-examine the lifetime. Do NOT use std::vector / new.

CPU Memory — ArenaAllocator

File: ZEngine/ZEngine/Core/Memory/Allocator.h

struct ArenaAllocator
{
    void  Initialize(size_t size);
    void* Allocate(size_t size, size_t alignment = DEFAULT_ALIGNMENT);
    void* Resize(void* ptr, size_t old_size, size_t new_size, size_t alignment);
    void  CreateSubArena(size_t size, ArenaAllocator* out);
    void  Clear();     // reset cursor to 0, keep pages
    void  Shutdown();  // unmap pages
};

Allocate bumps a cursor — O(1), no locks. Virtual address space is reserved up-front; physical pages are committed on first write (RSS much lower than virtual reservation). Every Allocate calls secure_memset(ptr, 0, n) — negligible for small objects, dominant for large buffers (e.g. ~0.4 ms for a 16 MB decode buffer at 40 GB/s).

Resize extends in-place if the pointer is the most recent allocation; otherwise allocates a new block forward and copies, leaving the old block permanently dead. Safe for scratch arenas; a slow memory leak for long-lived growing containers.

CreateSubArena advances the parent cursor by size. The child manages its own cursor independently.

Key invariants

  • ArenaAllocator is not thread-safe — all arenas are carved on the main thread before any worker starts.
  • Arena release does not call destructors. Call ptr->~T() explicitly on objects owning OS/GPU handles.
  • ZReleaseScratch pairs must be released in strict LIFO order — see Scratch Arenas.

CPU Memory — PoolAllocator

File: ZEngine/ZEngine/Core/Memory/Allocator.h

Fixed-size free list backed by a single arena carve at init. Suited for objects of a single known size (entity slots, component handles) with individual lifetimes.

struct PoolAllocator
{
    void  Initialize(ArenaAllocator* arena, size_t total_size,
                     size_t chunk_size, size_t alignment = DEFAULT_ALIGNMENT);
    void* Allocate();   // O(1) — pop free-list head, zero chunk
    void  Free(void*);  // O(1) — push free-list head; asserts range + alignment
    void  Clear();      // O(capacity) — zero all chunks, rebuild free list
};

Free list links are stored inside free chunks — zero separate metadata. After Clear() or across alloc/free cycles, allocation order is LIFO-scrambled; sequential layout is only guaranteed at init.

Safety invariants

Check Enforced?
Free ptr in range Always-on assert
Free ptr chunk-aligned Always-on assert
Double-free Not detected — second Free corrupts the free list silently (issue #697)
Exhaustion Allocate returns nullptr — caller must check (issue #681)

When NOT to use PoolAllocator

  • Multiple object sizes — requires multiple pools or wasteful over-sizing to largest.
  • Capacity unknown at init — no in-place growth; growing requires a new arena carve.
  • Per-frame Clear() — O(capacity) traversal is too expensive.

CPU Memory — TLSFSlab

Status: Implemented. TLSFSlab is backed by mattconte/tlsf, protected by an internal atomic spinlock, and used by the asset manager, render-resource manager, bitmap helpers, and TLSF-aware Array/UnorderedHashMap variants.

TLSFSlab wraps mattconte/tlsf (already vendored via FetchContent) with a backing buffer carved from a parent ArenaAllocator. Fills the gap for variable-size, individually-freed allocations that neither Arena nor Pool can handle: texture decode buffers, asset metadata containers that grow unpredictably.

struct TLSFSlab {
    void   Init(ArenaAllocator* arena, size_t bytes);
    void*  Alloc(size_t n);         // O(1) worst-case — asserts on exhaustion
    void*  Realloc(void* ptr, size_t n);  // O(1) if in-place, O(n) copy otherwise
    void   Free(void* ptr);         // O(1) — coalesces with adjacent free blocks
    void   Shutdown();              // tlsf_destroy; does NOT free backing
    size_t Overhead() const;
};

The backing buffer is carved from the parent arena once at Init. Subsequent Alloc/Free never touch the arena. Internal fragmentation is bounded at ≤ 1.0625× requested size. Adjacent frees always coalesce — no fragmentation cliff over time.

Current texture/upload use

Texture/upload work can use worker slabs and the render-resource manager's upload slabs. The slab's lock makes a cross-thread Free safe; it is not a lock-free per-worker-only allocator. Individual importers and third-party decoders may still have their own allocations, so this is not a promise of zero process-heap activity.

Historical rollout record

Phase Target Scope
1 v0.5.0 target Per-worker upload slabs, TextureDeferral refactor, STBI_MALLOC override
2 v0.6.0 target AssetManager containers (typed allocator for Array<T> / UnorderedHashMap)
3 v1.0.0 Per-archetype ECS slab for variable-payload component types

Allocation Macros

File: ZEngine/ZEngine/ZEngineDef.h

Macro Equivalent Notes
ZKilo(n) uint64_t(n) * 1024 Always 64-bit — no overflow
ZMega(n) uint64_t(n) * 1024² Always 64-bit
ZGiga(n) uint64_t(n) * 1024³ Always 64-bit
ZPushArray(arena, T, count) arena->Allocate(count * sizeof(T), alignof(T)) Returns T*, no constructor
ZPushStruct(arena, T) ZPushArray(arena, T, 1) Returns T*, no constructor
ZPushStructCtor(arena, T) new (ZPushStruct(arena, T)) T() Placement-new, default constructor
ZPushStructCtorArgs(arena, T, ...) new (ZPushStruct(arena, T)) T(...) Placement-new with args

When to use each: - ZPushStruct / ZPushArray — POD structs, trivial types. - ZPushStructCtor — objects with non-trivial default constructor (CommandPool, Semaphore, …). - ZPushStructCtorArgs — objects requiring constructor arguments (GameWindow, VulkanDevice, …).

Calling delete on an arena-allocated pointer is undefined behavior. Call ptr->~T() explicitly, then set the pointer to nullptr.


Scratch Arenas

Short-lived per-call temporaries use a scratch arena to avoid polluting long-lived arenas.

sequenceDiagram
    participant Code as Caller
    participant SA as ZGetScratch / ZReleaseScratch
    participant TA as Thread-local arena pair [A, B]

    Code->>SA: ZGetScratch(&my_arena)
    SA->>TA: pick arena that is NOT &my_arena
    SA-->>Code: ScratchArena { .Arena = chosen, .checkpoint }
    Code->>Code: allocate temporaries from scratch.Arena
    Code->>SA: ZReleaseScratch(scratch)
    SA->>TA: reset chosen arena cursor to checkpoint

Rules: - Never store a pointer into a scratch arena past ZReleaseScratch. - Always pair ZGetScratch / ZReleaseScratch — no early returns between them. - ZGetScratch / ZReleaseScratch must be released in strict LIFO order. Releasing an outer scratch while an inner scratch is still live leaves the inner's save point stale — a subtle corruption that manifests later. - Each thread has its own arena pair — scratch arenas are not shared across threads.


Thread Safety

Allocator Thread-safe? Notes
ArenaAllocator No All arenas carved on main thread before workers start. Workers never call ArenaAllocator::Allocate after init.
PoolAllocator No All current pools are single-threaded (main thread or one render thread). CAS / spinlock needed if shared.
TLSFSlab Yes An internal atomic spinlock protects Alloc, Realloc, and Free, including cross-thread release.
GpuAllocator (VMA) Yes VMA handles its own synchronization internally.

Memory Budget

MemoryBudgetConfig in ZEngine/ZEngine/Core/Memory/MemoryManager.h defines the profile. Obelisk validates that profile against an 8 GiB configured-capacity limit before initialization. The actual configured totals are 7,604 MiB (Default), 7,868 MiB (Editor), and 6,324 MiB (Server); the largest slots are ImportPipeline (4,096 MiB), AssetManager (1,024 MiB by default and 1,280 MiB for the calibrated editor profile), VulkanDevice (1,024 MiB), and ECSScene (512 MiB). UIContext is 64 MiB by default and 128 MiB for the editor. Not every declared slot is materialized by every runtime mode.

Configured owners are separate virtual reservations: Windows starts them PAGE_NOACCESS, while macOS/Linux start them PROT_NONE. The allocator promotes only the required page range on each allocation (VirtualAlloc(MEM_COMMIT) or mprotect), with owner-local page tracking for nested child arenas. Consequently, configured Linux startup does not need one writable 8 GiB mapping; the test suite exercises it below a constrained RLIMIT_AS limit. Do not use MAP_NORESERVE alone as a replacement for explicit commit admission, because it can defer failure to a later uncontrolled write.

The GPU allocator's domains and the 384 MiB persistent environment-lighting gate are separate GPU policies, not deductions from this CPU profile. See design/memory-budget.md for the full live table.


Performance Comparison

Approximate cycle counts on a cache-warm allocation path (bookkeeping only — does not include memset(n) zeroing which scales linearly with size):

Allocator Alloc cost Free cost Fragmentation Best for
ArenaAllocator ~3–5 cyc + memset(n) N/A Zero Scratch, import, per-frame
PoolAllocator ~5–8 cyc + memset(chunk) ~5–8 cyc Zero Entity slots, fixed-size objects
TLSFSlab ~20–40 cyc ~20–40 cyc ≤ 1.0625× Upload buffers, growing containers
System heap (jemalloc) ~50–300 cyc ~50–300 cyc Accumulates Nothing on the hot path
System heap (ptmalloc) ~100–500 cyc ~100–500 cyc Accumulates Nothing on the hot path

The cycle counts are planning estimates, not a current benchmark suite. TLSFSlab is available for variable-size individual lifetimes; it is not a claim that all such allocations have been migrated.


Container Ownership Rules

File: ZEngine/ZEngine/Core/Containers/Array.h

Array<T> is move-only

Array<T> copy constructor and copy assignment are deleted. The arena owns the backing memory; a shallow copy would alias the same buffer. Moving transfers the pointer and nulls the source.

Array<T>(const Array&)             = delete;
Array<T>& operator=(const Array&)  = delete;
Array<T>(Array&& other) noexcept;
Array<T>& operator=(Array&& other) noexcept;

Passing conventions

void Inspect(const Array<uint32_t>& arr);   // read-only
void Mutate(Array<uint32_t>& arr);          // in-place mutation
void Consume(Array<uint32_t> arr);          // ownership transfer — caller std::move()

ArrayView<T> for non-owning slices

ArrayView<T> is a plain {T*, size_t} — freely copyable, no ownership semantics.

Growing containers and allocator choice

Every Array<T>::grow() that reallocates on an arena abandons the old block for that arena's lifetime. Pre-size stable containers. For a genuinely variable, individually released workload, use the implemented TLSF-aware container initialization rather than treating arena growth as a general allocator.

HashMap / UnorderedHashMap with move-only values

map.insert(key, std::move(my_array));        // rvalue overload for move-only values
for (auto& [k, v] : my_map) { v.push(42); } // reference — no copy

The insert(const K&, const V&) overload is gated with requires std::is_copy_assignable_v<V> — using it with a move-only value is a compile error.


GPU Memory — VMA Allocator

File: ZEngine/ZEngine/Core/Memory/GpuAllocator.h

GPU memory is managed by Vulkan Memory Allocator (VMA). GpuAllocator wraps VmaAllocator and exposes typed helpers:

BufferView  AllocateBuffer(VkDeviceSize, VkBufferUsageFlags, GpuMemoryDomain, const char* debug_name);
void        FreeBuffer(BufferView&);

BufferImage AllocateImage(VkImageCreateInfo&, GpuMemoryDomain, VkDevice,
                           VkImageAspectFlagBits, VkImageViewType, uint32_t layers, const char*);
void        FreeImage(BufferImage&, VkDevice);

BufferView and BufferImage hold raw VkHandles + VmaAllocation. They are not arena-allocated and must be freed explicitly before the device is destroyed.


GPU Memory Domains

graph LR
    DG["DeviceGeometry\nVMA_MEMORY_USAGE_AUTO\ndevice-local preferred → VRAM\nGlobal vertex/index/storage buffers"]
    DT["DeviceTexture\nVMA_MEMORY_USAGE_AUTO\ndevice-local preferred → VRAM\nTexture images"]
    HU["HostUniform\nVMA_MEMORY_USAGE_AUTO\nhost-visible required → BAR / shared\nTransformSB, DrawDataSB"]
    HS["HostStaging\nPersistent staging ring plus one-shot fallback\nSequential CPU writes"]
    HR["HostReadback\nMapped read-mostly allocations\nGPU-to-CPU readback"]

Rule: HostUniform includes the mapped per-frame heaps and is flushed when non-coherent. DeviceGeometry and DeviceTexture receive GPU copies through the staging ring or a one-shot fallback. RenderTarget intentionally follows the default allocator path rather than a custom domain pool.


Arena-Allocated Vulkan Objects

Arena Clear() does not call destructors. Objects holding VkCommandPool, VkSemaphore, etc. must have their destructor called explicitly before the device is destroyed.

flowchart TD
    A["Arena-allocated object owns VkHandle"]
    B["Subsystem Shutdown() / Deinitialize()"]
    C{"GPU-idle\nguaranteed?"}
    D["Direct: ptr→~T() → vkDestroy*\nat QueueWaitAll point"]
    E["Deferred: Device→DeferFree(entry)\ndrained when timeline value ≥ stamp"]
    F["ptr = nullptr"]

    A --> B --> C
    C -->|Yes| D --> F
    C -->|No| E --> F
Class Strategy Reason
CommandPool Direct Always freed at GPU-idle
FramebufferVNext Direct Called after QueueWaitAll
GraphicPipeline Direct Same
Semaphore Deferred Can be signalled; deferred prevents in-flight use
Fence Deferred Same

DeferredFreeQueue is a 2048-slot circular buffer, drained in Deinitialize() and Dispose().

Checklist for a new arena-allocated class holding a Vulkan handle: 1. Add an explicit destroy call in Shutdown() or Deinitialize(). 2. Decide: direct (GPU-idle guaranteed) or deferred. 3. Set the pointer to nullptr after destruction. 4. Never call delete on an arena-allocated pointer.


Platform Notes

macOS Apple Silicon — 16 KB pages

mprotect rounds to 16 KB boundaries. The arena uses sysconf(_SC_PAGE_SIZE) → 16384 on arm64. A single 1-byte first-allocation commits 16 KB of physical RAM (vs 4 KB on Linux/Windows). Creating many small arenas at startup is 4× more expensive in physical pages than on Linux.

macOS Apple Silicon — Unified Memory Architecture

CPU and GPU share the same physical memory pool. The renderer still uses Vulkan/VMA allocations and staging copies through MoltenVK; it does not expose a direct TLSF-to-MTLBuffer path.

ARM64 weak memory ordering

ARM64 (Apple Silicon, Linux ARM) uses a weakly ordered memory model. Concurrent ownership paths must use the C++ atomic/lock synchronization already present in TLSFSlab and queue handoffs; the old cross-thread slab-free data-race warning was resolved by that lock.

Linux — Transparent Huge Pages

Configured owners use PROT_NONE reservations and promote allocation page ranges with mprotect, so startup does not depend on permissive overcommit for one whole-root mapping. madvise for already committed hot ranges remains a separate optimization opportunity.

On Linux with THP = madvise, calling madvise(ptr, size, MADV_HUGEPAGE) on hot arenas promotes pages to 2 MB huge pages. TLB coverage improves from 4 KB × 512 entries = 2 MB to 2 MB × 512 = 1 GB per miss. Measurable win for dense ECS archetype iteration. No code change required beyond one madvise call in ArenaAllocator::Initialize for arenas larger than 2 MB.

Windows — Commit semantics

On Windows, VirtualAlloc(MEM_COMMIT) reserves pagefile space immediately (not demand-paged like POSIX mprotect). The engine only commits as the cursor advances — correct. However, auditing memory pressure on Windows requires checking pagefile reservation, not RSS, because the semantics differ from Linux.


Memory Profiler

File: ZEngine/ZEngine/Profiling/MemoryProfiler.h

Profiling::MemoryProfiler::TrackArena("AssetManager", &AssetArena);

MemoryProfiler records tracked-arena current and peak offsets and warns above 80% with a 60-second cooldown. A budgeted arena is registered only if profiling is enabled. The main-thread frame boundary calls MemoryProfiler::Update(), so the editor profiler shows a live current and peak feed for named CPU owners. These values do not represent VMA/driver memory, persistent environment resources, or render-graph transients; those are separate accounting domains.