Summary:
========
Implemented C5 ULTRA TLS cache pattern following the successful C6 ULTRA design:
- Phase 5-1: Free-side TLS cache + segment learning
- Phase 5-2: Alloc-side TLS pop for complete free+alloc cycle integration
Targets C5 class (129-256B) as next legacy reduction after C6 completion.
Key Changes:
============
1. NEW FILES:
- core/box/tiny_c5_ultra_free_box.h: C5 ULTRA TLS cache structure
- core/box/tiny_c5_ultra_free_box.c: C5 free path implementation (same pattern as C6)
- core/box/tiny_c5_ultra_free_env_box.h: ENV gating (HAKMEM_TINY_C5_ULTRA_FREE_ENABLED)
2. MODIFIED FILES:
- core/front/malloc_tiny_fast.h:
* Added C5 ULTRA includes
* Added C5 alloc-side TLS pop at lines 186-194 (integrated with C6)
* Added C5 free path at lines 333-337 (integrated with C6)
- core/box/tiny_ultra_classes_box.h:
* Added TINY_CLASS_C5 constant
* Added tiny_class_is_c5() macro
* Extended tiny_class_is_ultra() to include C5
- core/box/free_path_stats_box.h:
* Added c5_ultra_free_fast counter
* Added c5_ultra_alloc_hit counter
- core/box/free_path_stats_box.c:
* Updated stats dump to output C5 counters
- Makefile:
* Added core/box/tiny_c5_ultra_free_box.o to all object lists
3. Design Rationale:
- Exact copy of C6 ULTRA pattern (proven effective)
- TLS cache capacity: 128 blocks (same as C6 for consistency)
- Segment learning on first C5 free via ss_fast_lookup()
- Alloc-side pop integrated directly in malloc_tiny_fast.h hotpath
- Legacy fallback unification via tiny_legacy_fallback_free_base()
4. Expected Impact:
- C5 legacy calls: 68,871 → 0 (100% elimination)
- Total legacy reduction: ~53% of remaining 129,623
- Mixed workload: Minimal regression (C5 is smaller class, fewer allocations)
5. Stats Collection:
Run with: HAKMEM_TINY_C5_ULTRA_FREE_ENABLED=1 HAKMEM_FREE_PATH_STATS=1 ./bench_allocators_hakmem
Expected output:
[FREE_PATH_STATS] ... c5_ultra_free=68871 c5_ultra_alloc=68871 ... legacy_fb=60752 ...
[FREE_PATH_STATS_LEGACY_BY_CLASS] ... c5=0 ...
Status:
=======
- Code: ✅ COMPLETE (3 new files + 5 modified files)
- Compilation: ✅ Verified (no errors, only unused variable warnings unrelated to C5)
- Functionality: Ready to benchmark (ENV gating: default OFF, opt-in via ENV)
Phase Progression:
==================
✅ Phase 4-4: C6 ULTRA free+alloc (legacy C6: 137,319 → 0)
✅ Phase 5-1/5-2: C5 ULTRA free+alloc (legacy C5: 68,871 → 0 expected)
⏳ Phase 4.5: C4 ULTRA (34,727 remaining)
📋 Future: C3/C2 ULTRA if beneficial
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
47 lines
1.9 KiB
C
47 lines
1.9 KiB
C
#include "free_path_stats_box.h"
|
|
#include <stdio.h>
|
|
|
|
FreePathStats g_free_path_stats = {0};
|
|
|
|
// Helper function for pool_api.inc.h (to avoid inline include issues)
|
|
void free_path_stat_inc_pool_v1_fast(void) {
|
|
if (__builtin_expect(free_path_stats_enabled(), 0)) {
|
|
g_free_path_stats.pool_v1_fast++;
|
|
}
|
|
}
|
|
|
|
__attribute__((destructor))
|
|
static void free_path_stats_dump(void) {
|
|
if (!free_path_stats_enabled()) {
|
|
return;
|
|
}
|
|
|
|
fprintf(stderr, "[FREE_PATH_STATS] total=%lu c7_ultra=%lu c6_ultra_free=%lu c6_ultra_alloc=%lu c5_ultra_free=%lu c5_ultra_alloc=%lu small_v3=%lu v6=%lu tiny_v1=%lu pool_v1=%lu remote=%lu super_lookup=%lu legacy_fb=%lu\n",
|
|
g_free_path_stats.total_calls,
|
|
g_free_path_stats.c7_ultra_fast,
|
|
g_free_path_stats.c6_ultra_free_fast, // Phase 4-2
|
|
g_free_path_stats.c6_ultra_alloc_hit, // Phase 4-4
|
|
g_free_path_stats.c5_ultra_free_fast, // Phase 5-1
|
|
g_free_path_stats.c5_ultra_alloc_hit, // Phase 5-2
|
|
g_free_path_stats.smallheap_v3_fast,
|
|
g_free_path_stats.smallheap_v6_fast,
|
|
g_free_path_stats.tiny_heap_v1_fast,
|
|
g_free_path_stats.pool_v1_fast,
|
|
g_free_path_stats.remote_free,
|
|
g_free_path_stats.super_lookup_called,
|
|
g_free_path_stats.legacy_fallback);
|
|
|
|
// Phase 4-1: Legacy per-class breakdown
|
|
fprintf(stderr, "[FREE_PATH_STATS_LEGACY_BY_CLASS] c0=%lu c1=%lu c2=%lu c3=%lu c4=%lu c5=%lu c6=%lu c7=%lu\n",
|
|
g_free_path_stats.legacy_by_class[0],
|
|
g_free_path_stats.legacy_by_class[1],
|
|
g_free_path_stats.legacy_by_class[2],
|
|
g_free_path_stats.legacy_by_class[3],
|
|
g_free_path_stats.legacy_by_class[4],
|
|
g_free_path_stats.legacy_by_class[5],
|
|
g_free_path_stats.legacy_by_class[6],
|
|
g_free_path_stats.legacy_by_class[7]);
|
|
|
|
fflush(stderr);
|
|
}
|