-
Notifications
You must be signed in to change notification settings - Fork 2.2k
All issues
Issue creation is restricted in this repository
- #19 · ZacharyZcR opened
on Jul 10, 2026 19 - #537 · ZacharyZcR opened
on Jul 22, 2026 14
Issues
is:issue state:open
is:issue state:open
Search results
- Status: Open.#721 In JustVugg/colibri;
- Status: Open.#720 In JustVugg/colibri;
- Status: Open.#718 In JustVugg/colibri;
[Experiment] expert-transition-history placement policy vs gate-momentum — controlled A/B for hypothesis #1
discussionProposta / discussione aperta, non un taskProposta / discussione aperta, non un taskenhancementNew feature or requestNew feature or requesthelp wantedExtra attention is neededExtra attention is neededperformanceVelocità / tok-s / ottimizzazioniVelocità / tok-s / ottimizzazioniStatus: Open.#708 In JustVugg/colibri;[Bug]: OpenMP spin-wait tuning is inert on macOS — and applying it costs 2.2x decode on a 32 GB host
bugDifetto verificato nel codiceDifetto verificato nel codicehelp wantedExtra attention is neededExtra attention is neededmetalBackend Metal/AppleBackend Metal/AppleperformanceVelocità / tok-s / ottimizzazioniVelocità / tok-s / ottimizzazioniStatus: Open.#707 In JustVugg/colibri;[Performance]: Apple M1 Max 32 GB — 0.15 tok/s, and Metal is a no-op at ~10% expert residency
benchmarkDatapoint di misurazione hardwareDatapoint di misurazione hardwaredocsDocumentazioneDocumentazionemetalBackend Metal/AppleBackend Metal/AppleperformanceVelocità / tok-s / ottimizzazioniVelocità / tok-s / ottimizzazioniStatus: Open.#706 In JustVugg/colibri;The learning cache is engine-specific: .coli_usage has two incompatible writers and two engines that cannot produce it at all
discussionProposta / discussione aperta, non un taskProposta / discussione aperta, non un taskenhancementNew feature or requestNew feature or requestStatus: Open.#700 In JustVugg/colibri;[Performance]:
benchmarkDatapoint di misurazione hardwareDatapoint di misurazione hardwaremetalBackend Metal/AppleBackend Metal/AppleperformanceVelocità / tok-s / ottimizzazioniVelocità / tok-s / ottimizzazioniStatus: Open.#693 In JustVugg/colibri;[Bug]: deep speculative verify batches are not token-exact vs CPU on CUDA (near-tie flips + lower draft acceptance)
bugDifetto verificato nel codiceDifetto verificato nel codicecudaBackend CUDA/NVIDIABackend CUDA/NVIDIAqualityQualità del modello / quantizzazioneQualità del modello / quantizzazioneStatus: Open.#689 In JustVugg/colibri;[Performance]: GLM-5.2 on a single H200 (141 GB) + 235 GB RAM — 4.75 tok/s median, matching the 6x RTX 5090 reference class
benchmarkDatapoint di misurazione hardwareDatapoint di misurazione hardwarecudaBackend CUDA/NVIDIABackend CUDA/NVIDIAperformanceVelocità / tok-s / ottimizzazioniVelocità / tok-s / ottimizzazioniStatus: Open.#688 In JustVugg/colibri;[Bug]: CUDA_EXPERT_GB=auto fills all VRAM on single-GPU; lazy dense uploads then fail 60x and fall back to CPU silently
bugDifetto verificato nel codiceDifetto verificato nel codicecudaBackend CUDA/NVIDIABackend CUDA/NVIDIAStatus: Open.#687 In JustVugg/colibri;- Status: Open.#683 In JustVugg/colibri;