Skip to content
Navigation Menu
Sign in
Appearance settings
Platform
AI CODE CREATION
GitHub Copilot
Write better code with AI
GitHub Copilot app
Direct agents from issue to merge
MCP Registry
Integrate external tools
DEVELOPER WORKFLOWS
Actions
Automate any workflow
Codespaces
Instant dev environments
Issues
Plan and track work
Code Review
Manage code changes
Code Quality
Enforce quality at merge
APPLICATION SECURITY
GitHub Advanced Security
Find and fix vulnerabilities
Code security
Secure your code as you build
Secret protection
Stop leaks before they start
EXPLORE
Why GitHub
Documentation
Blog
Changelog
Marketplace
View all features
Solutions
BY COMPANY SIZE
Enterprises
Small and medium teams
Startups
Nonprofits
BY USE CASE
App Modernization
DevSecOps
DevOps
CI/CD
View all use cases
BY INDUSTRY
Healthcare
Financial services
Manufacturing
Government
View all industries
View all solutions
Resources
EXPLORE BY TOPIC
AI
Software Development
DevOps
Security
View all topics
EXPLORE BY TYPE
Customer stories
Events & webinars
Ebooks & reports
Business insights
GitHub Skills
SUPPORT & SERVICES
Documentation
Customer support
Community forum
Trust center
Partners
View all resources
Open Source
COMMUNITY
GitHub Sponsors
Fund open source developers
PROGRAMS
Security Lab
Maintainer Community
Accelerator
GitHub Stars
Archive Program
REPOSITORIES
Topics
Trending
Collections
Enterprise
ENTERPRISE SOLUTIONS
Enterprise platform
AI-powered developer platform
AVAILABLE ADD-ONS
GitHub Advanced Security
Enterprise-grade security features
Copilot for Business
Enterprise-grade AI features
Premium Support
Enterprise-grade 24/7 support
Pricing
Search
/
Sign in
Sign up
Appearance settings
You signed in with another tab or window.
Reload
to refresh your session.
You signed out in another tab or window.
Reload
to refresh your session.
You switched accounts on another tab or window.
Reload
to refresh your session.
Dismiss alert
{{ message }}
Uh oh!
There was an error while loading.
Please reload this page
.
ggml-org
/
llama.cpp
Public
Notifications
You must be signed in to change notification settings
Fork
21.7k
Star
124k
Code
Issues
716
Pull requests
1.3k
Discussions
Actions
Projects
Wiki
Security and quality
13
Insights
Additional navigation options
Code
Issues
Pull requests
Discussions
Actions
Projects
Wiki
Security and quality
Insights
Actions: ggml-org/llama.cpp
Actions
All workflows
Workflows
Make Release
Make Release
Publish Docker image
Publish Docker image
Release
Release
Update Winget Package
Update Winget Package
.github/workflows/build-and-test-hexagon.yml
.github/workflows/build-and-test-hexagon.yml
Build Actions Cache
Build Actions Cache
Build and Test - Hexagon Android (QDC)
Build and Test - Hexagon Android (QDC)
Build on RISCV Linux Machine by Cloud-V
Build on RISCV Linux Machine by Cloud-V
Build relocatable cmake package
Build relocatable cmake package
Check Pre-Tokenizer Hashes
Check Pre-Tokenizer Hashes
Show more workflows...
Management
Caches
Deployments
Release
Release
Actions
Loading...
Loading
Sorry, something went wrong.
Uh oh!
There was an error while loading.
Please reload this page
.
will be ignored since log searching is not yet available
Show workflow options
Create status badge
Create status badge
Loading
Uh oh!
There was an error while loading.
Please reload this page
.
release.yml
will be ignored since log searching is not yet available
2,500+ workflow runs
2,500+ workflow runs
Event
Filter by Event
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching events.
Status
Filter by Status
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching statuses.
Branch
Filter by Branch
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching branches.
Actor
Filter by Actor
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching users.
chat : pass reasoning_effort to template
Release
#3957:
Commit
7e4c0a9
pushed by
aldehir
1h 5m 51s
master
master
1h 5m 51s
View workflow file
sync : ggml
Release
#3956:
Commit
9b05354
pushed by
ggerganov
2h 19m 57s
master
master
2h 19m 57s
View workflow file
ggml : recurrent state rollback for ggml_ssm_scan (#26623)
Release
#3955:
Commit
1692f9e
pushed by
ggerganov
3h 28m 11s
master
master
3h 28m 11s
View workflow file
llama : allow virtual igpu devices (#26953)
Release
#3954:
Commit
4c1a0af
pushed by
ggerganov
29s
master
master
29s
View workflow file
server: allow accessing /metrics and /slots during llama_decode() (#2…
Release
#3953:
Commit
77918ca
pushed by
ngxson
58m 51s
master
master
58m 51s
View workflow file
tests : replace personal home directory paths with generic placeholde…
Release
#3952:
Commit
885c5bb
pushed by
CISC
38m 59s
master
master
38m 59s
View workflow file
sycl: fuse mul_mat(gate) + mul_mat(up) + GLU for q4_K dense FFN (#26779)
Release
#3951:
Commit
6509138
pushed by
Titaniumtown
1h 52m 58s
master
master
1h 52m 58s
View workflow file
ggml: force single thread on wasi (#25686)
Release
#3950:
Commit
c6f6a92
pushed by
ggerganov
1h 19m 45s
master
master
1h 19m 45s
View workflow file
sycl: fuse the gated-delta-net state writeback cpy (#26643)
Release
#3949:
Commit
3d93885
pushed by
ggerganov
57m 45s
master
master
57m 45s
View workflow file
dflash : clarify output logging of target_layer_ids (#27013)
Release
#3948:
Commit
2bacf9e
pushed by
danbev
52m 44s
master
master
52m 44s
View workflow file
common: apply CPU parameters across tools (#27026)
Release
#3947:
Commit
a94d563
pushed by
ngxson
4h 8m 16s
master
master
4h 8m 16s
View workflow file
OpenVINO: Qwen3.5, memory optimization, and test-recurrent-state-roll…
Release
#3946:
Commit
aee56b3
pushed by
ggerganov
4h 45m 46s
master
master
4h 45m 46s
View workflow file
[SYCL] Support host pinned mem to improve SYCL Host-to-Device Memory …
Release
#3945:
Commit
a97123e
pushed by
ggerganov
4h 36m 44s
master
master
4h 36m 44s
View workflow file
chat : fix LFM2 tool call arg name prefix ambiguity (#26960)
Release
#3944:
Commit
2606220
pushed by
ngxson
4h 38m 50s
master
master
4h 38m 50s
View workflow file
server : serve index.html with no-cache (#27006)
Release
#3943:
Commit
981184e
pushed by
ngxson
5h 12m 58s
master
master
5h 12m 58s
View workflow file
spec : auto-detect mtp draft model type (#27005)
Release
#3942:
Commit
1d2869c
pushed by
ngxson
7h 55m 24s
master
master
7h 55m 24s
View workflow file
metal : add TQ2_0 support (#26980)
Release
#3941:
Commit
4a84b0a
pushed by
ggerganov
7h 17m 41s
master
master
7h 17m 41s
View workflow file
common : auto-detect spec type from draft GGUF metadata (#26814)
Release
#3940:
Commit
f65e568
pushed by
CISC
7h 30m 36s
master
master
7h 30m 36s
View workflow file
spec: enable backend sampling for both dflash & dspark (#26958)
Release
#3939:
Commit
0d0bfcd
pushed by
ggerganov
8h 17m 52s
master
master
8h 17m 52s
View workflow file
ggml-cpu/ops: vectorize flash-attention V-cache F16 to F32 conversion…
Release
#3938:
Commit
eeae28b
pushed by
ggerganov
7h 36m 36s
master
master
7h 36m 36s
View workflow file
sycl: remove separate fp32 type promotion in gemm non-oneDNN path (#2…
Release
#3937:
Commit
154d57a
pushed by
ggerganov
6h 55m 57s
master
master
6h 55m 57s
View workflow file
sycl: fuse UNARY(silu|sigmoid|softplus) + MUL (#26411)
Release
#3936:
Commit
1ee1cd9
pushed by
ggerganov
6h 27m 36s
master
master
6h 27m 36s
View workflow file
sycl : Add DMMV ESIMD Q3_K kernel (#26251)
Release
#3935:
Commit
8efbf65
pushed by
ggerganov
6h 3m 10s
master
master
6h 3m 10s
View workflow file
sycl : enhance concat to support Q4_0, Q4_1, Q5_0, Q5_1, Q8_0 (#26800)
Release
#3934:
Commit
d415e65
pushed by
ggerganov
2h 28m 5s
master
master
2h 28m 5s
View workflow file
server: refactor + correctness fixes for metrics (#26920)
Release
#3933:
Commit
decaf50
pushed by
ServeurpersoCom
35m 59s
master
master
35m 59s
View workflow file
Previous
1
2
3
4
5
…
99
100
101
Next
You can’t perform that action at this time.