Releases · ggml-org/llama.cpp

31 Jul 18:29

7845240

b6050

Fix params bug in diffusion example (#14993)

Assets 15

31 Jul 17:45

github-actions

b6049

d6818d0

b6049

llama : allow other bufts when overriding to CPU, add --no-repack opt…

Assets 15

31 Jul 17:49

github-actions

b6048

e08a988

b6048

Vulkan: Fix minor debug mode issues (#14899)

* vulkan: fix debug mode issues

* vulkan: remove broken check_results GGML_OP_SET_ROWS support

Assets 15

31 Jul 15:59

github-actions

b6047

952a47f

b6047

mtmd : support MiniCPM-V 4.0 (#14983)

* support minicpm-v 4

* add md

* support MiniCPM-o 4.0

* add default location

* temp rm MiniCPM-o 4.0

* fix code

* fix "minicpmv_projector" default path

Assets 15

31 Jul 14:36

github-actions

b6045

94933c8

b6045

server : implement universal assisted decoding (#12635)

* llama-server : implement universal assisted decoding

* Erase prompt tail for kv-cache

* set vocab_dft_compatible in common_speculative

* rename ctx_main to ctx_tgt

* move vocab_dft_compatible to spec struct

* clear mem_dft, remove mem

* detokenize id_last for incompatible models

* update comment

* add --spec-replace flag

* accept special tokens when translating between draft/main models

* Escape spec-replace

* clamp draft result to size to params.n_draft

* fix comment

* clean up code

* restore old example

* log common_speculative_are_compatible in speculative example

* fix

* Update common/speculative.cpp

Co-authored-by: Georgi Gerganov <[email protected]>

* Update common/speculative.cpp

Co-authored-by: Georgi Gerganov <[email protected]>

* Update common/speculative.cpp

Co-authored-by: Georgi Gerganov <[email protected]>

---------

Co-authored-by: Georgi Gerganov <[email protected]>

Assets 15

31 Jul 14:33

github-actions

b6044

c1dacaa

b6044

llama : merge build_moe_ffn_from_probs function into build_moe_ffn (#…

Assets 15

31 Jul 14:20

github-actions

b6043

a9f77a8

b6043

server : add openai-style logit_bias support (#14946)

Signed-off-by: Lukas Straub <[email protected]>

Assets 15

31 Jul 14:09

github-actions

b6042

8a4a856

b6042

Add LLaDA 8b Diffusion model (#14771)

* Add support for Llada-8b: diffusion model

* Add README

* Fix README and convert_hf_to_gguf

* convert_hf_to_gguf.py: address review comments

* Make everything in a single example

* Remove model-specific sampling

* Remove unused argmax

* Remove braced initializers, improve README.md a bit

* Add diffusion specific gguf params in set_vocab, remove setting rope_theta and rms_norm_eps

* Remove adding the mask token

* Move add_add_bos_token to set_vocab

* use add_bool in gguf_writer.py

Assets 15

31 Jul 13:53

github-actions

b6041

11490b3

b6041

CANN: Improve loading efficiency after converting weights to NZ forma…

Assets 15

31 Jul 06:28

github-actions

b6040

66625a5

b6040

graph : reduce splits for recurrent and hybrid models (#14825)

* graph : avoid creating redundant s_copy views

* graph : comment the s_copy views

Assets 15

Releases: ggml-org/llama.cpp

b6050

Uh oh!

b6049

Uh oh!

b6048

Uh oh!

b6047

Uh oh!

b6045

Uh oh!

b6044

Uh oh!

b6043

Uh oh!

b6042

Uh oh!

b6041

Uh oh!

b6040

Uh oh!