Releases · ggerganov/llama.cpp

09 May 13:41

fd9f92b

b2828 Latest

Latest

llama : update llama_timings.n_p_eval setting (#7160)

This commit changes the value assigned to llama_timings.n_p_eval when
ctx->n_p_eval is 0 to be 1 instead of 1 which is the current value.

The motivation for this change is that if session caching is enabled,
for example using the `--prompt-cache main-session.txt` command line
argument for the main example, and if the same prompt is used then on
subsequent runs, the prompt tokens will not actually be passed to
llama_decode, and n_p_eval will not be updated by llama_synchoronize.

But the value of n_p_eval will be set 1 by llama_get_timings because
ctx->n_p_eval will be 0. This could be interpreted as 1 token was
evaluated for the prompt which could be misleading for applications
using this value.

Signed-off-by: Daniel Bevenius <daniel.bevenius@gmail.com>

Assets 19

cudart-llama-bin-win-cu11.7.1-x64.zip

293 MB 2024-05-09T13:41:42Z
cudart-llama-bin-win-cu12.2.0-x64.zip

413 MB 2024-05-09T13:41:48Z
llama-b2828-bin-macos-arm64.zip

40.3 MB 2024-05-09T13:41:56Z
llama-b2828-bin-macos-x64.zip

36.9 MB 2024-05-09T13:41:58Z
llama-b2828-bin-ubuntu-x64.zip

45.5 MB 2024-05-09T13:41:59Z
llama-b2828-bin-win-arm64-x64.zip

5.98 MB 2024-05-09T13:42:00Z
llama-b2828-bin-win-avx-x64.zip

6.56 MB 2024-05-09T13:42:01Z
llama-b2828-bin-win-avx2-x64.zip

6.53 MB 2024-05-09T13:42:01Z
llama-b2828-bin-win-avx512-x64.zip

6.55 MB 2024-05-09T13:42:02Z
llama-b2828-bin-win-clblast-x64.zip

7.73 MB 2024-05-09T13:42:03Z
Source code (zip)

2024-05-09T11:03:29Z
Source code (tar.gz)

2024-05-09T11:03:29Z

09 May 10:34

github-actions

b2826

4734524

b2826

opencl : alignment size converted from bits to bytes (#7090)

* opencl alignment size should be converted from bits to bytes

Reference: https://registry.khronos.org/OpenCL/specs/3.0-unified/html/OpenCL_API.html#CL_DEVICE_MEM_BASE_ADDR_ALIGN

> Alignment requirement (in bits) for sub-buffer offsets.

* Update ggml-opencl.cpp for readability using division instead of shift

Co-authored-by: Jared Van Bortel <cebtenzzre@gmail.com>

---------

Co-authored-by: Jared Van Bortel <cebtenzzre@gmail.com>

Assets 19

09 May 00:59

github-actions

b2824

4426e29

b2824

cmake : fix typo (#7151)

Assets 19

09 May 00:13

github-actions

b2822

bc4bba3

b2822

Introduction of CUDA Graphs to LLama.cpp (#6766)

* DRAFT: Introduction of CUDA Graphs to LLama.cpp

* FIx issues raised in comments

* Tidied to now only use CUDA runtime (not mixed with driver calls)

* disable for multi-gpu and batch size > 1

* Disable CUDA graphs for old GPU arch and with env var

* added missing CUDA_CHECKs

* Addressed comments

* further addressed comments

* limit to GGML_ALLOW_CUDA_GRAPHS defined in llama.cpp cmake

* Added more comprehensive graph node checking

* With mechanism to fall back if graph capture fails

* Revert "With mechanism to fall back if graph capture fails"

This reverts commit eb9f15fb6fcb81384f732c4601a5b25c016a5143.

* Fall back if graph capture fails and address other comments

* - renamed GGML_ALLOW_CUDA_GRAPHS to GGML_CUDA_USE_GRAPHS

- rename env variable to disable CUDA graphs to GGML_CUDA_DISABLE_GRAPHS

- updated Makefile build to enable CUDA graphs

- removed graph capture failure checking in ggml_cuda_error
  using a global variable to track this is not thread safe, but I am also not safistied with checking an error by string
  if this is necessary to workaround some issues with graph capture with eg. cuBLAS, we can pass the ggml_backend_cuda_context to the error checking macro and store the result in the context

- fixed several resource leaks

- fixed issue with zero node graphs

- changed fixed size arrays to vectors

- removed the count of number of evaluations before start capturing, and instead changed the capture mode to relaxed

- removed the check for multiple devices so that it is still possible to use a single device, instead checks for split buffers to disable cuda graphs with -sm row

- changed the op for checking batch size to GGML_OP_ADD, should be more reliable than GGML_OP_SOFT_MAX

- code style fixes

- things to look into
  - VRAM usage of the cudaGraphExec_t, if it is significant we may need to make it optional
  - possibility of using cudaStreamBeginCaptureToGraph to keep track of which ggml graph nodes correspond to which cuda graph nodes

* fix build without cuda graphs

* remove outdated comment

* replace minimum cc value with a constant

---------

Co-authored-by: slaren <slarengh@gmail.com>

Assets 19

08 May 23:52

github-actions

b2821

c12452c

b2821

JSON: [key] -> .at(key), assert() -> GGML_ASSERT (#7143)

Assets 19

08 May 22:56

github-actions

b2820

9da243b

b2820

Revert "llava : add support for moondream vision language model (#6899)"

This reverts commit 46e12c4692a37bdd31a0432fc5153d7d22bc7f72.

Assets 19

08 May 22:43

github-actions

b2818

26458af

b2818

metal : use `vm_allocate` instead of `posix_memalign` on macOS (#7078)

* fix: use `malloc` instead of `posix_memalign` in `ggml-metal.m` to make it not crash Electron proccesses

* fix: typo

* fix: use `vm_allocate` instead of `posix_memalign`

* fix: don't call `newBufferWithBytesNoCopy` with `NULL` when `ggml_metal_host_malloc` returns `NULL`

* fix: use `vm_allocate` only on macOS

Assets 19

08 May 18:14

github-actions

b2817

83330d8

b2817

main : add --conversation / -cnv flag (#7108)

Assets 19

08 May 17:43

github-actions

b2816

465263d

b2816

sgemm : AVX Q4_0 and Q8_0 (#6891)

* basic avx implementation

* style

* combine denibble with load

* reduce 256 to 128 (and back!) conversions

* sse load

* Update sgemm.cpp

* oops

oops

Assets 19

08 May 17:15

github-actions

b2815

911b390

b2815

server : add_special option for tokenize endpoint (#7059)

Assets 19

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Releases: ggerganov/llama.cpp

b2828

b2826

b2824

b2822

b2821

b2820

b2818

b2817

b2816

b2815