ggml-cuda: provide static workspace for cuBLAS handles (#26574)
* provide static workspace for cuBLAS handles * account for concurrent streams when using GGML_CUDA_GRAPH_OPT * drop cublas_handle overloads and remove direct cublasSetStream calls * Update ggml/src/ggml-cuda/common.cuh --------- Co-authored-by: Oliver Simons <osimons@nvidia.com>
A
Alexander Heisler committed
d9b6be07d0864ab09417b17ba36f9788087dd22c
Parent: 929d47a
Committed by GitHub <noreply@github.com>
on 8/20/2026, 7:27:51 AM