[executorch][cuda] Optimize short-query INT4 matvec kernels (#21503)
This PR was created by the merge bot to help merge the original PR into the main branch. ghstack PR number: https://github.com/pytorch/executorch/pull/21473 by @Gasoonjia ^ Please use this as the source of truth for the PR details, comments, and reviews ghstack PR base: https://github.com/pytorch/executorch/tree/gh/gasoonjia/178/base ghstack PR head: https://github.com/pytorch/executorch/tree/gh/gasoonjia/178/head Merge bot PR base: https://github.com/pytorch/executorch/tree/main Merge bot PR head: https://github.com/pytorch/executorch/tree/gh/gasoonjia/178/orig Differential Revision: [D114032326](https://our.internmc.facebook.com/intern/diff/D114032326/) @diff-train-skip-merge Co-authored-by: gasoonjia <gasoonjia@icloud.com>
P
pytorchbot committed
f1ab0f8204f51b9e956e56f5ca55bb316522d681
Parent: d632341
Committed by GitHub <noreply@github.com>
on 7/30/2026, 10:36:37 PM