Add Sparse VideoGen attention for Wan on TPUs
Implement opt-in SVG attention for Wan inference with per-head spatial and temporal routing, head-local token layouts, and TPU sparse-attention kernels. Round sparse support to hardware tiles while preserving the approximate pair budget, retain valid support for every query, and use a padding-only cleanup path for sequence tails. Add configuration and scheduling controls, distributed inference integration, cache and unsupported-mode guards, AOT cache identity handling, documentation, and CPU/TPU regression coverage.
R
Ravisri Valluri committed
db8c0cd45942cbbca9a23865a7ac1e356c62b12e
Parent: 1bc5481