Instructions to use kernels-community/vllm-flash-attn3 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Kernels
How to use kernels-community/vllm-flash-attn3 with Kernels:
# !pip install kernels from kernels import get_kernel # a version (or an explicit revision) is required; see the "Files and versions" tab for the available ones kernel = get_kernel("kernels-community/vllm-flash-attn3", version=1) - Notebooks
- Google Colab
- Kaggle
Support for TPU v5e-8?
#8 opened about 1 year ago
by
emmarosess
Support for B200s?
👀 5
3
#7 opened about 1 year ago
by
shriramc
using SlidingWindowLayer Cache will cause a crash
#5 opened about 1 year ago
by
mdabbah
Not able to find the compatible kernel
4
#4 opened about 1 year ago
by
rom7
attention sinks & backward
4
#3 opened about 1 year ago
by
acforvs
Support for sm120?
1
#2 opened about 1 year ago
by
Enigrand