[FP8][Kernel] Dynamic kv cache scaling factors computation #11906

gshtras · 2025-01-09T18:35:37Z

This PR deprecates loading kv cache scales from json in favor of adding the option to dynamically compute them based on the first real input to the attention layer.
Our tests showed that the dynamic range computed based on the first input to each layer is representative of the entire model, and the accuracy is comparable with scaling factors computed using Quark quantizer (such as in HF amd/*-FP8-KV models)

Accuracy measured using the P3L benchmark that allows measuring accuracy on decode steps, using the data in the kv cache

K and V scale parameters are made on-device tensors in order to allow changing their values after the graph has been captured. This also lays the foundation to using per-channel quantization with tensor-like scales.

The effect is most visible on models with dynamic value ranges outside of the scope of fp8e4m3, such as Quen2 7B:
Using dynamic calculation reduces the PPL score from 34.84 to 22.62

On LLama based models the improvement is much smaller, due to the fact that identity scales work just as well, but still can be in single digit percents, on par with using the scales from a quantized model

…ching (#317) * Changed _k_scale and _v_scale to tensors * fixed rocm paged attention with tensor kv scales * Added on the fly scale factor calculation * trying to fix attn metadata * fixed AttentionMetadata issue, updated description for calculate-kv-scales flag in arg_utils.py * Changed K and V scale constants * Removed unneeded comment * Changes to pass format.sh, also fixed lingering k_scale/v_scale : float * Fix for TP > 1 * Ran format.sh * Removed legacy kv_scale loading from the json file * Removed the outdated kv cache docs * Revert some unwanted changes --------- Co-authored-by: Gregory Shtrasberg <[email protected]> Signed-off-by: Gregory Shtrasberg <[email protected]>

* Using tensors in the explicit cache function calls from mllama implementation * Properly creating the tensor Signed-off-by: Gregory Shtrasberg <[email protected]>

Signed-off-by: Gregory Shtrasberg <[email protected]>

…utation_upstream

Signed-off-by: Gregory Shtrasberg <[email protected]>

…utation_upstream

github-actions · 2025-01-09T18:35:49Z

👋 Hi! Thank you for contributing to the vLLM project.
Just a reminder: PRs would not trigger full CI run by default. Instead, it would only run fastcheck CI which starts running only a small and essential subset of CI tests to quickly catch errors. You can run other CI tests on top of those by going to your fastcheck build on Buildkite UI (linked in the PR checks section) and unblock them. If you do not have permission to unblock, ping simon-mo or khluu to add you in our Buildkite org.

Once the PR is approved and ready to go, your PR reviewer(s) can run CI to test the changes comprehensively before merging.

To run CI, PR reviewers can do one of these:

Add ready label to the PR
Enable auto-merge.

🚀

micah-wil and others added 10 commits December 18, 2024 13:40

Mllama kv scale fix (#335)

ef181a9

* Using tensors in the explicit cache function calls from mllama implementation * Properly creating the tensor Signed-off-by: Gregory Shtrasberg <[email protected]>

format

f9645bf

Signed-off-by: Gregory Shtrasberg <[email protected]>

Fix for different attention types

6ad050e

Signed-off-by: Gregory Shtrasberg <[email protected]>

Merge remote-tracking branch 'origin/main' into kv_cache_dynamic_comp…

64668c6

…utation_upstream

Properly initializing the new field in the attn metadata (#337)

3eaca59

Signed-off-by: Gregory Shtrasberg <[email protected]>

Upstream doesn't have fp8 navi support yet

3a18c31

Signed-off-by: Gregory Shtrasberg <[email protected]>

Cannot reference tensor contents during graph capture

390bdaa

Signed-off-by: Gregory Shtrasberg <[email protected]>

Adjusted cpu implementation datatypes

681ceb3

Signed-off-by: Gregory Shtrasberg <[email protected]>

Merge remote-tracking branch 'origin/main' into kv_cache_dynamic_comp…

9721ece

…utation_upstream

gshtras requested review from tlrmchlsmth, WoosukKwon, DarkLight1337, ywang96, robertgshaw2-neuralmagic, njhill, comaniac, alexm-neuralmagic, zhuohan123 and youkaichao as code owners January 9, 2025 18:35

mergify bot added the documentation Improvements or additions to documentation label Jan 9, 2025

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

[FP8][Kernel] Dynamic kv cache scaling factors computation #11906

[FP8][Kernel] Dynamic kv cache scaling factors computation #11906

gshtras commented Jan 9, 2025 •

edited by github-actions bot

Loading

github-actions bot commented Jan 9, 2025

[FP8][Kernel] Dynamic kv cache scaling factors computation #11906

Are you sure you want to change the base?

[FP8][Kernel] Dynamic kv cache scaling factors computation #11906

Conversation

gshtras commented Jan 9, 2025 • edited by github-actions bot Loading

github-actions bot commented Jan 9, 2025

gshtras commented Jan 9, 2025 •

edited by github-actions bot

Loading