Refactor QAT to use common fake_quantize_affine primitive #527

andrewor14 · 2024-07-19T15:12:20Z

Summary: Currently there are two QAT quantizers, 8da4w and 4w. Today, these use different autograd functions to represent their fake quantization numerics, but this is not scalable because new QAT quantizers may introduce yet another divergent code path. To address this, this commit refactors both quantizers to use the common fake_quantize_affine QAT primitive.

Test Plan:
python test/quantization/test_qat.py

Reviewers: jerryzh168

Subscribers: jerryzh168, supriyar, msaroufim

Summary: Currently there are two QAT quantizers, 8da4w and 4w. Today, these use different autograd functions to represent their fake quantization numerics, but this is not scalable because new QAT quantizers may introduce yet another divergent code path. To address this, this commit refactors both quantizers to use the common fake_quantize_affine QAT primitive. Test Plan: python test/quantization/test_qat.py Reviewers: jerryzh168 Subscribers: jerryzh168, supriyar, msaroufim

pytorch-bot · 2024-07-19T15:12:23Z

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/ao/527

📄 Preview Python docs built from this PR

Note: Links to docs will display an error until the docs builds have been completed.

✅ No Failures

As of commit 8486207 with merge base 6dd82d8 ():
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

jerryzh168 · 2024-07-20T00:00:40Z

torchao/quantization/prototype/qat.py

@@ -25,7 +25,10 @@
    ZeroPointDomain,
 )
 from torchao.quantization.unified import TwoStepQuantizer
-from torchao.quantization.utils import get_group_qparams_symmetric
+from torchao.quantization.utils import (
+    _get_per_token_block_size,


if it's helpful we could have a general util like:

def get_block_size(granularity, **kw_params) -> Callable: if granularity == Granularity.PER_BLOCK: ... elif type == Granularity.PER_TOKEN: ... ... block_size = get_block_size(Granularity.PER_TOKEN)(x)

Sounds good, let's do that separately

Summary: Currently there are two QAT quantizers, 8da4w and 4w. Today, these use different autograd functions to represent their fake quantization numerics, but this is not scalable because new QAT quantizers may introduce yet another divergent code path. To address this, this commit refactors both quantizers to use the common fake_quantize_affine QAT primitive. Test Plan: python test/quantization/test_qat.py Reviewers: jerryzh168 Subscribers: jerryzh168, supriyar, msaroufim

andrewor14 requested a review from jerryzh168 July 19, 2024 15:12

facebook-github-bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jul 19, 2024

jerryzh168 reviewed Jul 20, 2024

View reviewed changes

jerryzh168 approved these changes Jul 20, 2024

View reviewed changes

andrewor14 merged commit 5787e9e into main Jul 22, 2024
13 checks passed

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Refactor QAT to use common fake_quantize_affine primitive #527

Refactor QAT to use common fake_quantize_affine primitive #527

andrewor14 commented Jul 19, 2024

pytorch-bot bot commented Jul 19, 2024 •

edited

Loading

jerryzh168 Jul 20, 2024 •

edited

Loading

andrewor14 Jul 22, 2024

Refactor QAT to use common fake_quantize_affine primitive #527

Refactor QAT to use common fake_quantize_affine primitive #527

Conversation

andrewor14 commented Jul 19, 2024

pytorch-bot bot commented Jul 19, 2024 • edited Loading

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/ao/527

✅ No Failures

jerryzh168 Jul 20, 2024 • edited Loading

Choose a reason for hiding this comment

andrewor14 Jul 22, 2024

Choose a reason for hiding this comment

pytorch-bot bot commented Jul 19, 2024 •

edited

Loading

jerryzh168 Jul 20, 2024 •

edited

Loading