<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Llm-Compressor on Sana Fayyaz</title><link>https://sanafayyaz315.github.io/tags/llm-compressor/</link><description>Recent content in Llm-Compressor on Sana Fayyaz</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Thu, 03 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://sanafayyaz315.github.io/tags/llm-compressor/index.xml" rel="self" type="application/rss+xml"/><item><title>Half the Size, Better Performance, Same Accuracy: Understanding W8A8 INT8 LLM Quantization</title><link>https://sanafayyaz315.github.io/posts/understanding-w8a8-int8-llm-quantization/</link><pubDate>Thu, 03 Sep 2026 00:00:00 +0000</pubDate><guid>https://sanafayyaz315.github.io/posts/understanding-w8a8-int8-llm-quantization/</guid><description>Compressing Llama 3.1 8B with SmoothQuant + GPTQ (W8A8 INT8): what changes under the hood, and what it costs and buys you in accuracy and serving performance.</description></item></channel></rss>