Image21-INT8 / CHANGES.md
ixim's picture
Fix INT8 CPU offload; publish memory audit and paired retests
ba48d54 verified
|
Raw History Blame Contribute Delete
1.18 kB

Modifications

Modified by ixim / iximbox: eligible linear weights converted from Qwen-Image-2.1 to bitsandbytes LLM.int8 INT8. Built with Qwen. Non-commercial research/evaluation under the accompanying Qwen Research License.

  • Converted eligible transformer and text encoder linear layers to LLM.int8, threshold 6.0.
  • Retained sensitive projections, normalization, embeddings, vision model and VAE in floating point.
  • Re-serialized component weights, shard indexes and configs; no fine-tuning.
  • Weight-file headers and modified JSON files contain modification notices.
  • Exact quantized module names and dtype counts are in the component reports.

Runtime and evaluation correction, 2026-09-21

  • Added an instance-local bitsandbytes auxiliary-tensor offload workaround and sequential component loading in scripts/runtime.py.
  • Re-ran the paired evaluation in isolated processes with startup, loading and per-image memory records.
  • Preserved the initial evaluation and failed 1024px editing outputs; added controlled editing diagnostics and a paired 2048px retest with VAE tiling.
  • Model weights, quantization configuration and inference-file fingerprints are unchanged.