Qwen-3.5-4B-ASVD-Healed

Qwen-3.5-4B-ASVD-Healed is a structurally compressed and optimized iteration of the base Qwen/Qwen3.5-4B model. It is the result of an advanced post-training neural surgery pipeline combining Activation-Aware Singular Value Decomposition (ASVD) with a rigorous recovery phase via LoRA Healing (Pruning-Aware Fine-Tuning), operating natively in bfloat16 precision.

Architecture Overview

Traditional large language model (LLM) compression often relies on static quantization (such as INT4 or INT8), which inherently introduces dequantization latency during inference. This approach instead focuses on permanently eliminating matrix-level redundancies—removing structural noise without requiring external decoders, custom kernels, or specialized runtimes.

The surgical process was executed on high-performance infrastructure (NVIDIA A40 48GB VRAM), utilizing the Wikitext-2 dataset for long-context calibration to maximize the retention of logical reasoning and routing capabilities.

The Compression Pipeline

  1. Activation Profiling (RMS Scaling): The base model was calibrated using 2048-token batches from Wikitext-2 to extract the Root Mean Square (RMS) of activations across critical linear layers (q_proj, k_proj, v_proj, o_proj, down_proj). These values acted as scaling factors to protect essential routing pathways and penalize inactive weights.
  2. ASVD Guided Amputation: SVD factorization was applied directly on the GPU, enforcing a strict 85% variance retention threshold ($W \approx U_k \Sigma_k V_k^T$). The layers were restructured into bottleneck sequential layers, resulting in a significant absolute reduction of total parameters.
  3. Neural Healing (LoRA): To reverse the aphasia and syntactic degradation caused by aggressive pruning, LoRA adapters (Rank 32, Alpha 64) were injected into all linear layers. The model underwent micro-batch gradient accumulation fine-tuning to rebuild severed neural bridges. Finally, the adapters were permanently merged back into the base weights (merge and unload), outputting a pure, standalone bfloat16 artifact.

Technical Specifications

  • Base Model: Qwen/Qwen3.5-4B
  • Precision: bfloat16 (Native)
  • Target Layers: q_proj, k_proj, v_proj, o_proj, down_proj
  • Variance Threshold: 85%
  • LoRA Configuration: Rank 32, Alpha 64, Dropout 0.05
  • Calibration Dataset: Wikitext-2 (2048-token context)

Usage

The model is fully compatible with the standard Hugging Face ecosystem and functions exactly like any causal language model. No custom inference code is required.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "Kodjaoglanian/Qwen-3.5-4B-ASVD-Healed"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id, 
    torch_dtype=torch.bfloat16, 
    device_map="auto"
)

prompt = "The most important concept in quantum physics is"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

with torch.no_grad():
    outputs = model.generate(
        **inputs, 
        max_new_tokens=60, 
        do_sample=True, 
        temperature=0.7
    )

print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Advantages and Limitations

Operational Advantages:

  • Native Inference: Zero runtime dequantization overhead. It works out-of-the-box with standard inference engines, vLLM, and pipelines.
  • VRAM Efficiency: The absolute reduction in matrix parameters frees up valuable VRAM, enabling larger context windows and denser batch sizes in edge AI or memory-constrained environments.

Limitations:

  • While general syntax, grammar, and logical reasoning were fully restored by the LoRA healing phase, the 85% variance retention threshold means that highly esoteric or niche factual recall may exhibit slight degradation compared to the uncompressed 4B base model. For production use cases demanding strict factual accuracy in specialized domains, downstream fine-tuning of this compressed artifact is advised.
Downloads last month
992
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Kodjaoglanian/Qwen-3.5-4B-ASVD-Healed

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(903)
this model
Quantizations
1 model

Collection including Kodjaoglanian/Qwen-3.5-4B-ASVD-Healed