Instructions to use Kodjaoglanian/Qwen-3.5-4B-ASVD-Healed with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Kodjaoglanian/Qwen-3.5-4B-ASVD-Healed with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Kodjaoglanian/Qwen-3.5-4B-ASVD-Healed") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Kodjaoglanian/Qwen-3.5-4B-ASVD-Healed") model = AutoModelForCausalLM.from_pretrained("Kodjaoglanian/Qwen-3.5-4B-ASVD-Healed", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Kodjaoglanian/Qwen-3.5-4B-ASVD-Healed with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Kodjaoglanian/Qwen-3.5-4B-ASVD-Healed" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Kodjaoglanian/Qwen-3.5-4B-ASVD-Healed", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Kodjaoglanian/Qwen-3.5-4B-ASVD-Healed
- SGLang
How to use Kodjaoglanian/Qwen-3.5-4B-ASVD-Healed with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Kodjaoglanian/Qwen-3.5-4B-ASVD-Healed" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Kodjaoglanian/Qwen-3.5-4B-ASVD-Healed", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Kodjaoglanian/Qwen-3.5-4B-ASVD-Healed" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Kodjaoglanian/Qwen-3.5-4B-ASVD-Healed", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Kodjaoglanian/Qwen-3.5-4B-ASVD-Healed with Docker Model Runner:
docker model run hf.co/Kodjaoglanian/Qwen-3.5-4B-ASVD-Healed
Qwen-3.5-4B-ASVD-Healed
Qwen-3.5-4B-ASVD-Healed is a structurally compressed and optimized iteration of the base Qwen/Qwen3.5-4B model. It is the result of an advanced post-training neural surgery pipeline combining Activation-Aware Singular Value Decomposition (ASVD) with a rigorous recovery phase via LoRA Healing (Pruning-Aware Fine-Tuning), operating natively in bfloat16 precision.
Architecture Overview
Traditional large language model (LLM) compression often relies on static quantization (such as INT4 or INT8), which inherently introduces dequantization latency during inference. This approach instead focuses on permanently eliminating matrix-level redundancies—removing structural noise without requiring external decoders, custom kernels, or specialized runtimes.
The surgical process was executed on high-performance infrastructure (NVIDIA A40 48GB VRAM), utilizing the Wikitext-2 dataset for long-context calibration to maximize the retention of logical reasoning and routing capabilities.
The Compression Pipeline
- Activation Profiling (RMS Scaling): The base model was calibrated using 2048-token batches from Wikitext-2 to extract the Root Mean Square (RMS) of activations across critical linear layers (
q_proj,k_proj,v_proj,o_proj,down_proj). These values acted as scaling factors to protect essential routing pathways and penalize inactive weights. - ASVD Guided Amputation: SVD factorization was applied directly on the GPU, enforcing a strict 85% variance retention threshold ($W \approx U_k \Sigma_k V_k^T$). The layers were restructured into bottleneck sequential layers, resulting in a significant absolute reduction of total parameters.
- Neural Healing (LoRA): To reverse the aphasia and syntactic degradation caused by aggressive pruning, LoRA adapters (Rank 32, Alpha 64) were injected into all linear layers. The model underwent micro-batch gradient accumulation fine-tuning to rebuild severed neural bridges. Finally, the adapters were permanently merged back into the base weights (merge and unload), outputting a pure, standalone
bfloat16artifact.
Technical Specifications
- Base Model:
Qwen/Qwen3.5-4B - Precision:
bfloat16(Native) - Target Layers:
q_proj,k_proj,v_proj,o_proj,down_proj - Variance Threshold: 85%
- LoRA Configuration: Rank 32, Alpha 64, Dropout 0.05
- Calibration Dataset: Wikitext-2 (2048-token context)
Usage
The model is fully compatible with the standard Hugging Face ecosystem and functions exactly like any causal language model. No custom inference code is required.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "Kodjaoglanian/Qwen-3.5-4B-ASVD-Healed"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto"
)
prompt = "The most important concept in quantum physics is"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
with torch.no_grad():
outputs = model.generate(
**inputs,
max_new_tokens=60,
do_sample=True,
temperature=0.7
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Advantages and Limitations
Operational Advantages:
- Native Inference: Zero runtime dequantization overhead. It works out-of-the-box with standard inference engines, vLLM, and pipelines.
- VRAM Efficiency: The absolute reduction in matrix parameters frees up valuable VRAM, enabling larger context windows and denser batch sizes in edge AI or memory-constrained environments.
Limitations:
- While general syntax, grammar, and logical reasoning were fully restored by the LoRA healing phase, the 85% variance retention threshold means that highly esoteric or niche factual recall may exhibit slight degradation compared to the uncompressed 4B base model. For production use cases demanding strict factual accuracy in specialized domains, downstream fine-tuning of this compressed artifact is advised.
- Downloads last month
- 992