File size: 4,528 Bytes
2855f1e
 
 
 
 
 
 
 
 
 
 
c861fb5
 
 
 
 
 
 
 
 
 
 
2855f1e
 
 
2093b08
2855f1e
2093b08
2855f1e
2093b08
 
 
2855f1e
2093b08
2855f1e
2093b08
2855f1e
2093b08
 
 
 
 
 
 
2855f1e
 
 
2093b08
 
 
 
 
 
2855f1e
2093b08
2855f1e
2093b08
 
 
2855f1e
2093b08
 
 
 
 
2855f1e
2093b08
 
 
 
 
 
 
 
 
 
2855f1e
 
 
2093b08
 
 
2855f1e
 
2093b08
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2855f1e
2093b08
c861fb5
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
cfe2d11
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
---
language:
- en
license: apache-2.0
library_name: gguf
tags:
- ruvltra
- sona
- adaptive-learning
- gguf
- quantized
- turboquant
- kv-cache-compression
- flash-attention
- speculative-decoding
- graph-rag
- hybrid-search
- vector-database
- ruvector
- diskann
- mamba-ssm
- colbert
pipeline_tag: text-generation
---

<div align="center">

# RuvLTRA Medium

[![License](https://img.shields.io/badge/License-Apache%202.0-blue.svg)](https://opensource.org/licenses/Apache-2.0)
[![HuggingFace](https://img.shields.io/badge/πŸ€—%20Hugging%20Face-Model-yellow)](https://huggingface.co/ruv/ruvltra-medium)
[![GGUF](https://img.shields.io/badge/Format-GGUF-green)](https://github.com/ggerganov/ggml/blob/master/docs/gguf.md)

**βš–οΈ Balanced Model for General-Purpose Tasks**

</div>

---

## Overview

RuvLTRA Medium provides the sweet spot between capability and resource usage. Ideal for desktop applications, development workstations, and moderate-scale deployments.

## Model Card

| Property | Value |
|----------|-------|
| **Parameters** | 1.1 Billion |
| **Quantization** | Q4_K_M |
| **Context** | 8,192 tokens |
| **Size** | ~669 MB |
| **Min RAM** | 2 GB |
| **Recommended RAM** | 4 GB |

## πŸš€ Quick Start

```bash
# Download
wget https://huggingface.co/ruv/ruvltra-medium/resolve/main/ruvltra-1.1b-q4_k_m.gguf

# Run inference
./llama-cli -m ruvltra-1.1b-q4_k_m.gguf \
  -p "Explain quantum computing in simple terms:" \
  -n 512 -c 8192
```

## πŸ’‘ Use Cases

- **Development**: Code assistance and generation
- **Writing**: Content creation and editing
- **Analysis**: Document summarization
- **Chat**: Conversational AI applications

## πŸ”§ Integration

### Rust
```rust
use ruvllm::hub::ModelDownloader;

let path = ModelDownloader::new()
    .download("ruv/ruvltra-medium", None)
    .await?;
```

### Python
```python
from llama_cpp import Llama
from huggingface_hub import hf_hub_download

model_path = hf_hub_download("ruv/ruvltra-medium", "ruvltra-1.1b-q4_k_m.gguf")
llm = Llama(model_path=model_path, n_ctx=8192)
```

### OpenAI-Compatible Server

```bash
python -m llama_cpp.server \
  --model ruvltra-1.1b-q4_k_m.gguf \
  --host 0.0.0.0 --port 8000
```

## Performance

| Platform | Tokens/sec |
|----------|------------|
| M2 Pro (Metal) | 65 tok/s |
| RTX 4080 (CUDA) | 95 tok/s |
| i9-13900K (CPU) | 25 tok/s |

---

**License**: Apache 2.0 | **GitHub**: [ruvnet/ruvector](https://github.com/ruvnet/ruvector)


---

## ⚑ TurboQuant KV-Cache Compression

RuvLTRA models are fully compatible with **TurboQuant** β€” 2-4 bit KV-cache quantization that reduces inference memory by 6-8x with <0.5% quality loss.

| Quantization | Compression | Quality Loss | Best For |
|-------------|-------------|--------------|----------|
| 3-bit | 10.7x | <1% | **Recommended** β€” best balance |
| 4-bit | 8x | <0.5% | High quality, long context |
| 2-bit | 32x | ~2% | Edge devices, max savings |

### Usage with RuvLLM

```bash
cargo add ruvllm    # Rust
npm install @ruvector/ruvllm   # Node.js
```

```rust
use ruvllm::quantize::turbo_quant::{TurboQuantCompressor, TurboQuantConfig, TurboQuantBits};

let config = TurboQuantConfig {
    bits: TurboQuantBits::Bit3_5, // 10.7x compression
    use_qjl: true,
    ..Default::default()
};
let compressor = TurboQuantCompressor::new(config)?;
let compressed = compressor.compress_batch(&kv_vectors)?;
let scores = compressor.inner_product_batch_optimized(&query, &compressed)?;
```

### v2.1.0 Ecosystem

- **Hybrid Search** β€” Sparse + dense vectors with RRF fusion (20-49% better retrieval)
- **Graph RAG** β€” Knowledge graph + community detection for multi-hop queries
- **DiskANN** β€” Billion-scale SSD-backed ANN with <10ms latency
- **FlashAttention-3** β€” IO-aware tiled attention, O(N) memory
- **MLA** β€” Multi-Head Latent Attention (~93% KV-cache compression)
- **Mamba SSM** β€” Linear-time selective state space models
- **Speculative Decoding** β€” 2-3x generation speedup

[RuVector GitHub](https://github.com/ruvnet/ruvector) | [ruvllm crate](https://crates.io/crates/ruvllm) | [@ruvector/ruvllm npm](https://www.npmjs.com/package/@ruvector/ruvllm)


---

## Benchmarks (L4 GPU, 24GB VRAM)

| Metric | Result |
|--------|--------|
| **Inference Speed** | 62.6 tok/s |
| **Model Load Time** | 1.1s |
| **Parameters** | 3B |
| **TurboQuant KV (3-bit)** | 10.7x compression, <1% PPL loss |
| **TurboQuant KV (4-bit)** | 8x compression, <0.5% PPL loss |

*Benchmarked on Google Cloud L4 GPU via `ruvltra-calibration` Cloud Run Job (2026-03-28)*