Why 2.6 models performs worse in your own coding bech than older 2.5 versions?

#9
by sarkaritamminen - opened

MiMo Coding Bench

All
Open
Proprietary
4 models
#	Model	Score	Size	Context	Cost	License
1	
Xiaomi
MiMo-V2.5-Pro
Xiaomi
0.737	1.0T	1.0M	$0.43 / $0.87	
2	
Xiaomi
MiMo-V2.5
Xiaomi
0.718	311B	1.0M	$0.17 / $0.34	
3	
Xiaomi
MiMo-V2.6-Pro
New
Xiaomi
0.632	1.0T	1.0M	$0.43 / $0.87	
4	
Xiaomi
MiMo-V2.6-Flash
New
Xiaomi
0.612	309B	1.0M	$0.14 / $0.28	

https://llm-stats.com/benchmarks/mimo-coding-bench

Care to elaborate?

This model performs very poorly in practice; it is all show and no substance. When I tested Qwen3.8-next-flash, deepseek-v4-flash-0731 and this model on the same task, Xiaomi’s model failed miserably. It is a model that is all show and no substance, not worth using, and it also suffers from infinite loops.

Sign up or log in to comment