AI & ML interests

Tools for creating and exploring datasets

Recent Activity

prithivMLmodsย 
posted an update 4 days ago
view post
Post
3339
OneDecision-VisionGuard-Demo is now available on Hugging Face Spaces!

๐Ÿค— Space: prithivMLmods/OneDecision-VisionGuard-Demo

This demo showcases the OneDecision-VisionGuard family of multimodal image classification models for detecting NSFW and other sensitive visual content, with structured JSON reasoning, improved accuracy, and better handling of edge cases such as sensitive imagery, uncensored analysis, scene descriptions, and classification reasoning.

๐Ÿ“ฆ Models: 27B, 9B, 4B โ€” prithivMLmods/OneDecision-VisionGuard-27B-SFT, prithivMLmods/OneDecision-VisionGuard-9B-SFT, prithivMLmods/OneDecision-VisionGuard-4B-SFT

โ†—๏ธ Collection: https://huggingface.co/collections/prithivMLmods/onedecision-visionguard

To learn more, visit the app page or the respective model pages.
alielfilali01ย 
posted an update 10 days ago
view post
Post
407
multi-agents orchestration is old news by now. Try multi-teams of agents ! that is a whole different nightmare ... This is the real signal about the sparks of AGI.
prithivMLmodsย 
posted an update 17 days ago
view post
Post
3911
Qwen-Image-2.1 Plug and Play LoRA App is now live on Hugging Face Spaces.

๐Ÿ”— Space: prithivMLmods/Qwen-Image-2.1-LoRAs-PnP

It supports standard inference, 4-step Turbo inference, custom LoRA lazy repacks, and LoRA Plug and Play (PnP), all in one setting!

๐Ÿ”— Qwen-Image-2.1 Image-to-Image LoRAs: https://huggingface.co/collections/prithivMLmods/qwen-image-21-image-to-image-loras

๐Ÿ”— GitHub: https://github.com/PRITHIVSAKTHIUR/Qwen-Image-2.1-LoRAs-PnP

To learn more, visit the app page or the respective model pages.
prithivMLmodsย 
posted an update 25 days ago
view post
Post
844
VisionGuardrail EVO-2, a multimodal image-classification content-safety model based on Qwen/Qwen3.8-27B, is now available on the Hub!

Stricter image classification than before, with a dense 27-billion-parameter multimodal model, more precise reasoning, and improved captions for classifying visual media.

โž  Models: prithivMLmods/VisionGuardrail-Evo2-27B, prithivMLmods/VisionGuardrail-Evo2-27B-GGUF

โž  Collection: https://huggingface.co/collections/prithivMLmods/visionguardrail-evo2

โž  Previous Models: https://huggingface.co/collections/prithivMLmods/visionguardrail-collection

โคท To learn more, visit the app page or the respective model pages.
prithivMLmodsย 
posted an update 28 days ago
view post
Post
492
Scribble-Board-Fast is a sketch-to-image workspace powered by Klein-9B, transforming doodles, brush strokes, stickers, and uploaded images into high-fidelity visuals with 4-step distilled sampling.

> Space: prithivMLmods/Scribble-Board-Fast
> GitHub: https://github.com/PRITHIVSAKTHIUR/Scribble-Board-Fast

> To learn more, visit the app page or the respective model pages.
prithivMLmodsย 
posted an update about 1 month ago
view post
Post
3872
VisionGuardrail, a multimodal content-safety classifier based on Qwen3.5, is now available on Hugging Face in 4B and 9B variants. It is a direct upgrade to ImageShield-MMCF, providing improved parental controls through conservative visual content-safety filtering.

More About:
โž  hf.co/blog โ€” https://huggingface.co/blog/prithivMLmods/vision-guardrail-mini-blog

โž  Models:
โœฆ VisionGuardrail-4B: prithivMLmods/VisionGuardrail-4B
โœฆ VisionGuardrail-9B: prithivMLmods/VisionGuardrail-9B

โž  Dataset:
โœฆ ImageShield-Guardrail-Pro: prithivMLmods/ImageShield-Guardrail-Pro

โคท To learn more, visit the app page or the respective model pages.
prithivMLmodsย 
posted an update about 1 month ago
view post
Post
3109
ImageShield-MMCF โ€” Multimodal Content Filter is a multimodal content-safety classifier built on top of Qwen3.5 and is now available on Hugging Face!

This is the preview initial version (v1.0) of the model, designed to classify visual content as Safe or Unsafe, with a particular focus on detecting Not Safe for Work (NSFW) and other potentially sensitive visual content.

The demo is implemented in the prithivMLmods/opencaption-4b-vl-sft Space, which serves as an active content-safety layer for computer vision tasks. It helps block Not Safe for Work (NSFW) content generation and paves the way for more meaningful and responsible creativity.

โŠน ImageShield-MMCF-0.8B: prithivMLmods/ImageShield-MMCF-0.8B
โŠน ImageShield-MMCF-2B: prithivMLmods/ImageShield-MMCF-2B
  • 2 replies
ยท
prithivMLmodsย 
posted an update about 2 months ago
view post
Post
5295
The Qwen3.8 27B demo for object grounding is now available on Hugging Face Spaces.

It features three tasks: Object Detection (Bounding Boxes), Point Localization (Keypoints), and Spatial Guidance (Path Mapping).

Try it now: prithivMLmods/Qwen3.8-27B-Object-Detection
prithivMLmodsย 
posted an update 2 months ago
view post
Post
5575
Made a demo for Text/Image-to-3D Video and Image-to-3D Video asset generation using TRELLIS.2. It is paired with Z-Image-Turbo to accelerate the input image preprocessing pipeline, streamlining the Image-to-3D workflow. The generated GLB (GL Transmission Format) files are converted into MP4 (MPEG-4) videos, making them easy to preview and share. Try it now on Hugging Face Spaces.๐Ÿค—

โž  Image-to-3D-Video-Asset-Generator: prithivMLmods/Image-to-3D-Video-Asset-Generator
โž  collection: https://huggingface.co/collections/prithivMLmods/multimodal-implementations
โž  github: https://github.com/PRITHIVSAKTHIUR/Image-to-3D-Video-Asset-Generator

โคท To learn more, visit the app page or the respective model pages.
mmhamdyย 
posted an update 3 months ago
view post
Post
237
Decades before the modern scaling laws, this paper showed that neural networks behavior under scale follows remarkably predictable laws.

In 1993, researchers at Bell Labs were grappling with a constraint that feels entirely familiar (and contemporary): datasets were outgrowing the available hardware, and training a model to the end was becoming too expensive. To evaluate an architectural tweak to a state-of-the-art model (at the time it was LeNet) on 60,000 samples meant burning up to three weeks of compute time.

To save compute, people would train candidate architectures on small subsets of the data, assuming that the top performer at small scale would remain the top performer at full scale. But with our future wisdom, we know this is not the case.

In "Learning Curves: Asymptotic Values and Rate of Convergence (NeurIPS 93)", using insights from statistical mechanics, they proposed a practical and principled method for predicting the performance of classifiers trained on large datasets (at the time, models were assumed to be large enough). The method was based on a simple power-law modeling of the expected training and test errors.

It is often noted that many of today's breakthroughs in AI and deep learning are actually decades-old concepts that simply lacked the computational power to be tested at the time. While there is some truth to that, it highlights a more valuable lesson: there is immense worth in revisiting early literature and reflecting on foundational ideas we may have prematurely left behind.

So, go explore and find your own inspiration. The current trend has enough champions already!
  • 2 replies
ยท
eienmojikiย 
posted an update 4 months ago
mmhamdyย 
posted an update 4 months ago
view post
Post
339
It has been more than a decade now since the knowledge distillation paper came out.

Knowledge Distillation (KD) is one of my favorite topics, but I have to confess that I'm not a huge fan of the term because I find it confusing (or at least, it has became so over time).

The idea behind KD is not novel; it was there almost a decade before the paper came out (and arguably even a decade before that, back to 1990-91). But this paper is the one that clicked, the one that made the topic much more popular and introduced it to a broader audience.

First, the timing and the authors played a big role: we have Geoffrey Hinton, Oriol Vinyals, and Jeff Dean here. And second, Geoffrey Hinton is really good at idea branding: Model compression?! No, no, no! Let's call it "Knowledge Distillation" and use evocative terms such as "Dark Knowledge" to describe what is being transferred.

It's a great name, but as time has passed, the term became a bit of a relic. KD is no longer solely about compression (KD used to be introduced as a method for model compression, but now model compression is just one application of KD). And the other thing is that the word "distillation" implies some sort of potency here, that the student is somehow more powerful than the teacher, which is not the case (but many counterarguments could be made, for example, more powerful compared to another model trained with no teacher)

Nevertheless, the paper is incredibly well-written, short, and fun to read. It's one of few papers that I read several times. Check it out, and maybe share your thoughts on the topic with us here!

If you had to choose another name for Knowledge Distillation, what would it be?

  • 6 replies
ยท
Abhaykoulย 
posted an update 4 months ago
view post
Post
491
Shipped v0.1.2 of vtx โ€” a minimalist coding agent for the terminal.

Most agentic CLIs ship 10k+ token system prompts. Vtx is ~2,200. Less prompt overhead means more room for your code in the model's context window.

Vtx is a from-scratch Python implementation of the design philosophy behind pi-mono โ€” same principles, pure Python, no transpiled runtime.

What ships out of the box:

โ†’ Textual TUI + headless CLI (vtx -p "fix the failing test")
โ†’ 49 LLM provider gateways, all declared in a single provider.yaml
โ†’ 5 core tools (read / edit / write / bash / find) plus web search and fetch
โ†’ Session tree with compaction, handoff, and resume
โ†’ AGENTS.md / CLAUDE.md auto-discovery
โ†’ Skills system โ€” drop SKILL.md files in .agents/skills/ and they become slash commands
โ†’ Two OAuth flows (GitHub Copilot device flow, OpenAI Codex PKCE)
โ†’ Two-mode permissions: prompt (default) or auto, with a safe-command allowlist

This release adds a proper extension system. Register new LLM-callable tools, intercept tool calls, hook lifecycle events, and add slash commands from a single register(api) function in a Python file under ~/.vtx/agent/extensions/. Extensions can override built-in tools by name and chain handler logic across subscribers.

Apache 2.0. uv tool install vtx-coding-agent and you're running.

GitHub: https://github.com/OEvortex/vtx-coding-agent
PyPI: https://pypi.org/project/vtx-coding-agent

Built in the open. Feedback, extensions, and PRs welcome.
mmhamdyย 
posted an update 4 months ago
view post
Post
2700
What if you could train a model on just 10 images instead of 60,000 and still get close to the same performance?

Traditional machine learning requires thousands, even millions, of data points to achieve high accuracy. But what if we could "distill" the entire dataset into just a few synthetic samples?

This is what Dataset Distillation offers. Unlike traditional knowledge distillation, we keep the model fixed and distill the knowledge contained in a massive training set into a tiny set of synthetic distilled images.

The goal is to train a model on this ultra-small set and achieve performance that almost matches what the same model would get when trained on the massive original dataset.

For example, training on only 10 distilled MNIST images (this is equivalent to a single image per class) yields 94% accuracy, compared to 99% when training on the full 60,000 images.

Interestingly, these distilled images look significantly different (as you can see in the image below) from natural images because they are optimized for model training rather than for matching the correct data distribution.

But that's not all.

Most importantly, this same method opens the door to a potent form of data poisoning. Because distilled images are specifically optimized for rapid learning, an attacker can create a tiny set of adversarial distilled images to cause a well-trained model to forget or misclassify a specific category.

What I find fascinating about dataset distillation is this: it mimics human-like learning by letting a model grasp a concept from a single example, but it does so using alien synthetic images that mean absolutely nothing to a human eye!

What about you? What are your thoughts on it?
  • 2 replies
ยท