llrehf

community
Activity Feed

AI & ML interests

None defined yet.

merveย 
posted an update about 12 hours ago
view post
Post
196
large AI labs have dropped so many open models last week ๐Ÿ”ฅ don't miss out on them

โ†’ Apple released on-device vision LMs apple/fastvlm-68ac97b9cd5cacefdd04872e & apple/mobileclip2-68ac947dcb035c54bcd20c47
โ†’ OpenGVLab released InternVL3.5, 32 new vision LMs with one based on gpt-oss! (OS) OpenGVLab/internvl35-68ac87bd52ebe953485927fb
โ†’ MSFT released a killer small TTS model (OS) microsoft/VibeVoice-1.5B

find more herehttps://huggingface.co/collections/merve/august-29-releases-68b5a3754cfb8abf59e2b486
sergiopaniegoย 
posted an update 6 days ago
view post
Post
320
It's now posible to do end-2-end ML without leaving the @huggingface Hub, by combining TRL + HF jobs + Trackio!!

๐ŸกWe just released a full guide explaining the process.

Go check it out!

๐Ÿ“– Guide: https://huggingface.co/docs/trl/main/en/jobs_training

๐Ÿ’ก Reminder: HF Jobs is only available for Pro, Team, or Enterprise plans. Yet another reason to upgrade
merveย 
posted an update 7 days ago
view post
Post
5783
first vision language model built off openai/gpt-oss-20b just dropped! ๐Ÿ”ฅ

InternVL3.5 comes with 32 models ๐Ÿคฏ pre-trained, fine-tuned, aligned in various sizes OpenGVLab/internvl35-68ac87bd52ebe953485927fb
comes with gpt-oss or Qwen3 for LLM part โคต๏ธ
  • 1 reply
ยท
sergiopaniegoย 
posted an update 20 days ago
sergiopaniegoย 
posted an update 21 days ago
view post
Post
388
New Zero-Shot Object Detectors in transformers! ๐Ÿฅฝ

Weโ€™ve added LLMDet and MM GroundingDINO, plus a demo Space to compare them with others ๐Ÿ–ผ๏ธ

Play with it: ariG23498/zero-shot-od
sergiopaniegoย 
posted an update 22 days ago
sergiopaniegoย 
posted an update 26 days ago
view post
Post
442
Latest TRL release brings major upgrades for multimodal alignment!

We dive into 3 new techniques to improve VLM post-training in our new blog:

๐ŸŒ‹ GRPO
๐ŸŽž๏ธ GSPO
๐Ÿ™ MPO
โž• vLLM integration for online training w/ transformers backend\

๐Ÿก Blog: https://huggingface.co/blog/trl-vlm-alignment
merveย 
posted an update 26 days ago
view post
Post
3229
GPT-4.1-mini level model right in your iPhone ๐Ÿคฏ

openbmb/MiniCPM-V-4 is only 4B while surpassing GPT-4.1-mini in vision benchmarks ๐Ÿ”ฅ

allows commercial use as well!
sergiopaniegoย 
posted an update 27 days ago
merveย 
posted an update 28 days ago
view post
Post
1118
we're all sleeping on this OCR model rednote-hilab/dots.ocr ๐Ÿ”ฅ

dots.ocr is a new 3B model with sota performance, support for 100 languages & allowing commercial use! ๐Ÿคฏ

single e2e model to extract image, convert tables, formula, and more into markdown ๐Ÿ“
try it MohamedRashad/Dots-OCR
sergiopaniegoย 
posted an update 28 days ago
view post
Post
3404
Want to learn how to align a Vision Language Model (VLM) for reasoning using GRPO and TRL? ๐ŸŒ‹

๐Ÿง‘โ€๐Ÿณ We've got you covered!!

NEW multimodal post training recipe to align a VLM using TRL in @HuggingFace 's Cookbook.

Go to the recipe ๐Ÿ‘‰https://huggingface.co/learn/cookbook/fine_tuning_vlm_grpo_trl

Powered by the latest TRL v0.20 release, this recipe shows how to teach Qwen2.5-VL-3B-Instruct to reason over images ๐ŸŒ‹
merveย 
posted an update 28 days ago
view post
Post
654
massive releases and tons of Flux 1. Krea LoRas past week!
here's some of the picks, find more models in collection ๐Ÿซก merve/releases-august-2-6890c14248203522b7d0267f

LLMs ๐Ÿ’ฌ
> Tencent dropped tencent/Hunyuan-7B-Instruct
> Qwen released Qwen/Qwen3-Coder-30B-A3B-Instruct, 30B MoE with 3B params for coding (OS)

vision/multimodal
> RedNote released rednote-hilab/dots.ocr - 3B OCR model (OS)
> Cohere released CohereLabs/command-a-vision-07-2025 - 112B (dense!) VLM for 6 languages
> StepFun-AI shipped stepfun-ai/step3 - 321B MoE VLM (OS)
> Skywork shipped Skywork/Skywork-UniPic-1.5B - new any-to-any model (image+text โ†’ image+text) (OS)
sergiopaniegoย 
posted an update 29 days ago
view post
Post
4500
Just included example scripts for aligning models using GSPO (including VLM example) ๐Ÿ™†โ€โ™‚๏ธ๐Ÿ™†โ€โ™‚๏ธ

GSPO is the latest RL alignment algo by @Alibaba_Qwen and it's already supported in the latest TRL v0.20 release.

Super-easy-to-get-started example scripts below, GO run them!๐Ÿ‘ฉโ€๐Ÿ’ป๐Ÿ‘ฉโ€๐Ÿ’ป

๐Ÿง‘โ€๐ŸŽจ Script: https://github.com/huggingface/trl/blob/main/examples/scripts/gspo.py
๐Ÿฆ„ VLM script: https://github.com/huggingface/trl/blob/main/examples/scripts/gspo_vlm.py
๐Ÿงฉ More TRL examples: https://huggingface.co/docs/trl/main/en/example_overview
๐Ÿง™โ€โ™‚๏ธ GSPO paper: Group Sequence Policy Optimization (2507.18071)
merveย 
posted an update about 1 month ago
sergiopaniegoย 
posted an update about 1 month ago
view post
Post
342
Did you miss this? ๐Ÿ‘“

๐Ÿง™โ€โ™‚๏ธvLLM + transformers integration just got upgraded with direct VLM support.

Select a VLM + model_impl=transformers and play via vLLM!
merveย 
posted an update about 1 month ago
view post
Post
3611
past week in open AI was insane ๐Ÿ”ฅ here's some of picks, find more here merve/releases-july-25-688768ca47fe3693407e02d1

๐Ÿ’ฌ LLMs & VLMs
> Qwen/Qwen3-235B-A22B-Thinking-2507 had a new update (OS)
> Qwen/Qwen3-Coder-480B-A35B-Instruct is out with 480B total 35B active params ๐Ÿคฏ (OS)
> AllenAI dropped an update to allenai/olmOCR-7B-0725 ๐Ÿ“
> InternLM released internlm/Intern-S1 - 235B Qwen3 MoE + 6B InternViT encoder (OS)
> OmniSVG/OmniSVG is a new SVG generation VLM (OS)

๐Ÿ–ผ๏ธ image/video/3D generation
> WanAI released Wan2.2 series - both T2V and I2V 14B models for high-quality video generation (OS) multimodalart/wan-22-688767e313337b434ed55112
> Tencent dropped tencent/HunyuanWorld-1 - image-to-3D scene generation model
  • 1 reply
ยท
sergiopaniegoย 
posted an update about 1 month ago
view post
Post
2661
We just released TRL v0.20 with major multimodal upgrades!

๐Ÿ‘๏ธ VLM support for GRPO (highly requested by the community!)
๐ŸŽž๏ธ New GSPO trainer (from @Qwen , released last week, VLM-ready)
๐Ÿ™ New MPO trainer (multimodal by design, as in the paper)

๐Ÿ“ Full release notes here: https://github.com/huggingface/trl/releases/tag/v0.20.0
merveย 
posted an update about 1 month ago
view post
Post
4366
๐Ÿคฏ 241B VLM with apache-2.0 license internlm/Intern-S1

internlm released Intern-S1: multimodal reasoning model based on 235B MoE Qwen3 and 6B InternViT ๐Ÿ˜

benchmarks look great (๐Ÿ‘‘ best model โœ… best open model)
sergiopaniegoย 
posted an update about 1 month ago
view post
Post
1203
Yet Another New Multimodal Fine-Tuning Recipe ๐Ÿฅง

๐Ÿง‘โ€๐Ÿณ In this @HuggingFace Face Cookbook notebook, we demonstrate how to align a multimodal model (VLM) using Mixed Preference Optimization (MPO) using trl.

๐Ÿ’ก This recipe is powered by the new MPO support in trl, enabled through a recent upgrade to the DPO trainer!

We align the multimodal model using multiple optimization objectives (losses), guided by a preference dataset (chosen vs. rejected multimodal pairs).

Check it out! โžก๏ธ https://huggingface.co/learn/cookbook/fine_tuning_vlm_mpo
  • 2 replies
ยท
merveย 
posted an update about 1 month ago
view post
Post
816
so many open LLMs and image LoRAs dropped past week, here's some picks for you ๐Ÿซก merve/releases-july-18-687e3fbd2ab9b39c51f9238b

LLMs
> ByteDance released a bunch of translation models called Seed-X-RM (7B) ByteDance-Seed/Seed-X-RM-7B
> NVIDIA released reasoning models of which 32B surpassing the giant Qwen3-235B with cc-by-4.0 license ๐Ÿ‘ nvidia/openreasoning-nemotron-687730dae0170059860f1f01
> LG released a new EXAONE model (32B) LGAI-EXAONE/EXAONE-4.0-32B

VLMs/any-to-any
> vidore/colqwen-omni-v0.1 is a new any-to-any retriever (MIT)
> HiDream-ai/HiDream-E1-1 is image+text in image+text out model (MIT)

LoRAs
> There's a bunch of LoRAs based on Flux Kontext, gotta check out the collection ๐Ÿค