EmbeddedLLM @EmbeddedLLM
Your open-source AI ally. We are committed to making production-grade AI inference as accessible and reliable as electricity, powered by vLLM. Joined October 2023-
Tweets530
-
Followers1K
-
Following1K
-
Likes690
Singapore has come a long way. 🇸🇬 From AI adoption to AI infrastructure, the local ecosystem is now contributing to the layers production AI depends on: @PyTorch, @vllm_project, inference, sovereign AI, and open-source infra. Proud to see @RedHat_AI, @inferact, and @EmbeddedLLM building alongside APAC AI community.
The inaugural PyTorch Meetup Singapore brought together engineers, researchers, and community builders to talk about everything from vLLM project updates to the broader question of sovereign intelligence. Read the full technical recap and find presentation slides in our latest
Incredible collaboration from the team! Beyond basic inference support, we also have day-0 speculator and RL support🔥
🎉 Congrats to @MiniMax_AI on releasing MiniMax M3! Frontier coding and agentic capabilities, native image and video input, computer use, and a 1M-token context window, all in a single open model. At the heart of M3 is MSA, a new sparse attention architecture: instead of
🎉 Congrats to @MiniMax_AI on releasing MiniMax M3! Frontier coding and agentic capabilities, native image and video input, computer use, and a 1M-token context window, all in a single open model. At the heart of M3 is MSA, a new sparse attention architecture: instead of attending densely over the full KV cache, each query scores 128-token KV blocks and runs attention only over the top blocks. That is what makes 1M-token context practical to serve. M3 runs in vLLM with day-0 support, verified on NVIDIA and AMD hardware: ✨ MSA sparse attention with dedicated prefill and decode kernels ✨ 1M-token context serving with prefix caching and chunked prefill ✨ BF16 and MXFP8 checkpoints, with MoE backends for both Hopper and Blackwell ✨ Native multimodal input (image + video) ✨ Tool calling, reasoning parsing, and thinking-mode control for agent workloads Day-0 support like this is a true team effort. Grateful to the teams at @MiniMax_AI, @NVIDIAAI, @AIatAMD, and @inferact, and to the vLLM community for making it happen. 🙏 Deep dive into the implementation, kernel work, and deployment recipes: 🔗 vllm.ai/blog/2026-06-1…
MiniMax M3, Open-Weight, Now On Hugging Face , with only ~428B parameters and ~23B activated parameters Weights: huggingface.co/MiniMaxAI/Mini… MiniMax Sparse Attention: huggingface.co/papers/2606.13…
vime is a reference implementation for one reason only: make @vllm_project the best rollout engine for RL. This helps us better optimize vLLM for the whole ecosystem like @NovaSkyAI SkyRL, @PrimeIntellect Prime-RL, @nvidia NeMo-RL, @verl_project, and more! A wise man in leather jacket said: "We don't build PowerPoint slides and ship the chips. We build a whole data center. And until we get the whole data center built up, how do you know the software works? how do you know your fabric works?" - @NoPriorsPod
Today we're excited to introduce vime — a simple, stable, and efficient RL framework for LLM post-training in the vLLM ecosystem. Built on slime's proven training design and powered by vLLM inference, vime brings another strong option to the growing vLLM post-training ecosystem.
Congrats to @GoogleDeepMind on DiffusionGemma 🎉 A 26B diffusion language model on the Gemma4 backbone, and the first dLLM natively supported in vLLM. It denoises 256-token blocks in parallel instead of generating one token at a time: 1200+ output tok/s at batch size 1 on a single H200 (FP8). Built on model runner v2's ModelState plus the existing speculative decoding path, with minimal scheduler or runner changes. FP8 and NVFP4 checkpoints are on the @RedHat_AI hub. Thanks to the @GoogleDeepMind, @RedHat_AI, and @NVIDIAAI teams! 🔗 vllm.ai/blog/2026-06-1…
Meet DiffusionGemma! An experimental open model that explores a fast approach to text generation, released under an Apache 2.0 license. Moving beyond sequential, token-by-token processes to generate entire blocks of text simultaneously. Here’s what’s new with DiffusionGemma: 👇
Today we're excited to introduce vime — a simple, stable, and efficient RL framework for LLM post-training in the vLLM ecosystem. Built on slime's proven training design and powered by vLLM inference, vime brings another strong option to the growing vLLM post-training ecosystem. Our goal isn't a one-size-fits-all framework. We want users with different needs to find the right vLLM-ecosystem choice for their workflows—whether that's vime, NeMo RL, OpenRLHF, verl, or others. More choice. More interoperability. More innovation. Learn more: vllm.ai/blog/2026-06-0… #LLM #RLHF #PostTraining #vLLM
🎉 Meet vLLM-Omni v0.22.0, a major upgrade for omnimodal world models and production-grade multimodal serving. 🌍 Day-0 @NVIDIAAI Cosmos 3 world models: text, image, audio, video, and action, in and out. 🤖 Robot serving: DreamZero + OpenPI realtime API. 🎙️ Production TTS: Qwen3-TTS, Qwen3-Omni, VoxCPM2 and more. 🎨 Faster image/video/diffusion: Wan 2.2, HunyuanVideo 1.5, LTX-2.3. ⚡ Broader quantization (FP8/INT8, MXFP4/MXFP8, W4A16, ModelOpt) and hardware coverage. 339 commits, 124 contributors, 52 of them new. Thank you all. 🙌 🔗 github.com/vllm-project/v…
🎉 The vLLM community just got a free course, built by @RedHat_AI with @DeepLearningAI. It walks through the full optimize → deploy → benchmark lifecycle for serving open models. Three labs, each on a live vLLM server: - Compress: quantize a Qwen model with LLM Compressor, then measure the size vs. accuracy tradeoff - Serve: deploy with vLLM's OpenAI-compatible API and watch continuous batching, PagedAttention, and prefix caching in the live metrics - Benchmark: simulate traffic with GuideLLM and check quality with lm-eval A lot of the work went into visualizing what actually happens under inference, thanks to @cedricclyburn: how tokens flow through the model, how the KV cache grows in GPU memory, and what changes when you move from FP16 to INT8/INT4. ~1.5 hours, 9 lessons, 3 labs. Free on DeepLearning.AI. 📝 Read more: vllm.ai/blog/2026-06-0…
New short course: Fast & Efficient LLM Inference with vLLM, built in partnership with @RedHat and taught by @cedricclyburn. Learn to quantize an open-source LLM, serve it with vLLM, and benchmark your deployment across speed, cost, and accuracy. Free to enroll:
Amazing work! More and more RL frameworks are using vLLM as default. @vllm_project along with @anyscalecompute and @NovaSkyAI revamped weight syncing and improved wide-ep deployment for rollout!
Excited to share some of our work on improving vLLM for RL! A number of RL frameworks, including SkyRL, use vLLM for inference, and we’ve noticed some common problems: 1. Weight syncing between training and inference is implemented in an ad-hoc fashion and duplicated across
We've shipped two major upgrades for RL✨! 1. Native weight syncing APIs: Standardizes weight transfer, provides optimized implementations for NCCL and CUDA IPC out of the box, and also lets frameworks easily bring their own. 2. Improved pause/resume for Async RL: Careful coordination between DP ranks so that engines don’t deadlock. Validated at scale in P/D, wide-EP setups! In collaboration with @anyscalecompute, @NovaSkyAI, and @RedHat. More and more RL frameworks are using vLLM as the default for inference, details in the blog 👇 vllm.ai/blog/2026-05-2…
🦀 rustifying vLLM, one part at a time, great work @BugenZhao!
🦀 The Rust frontend is officially merged into vLLM! As GPUs get faster, the frontend has become a real share of CPU time. The new Rust frontend is a drop-in alternative to the Python API server — same engine, same ZMQ boundary. Opt in with VLLM_USE_RUST_FRONTEND=1. Early
🦀 The Rust frontend is officially merged into vLLM! As GPUs get faster, the frontend has become a real share of CPU time. The new Rust frontend is a drop-in alternative to the Python API server — same engine, same ZMQ boundary. Opt in with VLLM_USE_RUST_FRONTEND=1. Early numbers: on a preprocess-heavy workload, ~837 req/s vs ~162 req/s for default Python — ~5x in a single process. A few design choices we're excited about: • Layered crates with clear boundaries • Stream-native pipeline — non-streaming for free • Builds on stable Rust Huge thanks to @BugenZhao from @inferact for introducing the work at @PyTorch Meetup Singapore. github.com/vllm-project/v…
Great cohosting this luncheon with @a16z and Mirendil at MLSys 2026 yesterday! 🙌 We brought together top researchers and AI systems engineers for an afternoon of rich conversations on @vllm_project, the frontier of inference, and where AI systems are headed next. Huge thanks to everyone who joined — the energy in the room was something else. This is exactly the kind of cross-pollination between labs, infra teams, and industry that pushes the whole stack forward. More to come. 👀 #MLSys2026 #vLLM
A vLLM MoE deployment's DP/EP topology used to be locked in at launch — scaling or swapping config meant a full restart, in-flight traffic dropped. Elastic Expert Parallelism changes that. One API call resizes a live deployment: curl -X POST localhost:8000/scale_elastic_ep \ -d '{"new_data_parallel_size": 16}' Under the hood: standby comm groups span the target topology, EPLB redistributes experts across the new EP group, and weights are transferred directly between GPUs over NVIDIA NVLink/RDMA. The same runtime reconfiguration path is what fault-tolerant serving needs: evict failed ranks, redistribute their experts, bring replacements back, no restart. Thanks to @NVIDIAAI, Sky Computing, @anyscalecompute, @RedHat_AI, and the community. 📖 vllm.ai/blog/2026-05-1…
🎉 Day-0 vLLM support for Command A+! Congrats to @cohere on their most powerful open-source model yet. 🧠 218B MoE / 25B active, Apache 2.0 🌍 Multimodal + 48 languages ⚡ Runs on as little as 2× H100s @ W4A4 Serve it now in vLLM! 🚀 📖 cohere.com/blog/command-a…
Introducing: Cohere Command A+ We’ve created our most powerful LLM yet, optimized it to run on as little hardware as possible, and released it open-source for all.
🎉 Congrats to the VeRL-Omni team on the pre-release of a general RL post-training framework for multimodal generative models. Built on verl + vllm-omni. vLLM-Omni handles the multimodal rollout with step-wise continuous batching and embedding caching; vLLM serves the VLM-as-judge / OCR reward model, overlapped with rollout and training. In the Qwen-Image OCR demo, moving the reward to its own GPU cuts per-step wall-clock by ~14%. Released: Qwen-Image with FlowGRPO / MixGRPO / GRPO-Guard. BAGEL and Qwen3-Omni-Thinker PR-ready. Excited to push multimodal generative RL forward together with VeRL-Omni and the broader community. 🙌 📖 vllm.ai/blog/2026-05-1… 🔗 github.com/verl-project/v…
@inferact Huge congrats on the second office🚀 Go vLLM!
$AMD is unstoppable now. @AIatAMD ROCm flywheel is spinning hard: persistent MI355X access for @vllm_project @EmbeddedLLM is proud to help power that loop. AI inference is infrastructure. And the next era of AI infra won’t be won by default distribution. It’ll be won by kernels, compilers, runtimes, and relentless execution. We’re here for that fight. 🚀
@AnushElangovan The shock came when on Day 0 DeepSeekv4 launch, since the community vLLM/SGLang maintainers only had access to NVIDIA GPUs, they were only able to add Day 0 NVIDIA GPU support. Since then, AMD has finally priotitzed with actions and just not words by contributing an 2.5million
Great read from the @RedHat_AI team — a comprehensive investigation into TurboQuant in vLLM, with FP8 and BF16 as reference baselines: 4 models (30B to 200B+, decoder-only and MoE) and 5 benchmarks covering long-context retrieval and reasoning, all on the stable vLLM 0.20.2 release. If you're considering TurboQuant for your workload, this is the data to start from. 📝 vllm.ai/blog/turboquant
TurboQuant has drawn a lot of attention recently, but the accompanying evals didn't tell the full story. So we ran what I believe is the first comprehensive study of TurboQuant: where it helps, where it falls short, and how it impacts accuracy, latency, and throughput.
kolade @akoladefaj
83 Followers 387 Following Backend & AI Engineer. Distributed systems, reliability & AI infrastructure. Building @relierdev zero-job-loss for Celery.
L4L4-K @lalacoding
6 Followers 404 Following Tokyo-based Software Engineer focused on AI integration, data workflows, performance software, and computer vision. Python, C++, CUDA.
Darius Tan @dariustan_
33 Followers 172 Following
autodidac @autodidaclzfm
0 Followers 3K Following
Matthew Gordon @46emma538
13 Followers 332 Following
Yuma Ichikawa @yuma_1_or
3K Followers 170 Following Fujitsu (Senior Research Manager), RIKEN AIP, Ph.D. (Univ. Tokyo), The one in the articles? Just a lookalike. I’m actually just doing Gaussian integrals😎
Tang Yanfeng @TangYanfeng
12 Followers 405 Following
Towards a Philosophy ... @goatsintheshell
49 Followers 896 Following
m ahu @mahujam
5 Followers 396 Following
しんちろ @sinchir0
2K Followers 1K Following 一つのことをやりきる / 機械学習エンジニア / Kaggle 2x(Competitions, Notebooks) Master / 共著に「Kaggleではじめる大規模言語モデル入門」「Polarsとpandasで学ぶ データ処理アイデアレシピ55」/ マラソン サブ5.5
FlagOS Community @FlagOS_Official
263 Followers 355 Following An open-source system software stack for AI. Bridging Model–System–Chip layers. Build once, run across diverse hardware.
さるもく|製造... @sarumokueito
73 Followers 265 Following 製造業の受注処理まわりをPythonで半自動化しています。 ノンプログラマーですが、OCR / pandas / Excel自動化 / tkinter を試行中。 現場の手間とミスを減らす改善が好きです。
~JoyCode @AgoroJocelyn
34 Followers 519 Following Passionné de data, amoureux de la science.C/C++,SQL,Python,Django,VueJs, Ionic, ImbaJs😍. Définitivement tombé amoureux de JavaScript❤️
Or Perlman @OruP46
15 Followers 221 Following CTO of https://t.co/IjW3YtIyZk|AI chat & voice for websites|19 yrs shipping code|Tokyo|Father of one|On software craft, AI, and the tech I can't stop reading about
bing @bing29759125
1 Followers 132 Following
シオン@大学用�... @shion1010777
36 Followers 58 Following 18/Male/HOSEI GIS/Table tennis/Ikimonogakari/Comedy/Follow me
Anonymous_joker @anonymo_joker
34 Followers 2K Following
万千 @wanqian_nilk
7 Followers 741 Following
Name Cannot be Blank @modomains
336 Followers 4K Following
SYUN@笑う門には�... @syun88AI
2K Followers 7K Following 27歳の台湾人、AI,ROBOTICS,CV*投稿感想や呟きは、関わってきた組織や会社とは一切関係ありません。 自己紹介https://t.co/peGTZFCcJf
Dd @Ddzfjt
4 Followers 70 Following
Mohamed Zayed @MoZayed007
215 Followers 3K Following Research Engineer On my journey of 10K hrs, Opinions are my own.
mauryaland @mopodono
83 Followers 2K Following
Gary Oikawa @GaryOikawa
578 Followers 3K Following
Kangying L.(Connie) @timcanby
294 Followers 2K Following ML/SWE kaonavi ←SB Intuitions←https://t.co/m94NkbccQb←RecursiveAI(joined pj: https://t.co/CyM5PgT2Pn) ←JSPS DC2(図書館情報学&DH)Ritsumei| Women in Tech🙌New journey ▶️
Hikari∣LocalLLM⚡ @Hikari_07_jp
2K Followers 566 Following 2× RTX PRO 6000 + TR Pro 9965WX | Daily local LLM experiments on real silicon. RepE tuning • vLLM • quantization. Building intelligence I actually own.
じんや @kusojin_nisei
20 Followers 184 Following
Rafique Shaik @shaiktweetss
416 Followers 4K Following TPM at Google Japan|| Im lost on a quest for an utopian peace. 人生はダイジョウブじゃない #teampixel
subabusucu @subabusucu
0 Followers 118 Following
Kürşat Kılıç @KursatKilic_
146 Followers 2K Following BSc @muglaedutr MSc @PoliTOnews PhD @syudaikouhou MEXT Research Fellow @mextjapan ex-TUM Boring @TUM_Boring
Mohammad Elnaqa @MohammadElNaqa
0 Followers 102 Following
Roshan Rajan @roshanjrajan
20 Followers 654 Following
AMDYES @9999x9d
198 Followers 4K Following
波妞PONYO @ponyodong
11K Followers 707 Following AIGC Video Creator & Animation IP Strategist 🐷 前 Peppa Pig China Mktg. 🎨 专注 AI 视觉叙事与漫画 IP 孵化 💡 分享高阶 Prompt 与工作流实战 商业咨询PONYODONG
Jackmin @jackminong
2K Followers 929 Following making sand smarter @PrimeIntellect 🇺🇸 Previously @JinaAI_ 🇩🇪 @MoneyLion 🇲🇾
serein @you1873118
13K Followers 7K Following I hope tomorrow will be better! Bluebird Club Just here for the memes Sharing everything fun
SeeInX (AI Art) @seeinx_aiart
12K Followers 2K Following AI Art | ComfyUI | Realistic Style I still have so much more I want to show you..
Sigrid Jin 🌈🙏 @realsigridjin
15K Followers 1K Following member of prompt and pray 😭✌️ @ubc @thisissigrid experiencing context rot but 🇨🇦 🇰🇷 proudly korean-canadian
Vigo Zhao @VigoCreativeAI
7K Followers 348 Following Designer × AI Engineer | Ambassador @Alibaba_Qwen | Cut e-commerce AI costs ~60%|Reverse-engineering prompts from a single image
GENEL | AIを用い�... @genel_ai
18K Followers 180 Following MO3IC所属|AIを用いた動画制作|企業研修も| 検証: 画像生成AI / 動画生成AI / 音声生成AI| ゼロからAIでCMや動画が作れる解説はnoteで📒 Kling AI Creative Partner
ミロ @ml0_1337
4K Followers 324 Following AI SaaSの開発会社でCEOやってます。フォローをするとAI時代に生きる起業家としてのリアルが見れます。技術顧問等の相談はDMからお願いします。大学中退の高卒→受託事業→AI SaaS開発・運営、外部CTO/技術顧問
TonoKen3🤖Local-LLM... @Tono_Ken3
2K Followers 1K Following Founder/CEO of Lna-Lab K.K. Former publishing exec building local LLMs, AI agents & Physical AI for human-AI coexistence with RTX PROs (SM120).Civilization OS🤖
からあげ @karaage0703
30K Followers 2K Following エンジニア/AIのお仕事してます/からあげ大好き/はてなブログ書いてます/色々作ってます/からあげは概念/Amazonのアソシエイトとして適格販売により収入を得ています
株式会社フィッ... @Fixstars_JP
5K Followers 524 Following フィックスターズ公式アカウント。 「Speed up your AI」をスローガンに、優秀なエンジニア達がお客様のAI活用とAI開発を加速しています。グループ全体の最新情報をお届けします。 【東証プライム:3687】米国でのビジネスはこちら → @Fixstars_US
ハカセ アイ(Ai-H... @ai_hakase_
14K Followers 826 Following 生成AIと猫への愛50%ずつで構成されています、葉加瀬あいです🐾 🧠クリエイター向けAI解説:YouTube、Note、X 🤖AIエージェント開発:自社、受託 などを行う、小さな会社を経営しております。 PC周り、ジム、ゴルフ場が最近の住処です😽✨ お役立ち投稿は「ハイライト」欄にまとめてます🥳
Hiroki Yamamoto @tereka114
4K Followers 813 Following Acroquest Technology Co., Ltd/Data Scientist/Kaggle Grandmaster/CV/SA/C++/Python/Kaggle https://t.co/ANLAF8Pq7K お仕事などのご依頼、ご相談はDMorMLにてご連絡ください。
Yuma Ichikawa @yuma_1_or
3K Followers 170 Following Fujitsu (Senior Research Manager), RIKEN AIP, Ph.D. (Univ. Tokyo), The one in the articles? Just a lookalike. I’m actually just doing Gaussian integrals😎
Qubitium @qubitium
1K Followers 4K Following Building GPT-QModel, ModelCloudAI. Contributor to stuff you are probably using.
Chew Kok Wah @chewkokwah
29 Followers 1K Following Make the world a better place 创造美好新世界 Wujudkan dunia baru yang lebih indah
Robert Shriver Barnes @sailorbob74133
84 Followers 53 Following The righteous will rejoice when he sees the vengeance; he will bathe his feet in the blood of the wicked. Yonah Cp. 4: should not I have pity on Nineveh...?
Kalyan @nkalyanv99
66 Followers 105 Following 1st year PhD Student @TU_Muenchen. On a random walk through the subfields of ML.
Enrique Guerra 🇺�... @EnriqueGuerraF
167 Followers 2K Following PhD in ML @MonashUni - Libertarian, e/acc.
~JoyCode @AgoroJocelyn
34 Followers 519 Following Passionné de data, amoureux de la science.C/C++,SQL,Python,Django,VueJs, Ionic, ImbaJs😍. Définitivement tombé amoureux de JavaScript❤️
Ally @treasureh8nter
2K Followers 523 Following 🌏 UK-Taiwan | Ex-Tokyo Banking (16yrs) 💹 AMD & AI Investor 🚀 XRP Enthusiast 🔎 Visual Data & Market Insights 👇 “Follow” for cutting-edge tech & finance!
Terry_ray @Terryra00832493
31 Followers 118 Following Student life: No steady income yet, just dipping into fund investments. Tech junkie devouring industry articles daily.
Shyamji Tiwari @Shyamji33892908
25 Followers 107 Following
Hamid Shojanazeri @Nazeri2010
390 Followers 924 Following ml @pytorch model optimization, distributed training.
Graver256 @Graver256
26 Followers 270 Following
Amil Agrawal @amilabsolute
26 Followers 829 Following applied research @snap | cs + math @brownu | prev robots @raytheon, inference opt @amd | 20
Anmol soin @anmol_soin
151 Followers 891 Following Not much into tweet ;) I better do insta - @anmol_amo ✌️ Peace out.
Sourav Chakraborty @souravzzz
62 Followers 451 Following
Doplano @_doplano_
97 Followers 4K Following
Wes Henderson @wesjh_
98 Followers 357 Following Focused on Enterprise AI @ AMD |Former IP Lawyer | Current AI strategist | Too many hobbies and interests to count. UMSL BioChem| @SLULaw | @UTAustin AI/ML
deadend890 @deadend672
2 Followers 2 Following
Sam Luu @TheRea10G
8 Followers 50 Following
DanNeo S.S. ✮ 𓂃�... @DanNEO_SS
2K Followers 2K Following Wordsmith, Indie writer and creator of the Dark Neology. Available on Amazon and Kindle. #indieauthor #DarkNeology🏴
Razine Moundir Ghoarb @RazineMG
86 Followers 1K Following atoms to abstractions, one token at a time.
Nolan @Nolan1831749
8 Followers 25 Following
vatsal choksi @vatsalchoksi7
18 Followers 66 Following
Stephan90 @Stephan9015
606 Followers 190 Following
aCE @AlybiSaCE
3K Followers 3K Following @? | Representing the best @EpikWhale @anonzr @Crylix @fnmoneymaker @regsita & more
Davit Khachatryan @DavitKhachatry
48 Followers 166 Following
Ethai Reubinoff @EthaiReubinoff
515 Followers 179 Following 🗽/acc - Techno Libertarian Utopia - GPUs BTC and Guns - Money, Tech, Markets, Freedom, Knowledge. Early $PLTR, $TSLA, $AMD, $BTC investor.




































