On the Job Market | Final-Year CS PhD @ Northwestern University | Co-Founder, Cortices | Adobe Fellow | DAAD AINet Fellow

Evanston,IL
Heading to 🌉 SF for #COLM2026! I’ll be there Oct 6–9 to present our work DAPA 🎉 DAPA explores how to efficiently repair and strengthen the safety of foundation models, with a focus on practical and scalable safety alignment. My research interests lie at the intersection of LLM infrastructure and safety alignment, and I’m currently exploring LLM infrastructure safety—including safety challenges in efficient LLM systems such as recursive self-improvement (RSI), loop transformers, and speculative decoding. If you’re working on related topics, let’s chat over coffee! ☕️ Also welcome to join our second version of efficient reasoning workshop on Oct 9:) P.S. I’m on the job market and currently looking for 2027 Full-time Summer Research Scientist. Would love to connect! 🤝 Paper 📖: openreview.net/forum?id=6EJ2… Code 🔧: github.com/NWULIST/DAPA
5
126
Welcome to our COLM Workshop.
Everyone who comes to COLM, you are welcome to join our workshop on October 9, 2026 · Hilton Union Square, San Francisco!
33
Haozheng LUO retweeted
Everyone who comes to COLM, you are welcome to join our workshop on October 9, 2026 · Hilton Union Square, San Francisco!
1
6
28
7,660
Haozheng LUO retweeted
🤖 Reusing skills is one possible way for embodied agents to generalize. Manipulation varies but shares a few skills: wiping a table or a window is one skill. Many works give agents skills to plan with. But where do these skills, and their data, come from? 📼 Today, VLMs mostly label videos one at a time, and nothing carries over. People don't learn that way. As experience streams in, we spot skills we know, add new ones, and reuse them later. This streaming setting matters, but it has received little attention. 🧠 VLMs may become the brains of embodied agents. So we ask: from streaming experience, can they find reusable skills and keep one consistent skill library? 🧵 Video2Skill: From Streaming Experience to Reusable Embodied Skills Why it matters: 📈 It can label much more data for training agents. 🔁 It is planning in reverse. If a model can't find skills in what it has seen, how can it plan with them for something new? 📄 Paper: arxiv.org/abs/2609.36691 🌐 Project: andyzworks.github.io/video2s… 🤗 Data: huggingface.co/datasets/Ster…

ALT Animated loop: robot videos stream in; each event is tagged NEW (orange) and flies into a skill library, or REUSE (teal) and links to an existing skill; the library grows from grasp, transfer, place to pour.

2
5
10
324
Congrats to my collaborators! Attention Residual has emerged as an important technique in recent frontier models and top-tier research, but it can substantially amplify attention sinks and activation outliers.
🎉 OASIS is accepted to NeurIPS 2026! 🔗A follow-up to Kimi Attention Residuals @Kimi_Moonshot ⚠️ We find that the extra depth-wise Softmax in Attention Residuals can amplify attention sinks, activation outliers, and low-bit quantization errors. 💡 OASIS adds explicit null routes to both token and depth routing, then feeds token-level null signals back into depth routing. 📈 Across multiple models: ↓95.9% kurtosis / ↓82.0% W8A8 PPL / ↑42.1% W4A4 GSM8K 📚 On 12K context, OASIS also substantially recovers the performance drop of AttnResidual on RULER and LongBench. ✨ More robust Attention Residuals for low-bit inference and long context. Thanks to all the collaborators across 8 institutes! @robinluo1997 @HHarryD @Michael_Huang_W @NorthwesternU @illinoistech @RutgersU @UMich @UCLA @UCSD @TAMU @northwesterncs @iitcsdept @RutgersCS Paper Link: arxiv.org/pdf/2605.17887 Project Website: oasis-research.github.io/
3
59
Haozheng LUO retweeted
🎉 OASIS is accepted to NeurIPS 2026! 🔗A follow-up to Kimi Attention Residuals @Kimi_Moonshot ⚠️ We find that the extra depth-wise Softmax in Attention Residuals can amplify attention sinks, activation outliers, and low-bit quantization errors. 💡 OASIS adds explicit null routes to both token and depth routing, then feeds token-level null signals back into depth routing. 📈 Across multiple models: ↓95.9% kurtosis / ↓82.0% W8A8 PPL / ↑42.1% W4A4 GSM8K 📚 On 12K context, OASIS also substantially recovers the performance drop of AttnResidual on RULER and LongBench. ✨ More robust Attention Residuals for low-bit inference and long context. Thanks to all the collaborators across 8 institutes! @robinluo1997 @HHarryD @Michael_Huang_W @NorthwesternU @illinoistech @RutgersU @UMich @UCLA @UCSD @TAMU @northwesterncs @iitcsdept @RutgersCS Paper Link: arxiv.org/pdf/2605.17887 Project Website: oasis-research.github.io/
5
14
104
18,747
Haozheng LUO retweeted
🚀 New paper alert: On-Policy Self-Distillation without Any Supervision !! 🧠 Can a model truly on-policy “self-distill"? On-policy self-distillation removes the need for a stronger teacher, but still relies on an external ground-truth solution to make the self-teacher more capable. Our new work asks: can we remove supervision altogether? Can an LLM improve its reasoning entirely from its own generations—without ground-truth answers, verifiers, or a stronger teacher? ♻️ We introduce U-OPSD, a simple approach toward genuine self-distillation: For each problem, the model samples multiple solutions from itself: 🗳️ Agreement identifies a likely solution through majority vote 🔀 Disagreement reveals trajectories where the model may be going wrong 🎓 The pseudo solution conditions the self-teacher, which provides dense token-level guidance along those disagreeing trajectories U-OPSD turns self-consistency into supervision: agreement builds the teacher context, while disagreement determines where to distill. 🔓 No external supervision at all 📚 Unlabeled problems only: 🚫 GT solutions 🚫 Verifier 🚫 External strong teacher 🚫 Environment feedback 📊 Yet, across 5 math reasoning benchmarks and 6 Qwen3 settings, U-OPSD consistently improves the base model and matches or surpasses supervised SFT, GRPO, and OPSD. 🌙 Non-thinking mode: 📈 +8.5 / +10.7 over the 4B / 8B base models 🏆 +3.2 / +2.3 over GT-supervised OPSD 🚀 +7.0–11.3 over label-free self-rewarding RL methods including TTRL, RENT, and Intuitor 🧠 Thinking mode: 📈 +2.2 / +1.9 over the 4B / 8B base models 🤝 On par with GT-supervised OPSD and GRPO 🚀 +0.8–1.4 over label-free self-rewarding RL methods including TTRL, RENT, and Intuitor Meet U-OPSD 👇 📄 ArXiv Paper: arxiv.org/abs/2608.06296 💻 GitHub: github.com/williamium3000/u-… 🌐 Project Page: williamium3000.github.io/u-o… 🤗 Hugging Face Paper: huggingface.co/papers/2608.0… (Plz upvote if you can !! Amazing collaboration with @ Bingyang @JoLiang17 @tian_yunjie @Di Fu and Nuno ! Grateful for all the insightful discussions and exploration together! #LLM #OPD #OPSD #Distillation
16
14
851
Fast and Low-Cost Genomic Foundation Models via Outlier Removal 1. GERM is a new genomic foundation model that tackles a key problem in transformer-based genomics models: outliers in attention layers that hinder efficient adaptation and quantization. By removing these outliers, GERM becomes significantly more robust and lightweight. 2. Unlike traditional models like DNABERT-2, GERM replaces the standard attention mechanism with an outlier-free layer inspired by associative memory. This leads to a dramatic improvement in fine-tuning and post-training quantization, especially in resource-constrained environments. 3. GERM achieves up to 64.34% improvement in quantization robustness and 37.98% better performance in low-rank adaptation compared to DNABERT-2. It also reduces average kurtosis by 92.14% and the maximum infinity norm by 82.77%, key indicators of outlier suppression. 4. To make the model more accessible, the authors also introduce GERM-T, a continual learning approach that avoids retraining from scratch while still significantly reducing outliers and improving performance on low-end hardware. 5. GERM is efficient enough to fine-tune in 5 minutes on a single RTX 2080 Ti GPU. Even in extreme settings like 4-bit quantization, GERM retains competitive performance, unlike DNABERT-2 which collapses under such compression. 6. The study includes extensive benchmarking across 27 datasets, various quantization strategies (e.g., SmoothQuant, OmniQuant), and fine-tuning techniques (LoRA, QLoRA, LoftQ), showing that GERM consistently outperforms baseline models in both accuracy and hardware efficiency. 7. GERM can also generalize to other genomic foundation models like the Nucleotide Transformer, achieving over 50% improvement in low-rank adaptation and nearly 70% improvement in quantization robustness. 8. This work paves the way for deploying genomic foundation models on edge devices, making advanced sequence modeling more accessible to biomedical researchers without requiring massive computing infrastructure. 💻Code: github.com/MAGICS-LAB/GERM 📜Paper: arxiv.org/abs/2505.00598v1 #Genomics #TransformerModels #Quantization #LoRA #Bioinformatics #FoundationModels #MachineLearning #AI4Science #ComputationalBiology
1
13
37
6,169
Haozheng LUO retweeted
Or… what if we gave you $100 in Codex credits if you tell us what you love about GPT-5.6 Sol or why you switched? Tweet it, claim your gift, enjoy more usage. First 10k get the free tokens! switch-to-codex.openai.chatg…
codex should give one month free when you send them your claude cancelation screenshot
13,240
3,623
15,469
6,526,612
Haozheng LUO retweeted
🌟 Stella Sora Anime Expo 2026 Event Guide 🌟- Follow & Repost for a chance to win exclusive freebies! Here's your event guide to Stella Sora at AX2026. We can't wait to meet you! 🎁Follow & Repost the event guide post for a chance to win exclusive freebies! ▼How to Participate ① Follow @StellaSoraEN ② Repost this post ▼Event Period Jun. 09 – Jun. 23, 23:59 (UTC-7) ▼Event Prizes (shared among X & FB) We will randomly select 10 winners from X and Facebook. Each winner will receive one of each of the following: 🌟3D Lenticular Fan 🌟Tattoo Sticker 🌟Cooling Headband 🌟AX-exclusive Postcard #StellaSora #Yostar #AX2026
45
940
689
87,865
Haozheng LUO retweeted
🌟 Announcing the 2nd Workshop on Efficient Reasoning (ER) at @colm2026 — Oct 9! 📣 We welcome submissions! Submit your work here: openreview.net/group?id=colm… 🗓️ Deadline: July 12, 2026 (AoE) 🔗 Website: wdlctc.github.io/efficient-r… 💬 Topics include (but aren't limited to): 🔹 Multimodal, spatial & embodied reasoning under efficiency constraints 🔹 Curating high-quality reasoning datasets under resource constraints 🔹 Algorithmic innovations for efficient training & RL fine-tuning 🔹 Fast inference: pruning, compression, progressive generation, KV-cache tricks 🔹 Benchmarks & theory on time-/space-complexity and faithfulness 🔹 Systems to deploy long-CoT or on-device reasoning in the wild 🔹 Safety & robustness of efficient reasoning pipelines 🔹 Real-time applications in healthcare, robotics, autonomy, and more 🤝 We invite perspectives from ML, systems, natural & social sciences, and industry practitioners to rethink reasoning under tight compute, memory, latency, and cost budgets. Hope to see you there! 🚀
4
18
14,584
🚀 Launching Growing LLMs Beyond Boundaries (GLB) — a non-profit initiative to push LLMs beyond task-specific limits and into real-world impact. Goal: move from scaling → boundary-expanding systems (multimodal, adaptive, real-world).Join us: growing-llms-beyond-boundari…
28
Haozheng LUO retweeted
Cortices is mission control for AI agents, with support for Claude Code and Codex. Bring agents across laptops, servers, and workstations into one place. Chat with agents, review diffs, and schedule work from any browser or phone. Free, no waitlist: cortices.io
1
1
3
25
Haozheng LUO retweeted
🚨 New paper alert !! 🎥 Video VLMs are strong at high-level semantics and long-range temporal understanding. 🧠 JEPA is almost the opposite: better at dense, high-frequency dynamics, local physical consistency, and fast corrective control, but are less suited for rich semantic reasoning and long-horizon reasoning. We try to get the best of both: 🧩 A VLM as a cortex-like reasoner for semantics and long-horizon planning ⚡ A JEPA branch as a cerebellum-like controller for fine-grained dynamics, physical consistency, and rapid corrections Proudly, we present ThinkJEPA: a VLM-guided latent world model that FiLM-fuse the pyramid repr of VLMs encoding long-horizon semantic reasoning into the JEPA repr for fine-grained, physically consistent dynamics prediction. 🔗 Project: zhanghaichao.xyz/ThinkJEPA/ 📄 Paper: arxiv.org/pdf/2603.22281
7
67
352
18,020
We introduce TIDES, the first work showing that test-time scaling can be adversarially exploited.
By injecting small latent perturbations (DLT), we induce inference drift that amplifies with reasoning depth, causing longer reasoning traces to degrade performance. #safety
26
🧠 Make models reason safely, not just respond safely We introduce CRAFT (Contrastive Reasoning Alignment) to address a key gap: current safety methods mostly constrain outputs, while reasoning trajectories remain unaligned.
17