What do these badges mean?
- 🚀ShippingCode exists. Multiple GitHub repos already reference this paper — people are building on it.
- 📈ClimbingCitation velocity is rising. Researchers are starting to pick it up.
- 💤QuietPublished but no notable signal yet. Most papers live here — could become anything later.
- 🎭HypeHeavy social buzz but no shipping signal. The counter-signal — defer until Twitter/X data is wired up.
- 💤Quiet2608.23531·Aug 24, 2026·~13 mincs.CVcs.LG
Predicting Multiple Clinical Outcomes Related to Functional Recovery and Social Isolation Among Older Adults After Lower-Limb Fracture or Hip Replacement
Santosh Ray, Pratik K. Mishra, Ali Abedi, Charlene H. Chu, +2
⭐ 0 stars / 0 repos📚 0 citesELI5Researchers used smartwatch-like sensors and fitness trackers worn by older adults recovering from hip/knee surgery to predict multiple health outcomes at once—things like mobility, strength, and social isolation—rather than predicting each separately.
Problem solvedDoctors currently assess recovery from orthopedic surgery in isolated ways, missing how different aspects (mobility, strength, isolation) influence each other. This predicts them together from wearable data, enabling earlier intervention for struggling patients.
- 💤Quiet2608.23484·Aug 24, 2026·~11 mincs.AI
Multi-Modal Semantic Expansion with Constrained LLM Reranking for Conversational Music Recommendation
Naman Garg, Sarika Jain, George Fazekas
⭐ 0 stars / 0 repos📚 0 citesELI5A music recommendation system that listens to what you say in a conversation and suggests songs by searching through 7 different ways of understanding music (from audio to lyrics to what similar users liked), then ranks the best matches and generates natural responses using AI.
Problem solvedMusic platforms struggle to recommend songs based on natural conversation rather than just clicks or search—users want to say what they're in the mood for and get personalized suggestions, not navigate menus. This system combines multiple AI approaches to understand context and deliver better recommendations in real conversations.
- 💤Quiet2608.23435·Aug 24, 2026·~7 mincs.CVcs.AI
Towards Comprehensive Basketball Understanding
Yirong Hu, Jiayuan Rao, Yu Zhang, Shangzhe Di, +1
⭐ 0 stars / 0 repos📚 0 citesELI5A new benchmark tests whether AI can understand basketball games by combining multiple skills—identifying players, recognizing plays, looking up stats, and connecting them together. Most AI systems fail when you mix these skills, but a modular agent that chains specialized tools does much better.
Problem solvedCurrent AI video understanding systems are tested piecemeal (one skill at a time), but real-world sports analysis requires integrating multiple capabilities simultaneously. This benchmark reveals that multimodal models struggle when tasks demand coordinated reasoning across perception, retrieval, and domain knowledge.
- 💤Quiet2608.23373·Aug 24, 2026·~14 mincs.AIstat.AP
Modalities Should Talk to Each Other: Dual-Stream Multimodal Learning for Long-Horizon Influenza Forecasting
Seyed Mohammad Hossein Hashemi, Mohsen Hooshmand, Parvin Razzaghi
⭐ 0 stars / 0 repos📚 0 citesELI5A system that predicts flu outbreaks 3 months ahead by teaching news headlines and disease numbers to inform each other. Headlines provide context about what's happening, while disease numbers provide ground truth, making the prediction much more accurate than using either signal alone.
Problem solvedPublic health agencies struggle to forecast flu 12 weeks out because available data mixes reliable numbers with noisy news text. Traditional methods ignore the text or handle it poorly, missing the contextual signals headlines provide about outbreak conditions.
- 💤Quiet2608.23363·Aug 24, 2026·~8 mincs.CVcs.AIcs.LG
DF-MoE: Generalizable Deepfake Detection via Multimodal Sparse Mixture-of-Experts
Vlad Hondru, Florinel Alin Croitoru, Iuliana Georgescu, A. Sophia Koepke, +1
⭐ 0 stars / 0 repos📚 0 citesELI5This system detects fake videos by combining clues from multiple specialized AI models (like ones that track mouth movements, facial expressions, and voice emotion) and blending their outputs smartly so it generalizes better across different types of deepfakes.
Problem solvedDeepfake detectors trained on one type of fake video fail when tested on videos made with different methods. This approach extracts diverse signals from audio and video that are harder to fool consistently, making detectors work better across different deepfake generation techniques.
- 💤Quiet2608.23344·Aug 24, 2026·~13 mincs.LG
Towards Actionable Surgical Team Dynamics: from Teamwork to Counterfactual Annotations
Vincenzo Marco De Luca, Antonio Longa, Andrea Passerini
⭐ 0 stars / 0 repos📚 0 citesELI5Researchers created a dataset of real operating room recordings with detailed annotations about how surgical teams interact, communicate, and perform—plus imaginary alternate scenarios showing what could have happened differently if certain team members had acted differently.
Problem solvedSurgical teams need better ways to understand what causes coordination failures and poor outcomes in the OR. Existing data is messy and fragmented; this unified dataset with counterfactual annotations lets researchers and AI systems learn what behavioral changes could prevent surgical mishaps.
- 💤Quiet2608.21357·Aug 21, 2026·~7 mincs.AI
VIALS: A Benchmark for Visual Interpretation of Artifacts in the Life Sciences
Elaine Lau, Thanuka Udumulla, Lee Izhaki-Tavor, Francisco Guzmán, +2
⭐ 0 stars / 0 repos📚 0 citesELI5A new test set shows that AI vision models struggle to read scientific images like microscope photos, gel blots, and DNA maps—the kinds of pictures biologists look at every day to make decisions. These models work fine on everyday photos but fail when the images need real domain expertise to interpret.
Problem solvedBiotech labs need AI assistants that can actually understand scientific imagery, not just describe it. Current vision-language models can't reliably interpret the visual artifacts scientists depend on, making them unsuitable for real lab workflows where these images drive critical research decisions.
- 💤Quiet2608.21277·Aug 21, 2026·~9 mincs.LG
ConceptTS: LLM-Guided Concept Bottlenecks for Interpretable Multivariate Time-Series Forecasting
Yichen Jiang, Yueqiao Chen, Dongyu Liu
⭐ 0 stars / 0 repos📚 0 citesELI5This system teaches an AI to forecast future time-series data (like air quality over time) by first breaking down its reasoning into human-readable concepts—like 'pollution trend' or 'weather pattern'—which it learns from an LLM's suggestions. Instead of a black box, you can see exactly which concepts the model used to make each prediction.
Problem solvedTime-series forecasting models are accurate but opaque—you can't tell why they predicted what they did, making them risky for regulated domains like environmental monitoring or finance. This adds interpretability without sacrificing accuracy, so practitioners can audit predictions and spot when the model relies on suspicious patterns.
- 💤Quiet2608.20331·Aug 20, 2026·~11 mincs.CLcs.AIcs.CV
G-CARL: Grounded Checklist-Aligned Reward Learning for Patient-Oriented Medical Report Interpretation
Shiao Xie, Siyu Chen, Jianwei Lv, Bo Yuan, +2
⭐ 0 stars / 0 repos📚 0 citesELI5A system that helps doctors explain medical reports to patients in plain language. It uses a smart training method that checks facts against medical sources while making sure the explanation answers what the patient actually asked.
Problem solvedPatients struggle to understand medical reports, and current AI systems either get facts wrong or miss patient concerns. This paper fixes both problems at once by training models to be both accurate and patient-focused, verified by real clinicians.
- 💤Quiet2608.20320·Aug 20, 2026·~12 mincs.AIcs.CL
An Agentic Approach for Active Data Collection, Travel Behavior Modeling, and Weather-Sensitive Demand Prediction
Narges Ahmadi, Yubo Jiao, Jônatas Augusto Manzolli, Jiangbo Yu, +1
⭐ 0 stars / 0 repos📚 0 citesELI5A system uses three AI agents working together: one chats with commuters about how they'd travel in different weather, another processes the data, and a third predicts travel choices. They compare traditional statistical models against different-sized AI language models, finding that language models with pictures of the weather actually predict commute choices as well as traditional methods.
Problem solvedTravel researchers manually build surveys, then separately run prediction models on the results—it's fragmented and hard to audit. This creates a unified workflow where conversational AI collects the data, processes it, and makes predictions all together, making the whole pipeline traceable and reproducible.
- 💤Quiet2608.20237·Aug 20, 2026·~11 mincs.AI
Rule-Compliant Visual Spatial Planning for Multimodal Large Language Models
Yu Chen, Ting Lei, Yaoyi Li, Jia Cai, +2
⭐ 0 stars / 0 repos📚 0 citesELI5This paper tests whether AI vision models can follow rules while planning paths through mazes. Instead of just finding any route, the model must navigate while obeying natural-language constraints (like 'avoid red tiles'), which requires understanding both the visual layout and the rules simultaneously.
Problem solvedCurrent vision-language models struggle with constrained planning tasks where they must obey explicit rules. Real-world applications like robotics and game AI need models that can verify rule compliance and plan accordingly, but there's no good benchmark or technique to measure or improve this ability.
- 💤Quiet2608.20186·Aug 20, 2026·~14 mincs.LGq-bio.NC
Decoding silent reading from non-invasive EEG
Ingo Marquardt, Anthilia Alchanat, Priyanka Jain
⭐ 0 stars / 0 repos📚 0 citesELI5Researchers used EEG brain scans to decode which words a person is silently reading by training a neural network to match brain signals with word embeddings from a language model. They collected 49 hours of data from one person and showed the system can reliably identify words from brain activity alone.
Problem solvedBrain-computer interfaces for reading inner speech have been impossible to train because you can't get ground-truth data of what someone is thinking. This work uses silent reading as a scalable alternative to finally demonstrate that word-level information is recoverable from non-invasive EEG without surgically implanted electrodes.
- 💤Quiet2608.20116·Aug 20, 2026·~8 mincs.CL
When Text and Numbers Disagree: Evidence Arbitration in Large Language Models
Mattia Carletti, Edward Phillips, Fredrik K. Gustafsson, Patitapaban Palo, +4
⭐ 0 stars / 0 repos📚 0 citesELI5When an AI gets conflicting information—like text saying one thing and numbers saying another—this paper tests how it decides which to believe. Researchers built a fake scenario with intentional conflicts to see if models have biases toward text or numbers, and they found models use simple patterns (like trusting newer data) rather than carefully weighing evidence.
Problem solvedAI systems that use multiple tools and data sources often get contradictory signals. If an LLM blindly trusts one type of evidence over another, it could make bad decisions—like accepting a forecaster's prediction even when direct data contradicts it. This paper exposes those blind spots so we can fix them.
- 💤Quiet2608.20019·Aug 20, 2026·~11 mincs.AI
Contrastive Mixed Prompt Learning for Incomplete Multimodal Sentiment Analysis with Unseen Modality Combination
Kaixin Xu, NaiJin Liu, Yulin Kang, Tangyue Jin, +7
⭐ 0 stars / 0 repos📚 0 citesELI5When analyzing sentiment from multiple data types (text, image, audio), the model learns patterns from certain combinations but encounters different combinations at test time. This paper teaches the model to handle new combinations it's never seen before by using contrastive learning and flexible prompt templates.
Problem solvedReal-world systems often see modality combinations at test time (e.g., video+text) that weren't in training data (e.g., only audio+text). Existing models fail on these unseen combinations. This fixes generalization to handle any modality mix without retraining.
- 💤Quiet2608.19981·Aug 20, 2026·~8 mincs.CL
HealMed: Multilingual Evaluation of Large Language Models in Medicine
Yingjian Chen, Fan Gao, Sherry T. Tong, Haoyu Zhang, +41
⭐ 0 stars / 0 repos📚 0 citesELI5A team of doctors created a test with 9,000 medical questions across 9 languages to see how well AI models answer medical questions globally—and found that models perform much worse in some languages than others, especially less common ones.
Problem solvedMedical AI systems are being deployed worldwide, but we didn't have a reliable way to test if they actually work well across different languages. This benchmark shows which models are trustworthy for non-English patients and reveals that many models fail in low-resource languages.
- 💤Quiet2608.19128·Aug 19, 2026·~12 mincs.LG
Beyond Trial Averaging: Anchoring Neural and Visual Representations for Few-Repetition Brain-to-Image Retrieval
Zhenyao Cui, Siyuan Kan, Dingkun Liu, Dongrui Wu
⭐ 0 stars / 0 repos📚 0 citesELI5When people see images, their brain signals can be used to retrieve those images from a database. The trick is that noisy brain signals from just one or two viewings work poorly—you currently need 80+ repetitions. This paper fixes it by using a 'reference point' (an averaged signal) to pull both the noisy brain signal and the image toward each other simultaneously, like magnets meeting in the middle.
Problem solvedBrain-to-image decoding currently requires 80+ repeated stimulus presentations to work reliably, making it impractical for medical, assistive tech, and dream research applications. This work enables accurate retrieval from just 1–4 repetitions, dramatically reducing user burden and latency.
- 💤Quiet2608.16805·Aug 17, 2026·~12 mincs.CVcs.AI
Diagnosing Dense Same-Class Attribute Misbinding in Large Vision-Language Models
Yuanzhi Xu, Qian Gao, Jun Fan, Guohui Ding, +3
⭐ 0 stars / 0 repos📚 0 citesELI5Vision-language models can see multiple similar objects and their colors, but often mix up which color belongs to which object—like describing 'the red cat' when they meant 'the blue cat.' This paper builds a test to catch and measure exactly when this happens.
Problem solvedCurrent benchmarks can't tell the difference between a model genuinely failing to see something versus correctly seeing multiple objects but assigning attributes to the wrong one. This distinction matters for safety and trust in real-world applications.
- 💤Quiet2608.14446·Aug 14, 2026·~10 mincs.AI
Wyvern: An Agentic Framework for Generating Grounded Multimodal Reports
Beatrice Alessandra Motetti, Emilien Guandalino, Daniele Jahier Pagliari, Alessio Burrello, +3
⭐ 0 stars / 0 repos📚 0 citesELI5A system that automatically writes technical reports by coordinating multiple AI agents to gather information, create images and tables, fact-check claims against sources, and package everything into a single document with citations.
Problem solvedAI-generated content often makes up facts or lacks sources to back claims up. This system ensures reports are grounded in real information and properly cited, making them trustworthy enough for people to actually use in technical contexts.
- 💤Quiet2608.13560·Aug 13, 2026·~11 mincs.CVcs.AIcs.CL
AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design
Yaxin Luo, Haobin Jiang, Jialv Zou, Xu Huang, +10
⭐ 0 stars / 0 repos📚 0 citesELI5A system that automatically learns how to design better academic posters from research papers by repeatedly trying different design approaches, getting feedback, and improving its process — like a designer who gets better at their job the more posters they make.
Problem solvedConverting papers into visually effective posters is tedious and requires constant manual tweaking. This system automates the entire design workflow, learning from failures to improve over time, saving hours of manual work while producing higher-quality results than existing tools.
- 💤Quiet2608.13558·Aug 13, 2026·~13 mincs.AIcs.CL
OmniScientist: An Omni-Modal Omni-Discipline AI Scientist
Bobo Li, Hao Fei, Tianjie Ju, Mong-Li Lee, +1
⭐ 0 stars / 0 repos📚 0 citesELI5An AI system that reads raw scientific data (images, videos, audio, 3D models, tables, etc.) directly and runs through a complete research pipeline—generating ideas, running experiments, and writing papers—without losing information to preprocessing.
Problem solvedAI research assistants today work with pre-summarized numbers or text, missing crucial patterns in raw images, videos, and spatial data. This system processes all raw evidence throughout the research lifecycle, making scientific claims more rigorous and discoveries more reliable.
- 💤Quiet2608.13505·Aug 13, 2026·~12 mincs.LGcs.CLcs.CV
Intern-S2-Preview: Scientific Agentic Foundation Model
Lei Bai, Jiaqi Cao, Chiyu Chen, Guanzhou Chen, +121
⭐ 0 stars / 0 repos📚 0 citesELI5A large AI model trained to understand scientific papers, images, and data, then taught to use tools and run experiments to solve research problems step-by-step over long periods of time.
Problem solvedScientists need AI that can read mixed-format research (papers, charts, tables), reason about evidence, and autonomously run experiments—not just answer questions. This model is built from the ground up to handle those tasks.
- 💤Quiet2608.13463·Aug 13, 2026·~9 mincs.CVcs.AIcs.CL
MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification
Daniel Perkins, John Squires, Janou Milligan, Chandra Raskoti, +1
⭐ 0 stars / 0 repos📚 0 citesELI5A system that uses a smart language model to look at images and decide which of several different AI vision models should classify each one, depending on what type of image it is. It's like having multiple specialists and a smart dispatcher sending each case to the right expert.
Problem solvedVision models trained on one dataset often fail when you test them on different types of images or harder variations. This system can handle multiple image sources and difficulty levels without retraining, and you can update it just by changing prompts instead of retraining models.
- 💤Quiet2608.13283·Aug 13, 2026·~10 mincs.AI
Towards Context-Aware Clinical Motion Understanding in Daily Living at Home: Freezing of Gait Detection with Egocentric Vision
Vayalet Stefanova, Diwas Lamsal, Margot Genbrugge, Maxim Yudayev, +4
⭐ 0 stars / 0 repos📚 0 citesELI5A system uses video from a wrist-worn camera plus motion sensors to detect when Parkinson's patients involuntarily freeze while walking at home. The combination helps tell the difference between intentional stopping, reaching for something, and actual medical freezing.
Problem solvedDoctors need to detect freezing of gait in Parkinson's patients during real daily life, not just in controlled labs. Motion sensors alone can't distinguish between intentional pauses and pathological freezing, so adding egocentric vision provides the visual context needed for accurate diagnosis.
- 💤Quiet2608.12313·Aug 12, 2026·~12 mincs.CVcs.CL
AVA-Encoder: Towards Agent-Native Video Representation Learning
Chuyue Li, Jinpeng Yu, Haozhe Wang, Tian Xueyun, +6
⭐ 0 stars / 0 repos📚 0 citesELI5A system that converts videos into knowledge graphs (structured information with text, images, and audio linked together) that AI agents can understand and edit, then reconstructs the video back. It learns what makes good representations by having agents fix their mistakes using natural language feedback.
Problem solvedVideo AI systems today can't work with film-quality videos or easily edit them because they lack a structured, editable representation. Agents need a way to understand video content semantically and manipulate it precisely—not just process raw pixels.
- 💤Quiet2608.12262·Aug 12, 2026·~9 mincs.CVcs.AI
Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams
Weihao Bo, Shan Zhang, Yanpeng Sun, Jie Liu, +6
⭐ 0 stars / 0 repos📚 0 citesELI5A benchmark that tests how well AI models can understand scientific diagrams and convert them into code (like LaTeX), plus answer questions about them. Think of it like giving models a chemistry diagram and asking them to both write the code that draws it AND explain what's happening.
Problem solvedScientists need tools to automatically parse and edit scientific diagrams in documents, but existing AI models are bad at converting visual diagrams into executable code. This benchmark reveals which models can handle this task and where improvements are needed.
- 💤Quiet2608.07417·Aug 7, 2026·~11 mincs.CVcs.AI
I Seek You in Videos: Identity-Conditioned Queries for Person-Centric Video Reasoning
Shibo Gao, Chongxiao Wang, Chenglong Huang, Jie Ma, +9
⭐ 0 stars / 0 repos📚 0 citesELI5Given a photo of a person and a video, answer questions about what that specific person does in the video—like tracking them across scenes, understanding their actions, and reasoning about cause-and-effect. It's like pointing at someone in a photo and asking 'what is this person up to in this 5-minute clip?'
Problem solvedVideo AI systems today struggle to recognize and track the same person across complex, real-world videos with multiple people and scenes. This matters for security, content moderation, and any application that needs to understand what a specific individual is doing over time.
- 💤Quiet2608.07385·Aug 7, 2026·~12 mincs.LGcs.AIeess.AS
Omni-modal decomposition autoencoders learn full-stack wearable disentangled representations
Ioannis Ziogas, Ensieh Khazaei, Bilal Taha, Aamna Al Shehhi, +3
⭐ 0 stars / 0 repos📚 0 citesELI5A new AI system learns clean, interpretable patterns from many wearable sensors (like smartwatches with 30+ different sensors) by breaking down signals into separate layers, then uses those patterns for activity recognition, user identification, and generating synthetic sensor data—all in one lightweight model.
Problem solvedWearable devices have many different sensors but existing AI models either focus on one task (like activity recognition) or don't work efficiently on edge devices. This creates a brittle, inflexible setup requiring multiple separate models. OmniDecVAEs unifies activity recognition, identity recognition, and data generation in a single compact model that runs on-device.
- 💤Quiet2608.07349·Aug 7, 2026·~14 mincs.LG
Residual Algebra for Representation-Preserving Learning
Yao Wu
⭐ 0 stars / 0 repos📚 0 citesELI5When machine learning systems combine data from multiple sources, they usually just smoosh everything together and lose track of which source caused errors. This paper treats each data source as owning its own error signal, then chains together operations that respect or deliberately discard that ownership—like a relay race where each runner fixes only their leg's mistakes.
Problem solvedReal systems get features from different pipelines (databases, sensors, models) but standard approaches erase which source failed, making debugging and improvement hard. This method tracks error ownership through the pipeline, enabling better error correction and more interpretable learning—proven on stock trading data where it doubled returns.
- 💤Quiet2608.07282·Aug 7, 2026·~8 mincs.CLcs.CV
Gaze Behavior in Visual World Experiments Can be Modeled With Off-the-shelf Language-Vision Encoders
Rahul Murali Shankar, Titus von der Malsburg, Sebastian Padó
⭐ 0 stars / 0 repos📚 0 citesELI5Researchers used an off-the-shelf vision-language model (CLIP) to predict where people look when they see pictures and read text at the same time—and it matched human eye-tracking data without any special training.
Problem solvedMost AI research on human language processing ignores how people actually process language in real-world multimodal settings (like looking at pictures while reading). This work bridges that gap, showing that standard vision-language models can predict human visual attention patterns.
- 💤Quiet2608.05131·Aug 5, 2026·~10 mincs.CVcs.AI
OPD-V: Visual On-Policy Self-Distillation with Modality Balance
Aniri, Jinhe Bi, Peng Liao, Zengjie Jin, +4
⭐ 0 stars / 0 repos📚 0 citesELI5When AI models that understand both images and text learn from themselves, they often get lazy and ignore the visual info—relying too much on text instead. This paper teaches models to balance using both types of information by creating good and bad examples that show what happens when vision gets ignored, then uses that signal to fix the problem.
Problem solvedMultimodal models waste their ability to see images because text generation takes over during self-improvement training. This makes expensive visual reasoning features go unused, limiting how smart the model can actually become at understanding complex visual tasks.
- 💤Quiet2608.05103·Aug 5, 2026·~7 mincs.LGmath-phphysics.ao-ph
Multimodal Spatiotemporal Atmospheric Data Assimilation with Latent Flow-matching
Dibyajyoti Chakraborty, Romit Maulik
⭐ 0 stars / 0 repos📚 0 citesELI5Instead of traditional weather forecasting methods, this approach uses a video-generation model trained on past weather data to fill in missing parts of the atmosphere. You give it sparse observations (like a few weather stations), and it generates realistic complete weather states that are consistent over time.
Problem solvedWeather forecasting with incomplete observations is slow and rigid—traditional methods require expensive numerical simulations and can't easily handle different types of missing data. This method is faster, flexible across filtering/smoothing tasks, and works directly from sparse real-world observations without retraining.
- 💤Quiet2608.05026·Aug 5, 2026·~11 mincs.HCcs.AI
ArtAnno: Annotating Implicit Semantics in Artworks through LLM Agent-Driven Bidirectional Human-AI Augmentation
Xiaoyan Gu, Yifang Wang, Wenqing Zheng, Haozhong Liu, +5
⭐ 0 stars / 0 repos📚 0 citesELI5A system that helps art experts and novices label paintings by having AI suggest meanings while learning from human corrections in real-time, creating a two-way loop where humans teach the AI and the AI teaches back.
Problem solvedTagging artwork with deeper cultural and contextual meanings is slow, requires deep expertise, and existing AI tools force experts to manually fix mistakes rather than learning from corrections. This speeds up the process and lets knowledge accumulate.
- 💤Quiet2608.05000·Aug 5, 2026·~12 mincs.CVcs.LGcs.MM
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes
Junlin Han, Shengbang Tong, David Fan, Minghao Chen, +3
⭐ 0 stars / 0 repos📚 0 citesELI5Researchers figured out how to train AI models that work with both images and text together more efficiently. They discovered that mixing modalities early, using smart architecture choices, and matching data complexity to modality pairing saves 95% of compute while maintaining performance.
Problem solvedBuilding multimodal AI models is expensive and poorly understood—teams waste compute on bad architectural choices and timing decisions. This work provides proven recipes and principles to cut training costs dramatically while actually improving how well vision and language work together.
- 💤Quiet2608.04949·Aug 5, 2026·~9 mincs.CVcs.CLcs.IT
UG-UMRE: Uncertainty-Guided Modality Augmentation and Distributional Calibration for Unified Multimodal Relation Extraction
Bo Kong, Liruiz Jia, Yi Liang, Chao Liu, +4
⭐ 0 stars / 0 repos📚 0 citesELI5When you're trying to understand relationships between things in both images and text (like 'who is standing next to whom'), this system learns to ignore noisy, unreliable information and make the image and text data 'speak the same language' so they work better together.
Problem solvedCurrent multimodal relation extraction systems struggle with noisy data propagation and misalignment between how images and text represent information, leading to poor accuracy. This fixes both by measuring uncertainty to filter noise and forcing different data types into a shared statistical space.
- 💤Quiet2608.04926·Aug 5, 2026·~8 mincs.LGcs.AIcs.CL
Consistency-Driven Co-Evolution for Self-Supervised Cross-Representation Learning
Xuehang Guo, Pengyuan Li, Tom Hope, Tirthankar Ghosal, +2
⭐ 0 stars / 0 repos📚 0 citesELI5A system that learns to understand charts, tables, and code together by making sure they all agree with each other—like checking that a pie chart, a spreadsheet, and a Python script all describe the same data without needing labeled examples.
Problem solvedWorking with charts, tables, and code is expensive to label and hard to align because one chart could match many tables. This system learns alignments automatically by enforcing consistency across all three formats, saving annotation costs and improving accuracy.
- 💤Quiet2608.04902·Aug 5, 2026·~9 mincs.CVcs.LGcs.SD
Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation
Zehua Chen, Junyou Wang, Yuxuan Jiang, Zhenying Fang, +4
⭐ 0 stars / 0 repos📚 0 citesELI5This paper improves video-to-audio generation by focusing on what changes between consecutive frames rather than treating each frame independently. Instead of adding extra networks, they show that tracking motion and visual differences over time is the key insight that makes audio synthesis from video much better.
Problem solvedExisting video-to-audio systems either need extra supervision signals, additional neural networks, or complex reasoning pipelines to handle temporal information. This work shows you can get better results with simpler models by smartly using the differences between frames—making V2A generation faster and more practical to deploy.
- 💤Quiet2608.04010·Aug 4, 2026·~9 mincs.CVcs.CL
ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs
Yang Yang, Qinyu Zhao, Mouxiang Chen, Xiaohui Li, +4
⭐ 0 stars / 0 repos📚 0 citesELI5Instead of making multimodal AI models bigger or slower, this technique runs the vision and language parts in parallel branches that share the same underlying model, letting you allocate more computation to whichever part matters for your task.
Problem solvedMultimodal models waste resources with fixed computation splits between vision and language components, and scaling them up requires either huge memory or slow inference. This lets you tune the balance per task without growing the core model.
- 💤Quiet2608.03979·Aug 4, 2026·~10 mincs.CVcs.AI
Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent
Zhen Fang, Yu Zeng, Wenxuan Huang, Yiming Zhao, +16
⭐ 0 stars / 0 repos📚 0 citesELI5A video AI agent that watches continuous video clips and researches questions about them by actually using visual analysis tools and web search together—rather than just guessing from its training data or defaulting to text search.
Problem solvedCurrent multimodal AI agents skip over visual details in videos and rely on memorized facts instead of real tool use. This makes them bad at answering complex questions about video content that require cross-frame analysis and up-to-date information.
- 💤Quiet2608.03952·Aug 4, 2026·~13 mincs.AI
TACT: Taxonomy-Aligned Post-Training for Pedagogically Adaptive English Tutoring
Dongjie Yang, Siyan Lin, Leixian Shen, Rui Sheng, +2
⭐ 0 stars / 0 repos📚 0 citesELI5This paper teaches an AI tutor to help English learners by learning from real teaching strategies. Instead of just generating fluent responses, the tutor learns to pick the right teaching move (like asking clarifying questions vs. correcting errors) based on what the student does, similar to how human tutors adapt their approach.
Problem solvedMost AI English tutors sound natural but don't actually teach well—they don't adapt their strategy to each student's needs or difficulties. This work fixes that by grounding the tutor in real pedagogical principles from education research, making it actually effective at helping learners improve.
- 💤Quiet2608.02583·Aug 3, 2026·~11 mincs.CVcs.AIcs.CL
UEmbed: Unified Sparse and Dense Multimodal Embeddings
Tingyu Song, Mingxin Li, Yanzhao Zhang, Dingkun Long, +4
⭐ 0 stars / 0 repos📚 0 citesELI5A single AI model that creates two types of search fingerprints—one sparse (like specific keywords) and one dense (like semantic meaning)—for both text and images, so you can search across different modalities with one unified system instead of bolting together multiple models.
Problem solvedBuilding search systems currently requires separate models for text vs. images and often needs different models for keyword-based vs. semantic search. UEmbed solves both problems in one model, reducing complexity and making multimodal retrieval-augmented generation (RAG) systems easier to deploy and maintain.
- 💤Quiet2607.29602·Jul 31, 2026·~8 mincs.CLcs.AIcs.CV
FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models
Jeffrey M. Girard, Jason Z. Zheng, Jacqueline R. Vertino, Antony D'Avirro, +1
⭐ 0 stars / 0 repos📚 0 citesELI5A benchmark that tests whether AI models can tell if two people already know each other or are meeting for the first time, just by watching a 20-second video conversation. Turns out the best AI matches human accuracy, but they're using different strategies to get there.
Problem solvedThere's no standard way to evaluate whether AI can understand social familiarity from real interactions—a key part of human social intelligence. This benchmark lets researchers measure that capability across different AI models and modalities.
- 💤Quiet2607.28590·Jul 30, 2026·~12 mincs.CVcs.CL
VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation
Kangning Zhang, Yixing Li, Shuai Shao, Qingyao Li, +8
⭐ 0 stars / 0 repos📚 0 citesELI5When a teacher AI helps a student AI learn from images and text, the feedback mixes visual insights with language patterns. This method separates out just the visual part by asking: 'what would the teacher say with this image detail, versus without it?' and uses only that difference as training signal.
Problem solvedMultimodal distillation wastes teaching signal by mixing visual knowledge with linguistic biases and teacher quirks. This approach isolates pure visual evidence, so students learn what the images actually show rather than absorbing teacher artifacts.
- 💤Quiet2607.28580·Jul 30, 2026·~11 mincs.AI
DualG-MRAG: Decoupling Macro-Reasoning and Micro-Matching for Multimodal Retrieval-Augmented Generation
Jiacheng Tao, Qingyun Sun, Haonan Yuan, Ziwei Zhang, +1
⭐ 0 stars / 0 repos📚 0 citesELI5A system that retrieves information from both text and images to answer complex questions by using two separate graphs — one for understanding overall relationships, one for checking specific details — then guides generation by showing the model the exact reasoning path it used.
Problem solvedExisting systems that combine text and images for question-answering either miss connections across documents or get confused by too much visual detail. This approach separates high-level reasoning from detailed matching, reducing noise while keeping important evidence intact.
- 💤Quiet2607.28375·Jul 30, 2026·~10 mincs.AIcs.MM
HyperClaim: Fine-Grained Cross-Modal Hypergraph Reasoning for Video Misinformation Detection
Xiangbo Wang, Jiasheng Zhang, Xingtong Yu, Luoqiang Lei, +1
⭐ 0 stars / 0 repos📚 0 citesELI5A system that detects fake videos by carefully examining how words in claims connect to specific moments and images in the video, rather than treating the whole video as one blob. It uses a hypergraph (a network that can connect multiple things at once) to capture these fine-grained relationships across text and video.
Problem solvedVideo misinformation spreads fast and is hard to fact-check because existing detectors either look at videos too broadly or require external fact-checking tools. This method pinpoints exactly which frames and text snippets contradict a claim, making detection faster and more transparent.
- 💤Quiet2607.28374·Jul 30, 2026·~11 mincs.LG
LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger
Enjun Du, Hange Zhou, Chenxu Du, Siyi Liu, +3
⭐ 0 stars / 0 repos📚 0 citesELI5A system that forces AI agents answering visual questions to cite their sources like a lawyer would—every claim must trace back to what the agent actually saw or retrieved, preventing it from making stuff up or getting lucky with wrong-but-matching answers.
Problem solvedCurrent visual QA agents can produce right answers for wrong reasons (made-up details, language biases, or lucky errors), making it impossible to trust them for real applications. This system audits the entire reasoning chain to ensure answers are genuinely grounded in evidence.
- 💤Quiet2607.27109·Jul 29, 2026·~7 mincs.SDcs.AI
MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning
Weijie Wu, Junbo Li, Lin Li, Jun Fang, +1
⭐ 0 stars / 0 repos📚 0 citesELI5A benchmark that tests how well AI models describe audio by checking whether their descriptions cover important details (like what instruments are playing, the mood, the setting) and whether those details are actually accurate.
Problem solvedAudio AI models now generate longer, more detailed descriptions, but existing tests only measure overall quality—they can't tell you what information the model is missing or getting wrong. MMAC fixes this by examining 15 specific aspects of audio descriptions.
- 💤Quiet2607.26042·Jul 28, 2026·~8 mincs.CVcs.LG
VetClaw: An Edge-Cloud Multimodal Agentic System for Veterinary Disease Screening
Syed Mhamudul Hasan, Anas AlSobeh, Hussein Zangoti, Abdur R. Shahid
⭐ 0 stars / 0 repos📚 0 citesELI5A vet system that combines a camera and a chatbot-like AI to help identify animal diseases. You take a photo and describe symptoms, the AI analyzes both, and the system decides whether to flag it as urgent or safe to wait.
Problem solvedVets in remote or under-resourced areas can't easily screen animals for diseases early, and static AI models that just look at images aren't reliable enough for real diagnosis. This system adds workflow logic and safety checks to make AI predictions actually useful for decision-making.
- 💤Quiet2607.26023·Jul 28, 2026·~12 mincs.AI
CHARM: A Multimodal Graph Foundation Model with Hierarchical Context Modeling for Zero-Shot Transfer
Ankang Yang, Jitao Zhao, Di Jin, Yuxiao Huang, +1
⭐ 0 stars / 0 repos📚 0 citesELI5A model that learns from graphs where nodes have images, text, and other data types together, then applies that knowledge to brand new graphs without needing retraining. It organizes node information into meaningful chunks that capture relationships across different data types.
Problem solvedCurrent graph models either need retraining on new domains or only handle text. Real-world graphs mix images, text, and other media, but no foundation model could transfer knowledge across multimodal graphs without expensive labeled data and fine-tuning.
- 💤Quiet2607.25961·Jul 28, 2026·~11 mincs.CVcs.AI
Knowledge-Guided Multimodal Reasoning over Interacting Streams for Video-Level Ambivalence and Hesitancy Recognition
Podakanti Satyajith Chary, Barath Parthiban, Pranesh Velmurugan, Adeeba Khan, +1
⭐ 0 stars / 0 repos📚 0 citesELI5A system that detects when people are conflicted or hesitant (shown through mismatches in their facial expressions, tone, words, and body language) by watching video and having an AI reason through the mixed signals to decide if someone is ambivalent.
Problem solvedHealth behavior change programs need to identify when people are genuinely conflicted or hesitant so they can intervene early—but detecting this requires reading subtle disagreements across how someone looks, sounds, and speaks simultaneously, which humans and basic AI struggle with.
- 💤Quiet2607.24743·Jul 27, 2026·~13 mincs.CVcs.AIcs.CL
ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding
Hangjie Yuan, Yichen Qian, Zhiwei Tang, Xianzhe Xu, +20
⭐ 0 stars / 0 repos📚 0 citesELI5A specialized AI system that looks at medical images (X-rays, CT scans, 3D volumes) and explains what it sees in clinical language, much like a radiologist would. It handles both flat 2D images and complex 3D scans in one unified model.
Problem solvedExisting medical AI struggles to process diverse image types together and evaluate results the way actual doctors do. ClinFusion lets hospitals deploy a single model for medical imaging that radiologists can trust, with evaluation metrics that match clinical reality.