Can Argus Judge Them All? Comparing VLMs Across Domains
Fuente:
arXiv
Saved in:
| Main Authors: | Joshi, Harsh, Kashyap, Gautam Siddharth, Ali, Rafiq, Shabbir, Ebad, Jain, Niharika, Jain, Sarthak, Gao, Jiechao, Naseem, Usman |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Revealing the Truth with ConLLM for Detecting Multi-Modal Deepfakes
by: Kashyap, Gautam Siddharth, et al.
Published: (2026)
by: Kashyap, Gautam Siddharth, et al.
Published: (2026)
Can Large Language Models Make Everyone Happy?
by: Naseem, Usman, et al.
Published: (2026)
by: Naseem, Usman, et al.
Published: (2026)
LLMs on a Budget? Say HOLA
by: Siddiqui, Zohaib Hasan, et al.
Published: (2025)
by: Siddiqui, Zohaib Hasan, et al.
Published: (2025)
Truth, Trust, and Trouble: Medical AI on the Edge
by: Azeez, Mohammad Anas, et al.
Published: (2025)
by: Azeez, Mohammad Anas, et al.
Published: (2025)
Are Large Language Models Economically Viable for Industry Deployment?
by: Mohammad, Abdullah, et al.
Published: (2026)
by: Mohammad, Abdullah, et al.
Published: (2026)
Do Large Language Models Reflect Demographic Pluralism in Safety?
by: Naseem, Usman, et al.
Published: (2026)
by: Naseem, Usman, et al.
Published: (2026)
Are Aligned Large Language Models Still Misaligned?
by: Naseem, Usman, et al.
Published: (2026)
by: Naseem, Usman, et al.
Published: (2026)
ChildGuard: A Specialized Dataset for Combatting Child-Targeted Hate Speech
by: Kashyap, Gautam Siddharth, et al.
Published: (2025)
by: Kashyap, Gautam Siddharth, et al.
Published: (2025)
How Can Multimodal Remote Sensing Datasets Transform Classification via SpatialNet-ViT?
by: Kashyap, Gautam Siddharth, et al.
Published: (2025)
by: Kashyap, Gautam Siddharth, et al.
Published: (2025)
Do Clinical Question Answering Systems Really Need Specialised Medical Fine Tuning?
by: Ray, Sushant Kumar, et al.
Published: (2026)
by: Ray, Sushant Kumar, et al.
Published: (2026)
We Think, Therefore We Align LLMs to Helpful, Harmless and Honest Before They Go Wrong
by: Kashyap, Gautam Siddharth, et al.
Published: (2025)
by: Kashyap, Gautam Siddharth, et al.
Published: (2025)
AlignCultura: Towards Culturally Aligned Large Language Models?
by: Kashyap, Gautam Siddharth, et al.
Published: (2026)
by: Kashyap, Gautam Siddharth, et al.
Published: (2026)
Too Helpful, Too Harmless, Too Honest or Just Right?
by: Kashyap, Gautam Siddharth, et al.
Published: (2025)
by: Kashyap, Gautam Siddharth, et al.
Published: (2025)
When the Model Said 'No Comment', We Knew Helpfulness Was Dead, Honesty Was Alive, and Safety Was Terrified
by: Kashyap, Gautam Siddharth, et al.
Published: (2026)
by: Kashyap, Gautam Siddharth, et al.
Published: (2026)
They Said Memes Were Harmless-We Found the Ones That Hurt: Decoding Jokes, Symbols, and Cultural References
by: Tripathi, Sahil, et al.
Published: (2026)
by: Tripathi, Sahil, et al.
Published: (2026)
Beyond Retrieval: Joint Supervision and Multimodal Document Ranking for Textbook Question Answering
by: Alawwad, Hessa, et al.
Published: (2025)
by: Alawwad, Hessa, et al.
Published: (2025)
With Argus Eyes: Assessing Retrieval Gaps via Uncertainty Scoring to Detect and Remedy Retrieval Blind Spots
by: Taghavi, Zeinab Sadat, et al.
Published: (2026)
by: Taghavi, Zeinab Sadat, et al.
Published: (2026)
CUE-R: Beyond the Final Answer in Retrieval-Augmented Generation
by: Jain, Siddharth, et al.
Published: (2026)
by: Jain, Siddharth, et al.
Published: (2026)
LLM Agents Improve Semantic Code Search
by: Jain, Sarthak, et al.
Published: (2024)
by: Jain, Sarthak, et al.
Published: (2024)
Argus: Evidence Assembly for Scalable Deep Research Agents
by: Zhang, Zhen, et al.
Published: (2026)
by: Zhang, Zhen, et al.
Published: (2026)
Structural Representation Learning and Disentanglement for Evidential Chinese Patent Approval Prediction
by: Shan, Jinzhi, et al.
Published: (2024)
by: Shan, Jinzhi, et al.
Published: (2024)
Evaluating Multimodal Large Language Models on Educational Textbook Question Answering
by: Alawwad, Hessa A., et al.
Published: (2025)
by: Alawwad, Hessa A., et al.
Published: (2025)
FastLane: Efficient Routed Systems for Late-Interaction Retrieval
by: Kumar, Ramnath, et al.
Published: (2026)
by: Kumar, Ramnath, et al.
Published: (2026)
Revisiting Document-Level Relation Extraction with Context-Guided Link Prediction
by: Jain, Monika, et al.
Published: (2024)
by: Jain, Monika, et al.
Published: (2024)
Recall Them All: Retrieval-Augmented Language Models for Long Object List Extraction from Long Documents
by: Singhania, Sneha, et al.
Published: (2024)
by: Singhania, Sneha, et al.
Published: (2024)
Evaluating the Efficacy of Open-Source LLMs in Enterprise-Specific RAG Systems: A Comparative Study of Performance and Scalability
by: B, Gautam, et al.
Published: (2024)
by: B, Gautam, et al.
Published: (2024)
MSynFD: Multi-hop Syntax aware Fake News Detection
by: Xiao, Liang, et al.
Published: (2024)
by: Xiao, Liang, et al.
Published: (2024)
One for Dozens: Adaptive REcommendation for All Domains with Counterfactual Augmentation
by: Luo, Huishi, et al.
Published: (2024)
by: Luo, Huishi, et al.
Published: (2024)
First Steps, Lasting Impact: Platform-Aware Forensics for the Next Generation of Analysts
by: Jain, Vinayak, et al.
Published: (2026)
by: Jain, Vinayak, et al.
Published: (2026)
MFBE: Leveraging Multi-Field Information of FAQs for Efficient Dense Retrieval
by: Banerjee, Debopriyo, et al.
Published: (2023)
by: Banerjee, Debopriyo, et al.
Published: (2023)
Judging the Judges: A Collection of LLM-Generated Relevance Judgements
by: Rahmani, Hossein A., et al.
Published: (2025)
by: Rahmani, Hossein A., et al.
Published: (2025)
All Roads Lead to Rome: Unveiling the Trajectory of Recommender Systems Across the LLM Era
by: Chen, Bo, et al.
Published: (2024)
by: Chen, Bo, et al.
Published: (2024)
Evaluating the Robustness of Dense Retrievers in Interdisciplinary Domains
by: Chaturvedi, Sarthak, et al.
Published: (2025)
by: Chaturvedi, Sarthak, et al.
Published: (2025)
World Food Atlas Project
by: Rostami, Ali, et al.
Published: (2025)
by: Rostami, Ali, et al.
Published: (2025)
Vote-in-Context: Turning VLMs into Zero-Shot Rank Fusers
by: Eltahir, Mohamed, et al.
Published: (2025)
by: Eltahir, Mohamed, et al.
Published: (2025)
From Facts to Conclusions : Integrating Deductive Reasoning in Retrieval-Augmented LLMs
by: Mishra, Shubham, et al.
Published: (2025)
by: Mishra, Shubham, et al.
Published: (2025)
One Model for All: Large Language Models are Domain-Agnostic Recommendation Systems
by: Tang, Zuoli, et al.
Published: (2023)
by: Tang, Zuoli, et al.
Published: (2023)
Impacts of Mainstream-Driven Algorithms on Recommendations for Children Across Domains: A Reproducibility Study
by: Ungruh, Robin, et al.
Published: (2025)
by: Ungruh, Robin, et al.
Published: (2025)
Enhancing Research Information Systems with Identification of Domain Experts
by: Shahi, Gautam Kishore, et al.
Published: (2024)
by: Shahi, Gautam Kishore, et al.
Published: (2024)
MaiBERT: A Pre-training Corpus and Language Model for Low-Resourced Maithili Language
by: Yadav, Sumit, et al.
Published: (2025)
by: Yadav, Sumit, et al.
Published: (2025)
Similar Items
-
Revealing the Truth with ConLLM for Detecting Multi-Modal Deepfakes
by: Kashyap, Gautam Siddharth, et al.
Published: (2026) -
Can Large Language Models Make Everyone Happy?
by: Naseem, Usman, et al.
Published: (2026) -
LLMs on a Budget? Say HOLA
by: Siddiqui, Zohaib Hasan, et al.
Published: (2025) -
Truth, Trust, and Trouble: Medical AI on the Edge
by: Azeez, Mohammad Anas, et al.
Published: (2025) -
Are Large Language Models Economically Viable for Industry Deployment?
by: Mohammad, Abdullah, et al.
Published: (2026)