Energy-Based Transformers are Scalable Learners and Thinkers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gladstone, Alexi, Nanduru, Ganesh, Islam, Md Mofijul, Han, Peixuan, Ha, Hyeonjeong, Chadha, Aman, Du, Yilun, Ji, Heng, Li, Jundong, Iqbal, Tariq |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Cognitively Inspired Energy-Based World Models
von: Gladstone, Alexi, et al.
Veröffentlicht: (2024)
von: Gladstone, Alexi, et al.
Veröffentlicht: (2024)
Embodied Referring Expression Comprehension in Human-Robot Interaction
von: Islam, Md Mofijul, et al.
Veröffentlicht: (2025)
von: Islam, Md Mofijul, et al.
Veröffentlicht: (2025)
Assessing the Creativity of Large Language Models: Testing, Limits, and New Frontiers
von: Schapiro, Samuel, et al.
Veröffentlicht: (2026)
von: Schapiro, Samuel, et al.
Veröffentlicht: (2026)
DM-Codec: Distilling Multimodal Representations for Speech Tokenization
von: Ahasan, Md Mubtasim, et al.
Veröffentlicht: (2024)
von: Ahasan, Md Mubtasim, et al.
Veröffentlicht: (2024)
FuseCodec: Semantic-Contextual Fusion and Supervision for Neural Codecs
von: Ahasan, Md Mubtasim, et al.
Veröffentlicht: (2025)
von: Ahasan, Md Mubtasim, et al.
Veröffentlicht: (2025)
Breaking Down the Defenses: A Comparative Survey of Attacks on Large Language Models
von: Chowdhury, Arijit Ghosh, et al.
Veröffentlicht: (2024)
von: Chowdhury, Arijit Ghosh, et al.
Veröffentlicht: (2024)
SloMo-Fast: Slow-Momentum and Fast-Adaptive Teachers for Source-Free Continual Test-Time Adaptation
von: Iftee, Md Akil Raihan, et al.
Veröffentlicht: (2025)
von: Iftee, Md Akil Raihan, et al.
Veröffentlicht: (2025)
MathMist: A Parallel Multilingual Benchmark Dataset for Mathematical Problem Solving and Reasoning
von: Sobhani, Mahbub E, et al.
Veröffentlicht: (2025)
von: Sobhani, Mahbub E, et al.
Veröffentlicht: (2025)
From Fog to Failure: The Unintended Consequences of Dehazing on Object Detection in Clear Images
von: Kumar, Ashutosh, et al.
Veröffentlicht: (2025)
von: Kumar, Ashutosh, et al.
Veröffentlicht: (2025)
Generative Data Augmentation using LLMs improves Distributional Robustness in Question Answering
von: Chowdhury, Arijit Ghosh, et al.
Veröffentlicht: (2023)
von: Chowdhury, Arijit Ghosh, et al.
Veröffentlicht: (2023)
Ready to Translate, Not to Represent? Bias and Performance Gaps in Multilingual LLMs Across Language Families and Domains
von: Sayeedi, Md. Faiyaz Abdullah, et al.
Veröffentlicht: (2025)
von: Sayeedi, Md. Faiyaz Abdullah, et al.
Veröffentlicht: (2025)
Breaking Language Barriers: A Question Answering Dataset for Hindi and Marathi
von: Sabane, Maithili, et al.
Veröffentlicht: (2023)
von: Sabane, Maithili, et al.
Veröffentlicht: (2023)
Unboxing Occupational Bias: Grounded Debiasing of LLMs with U.S. Labor Data
von: Gorti, Atmika, et al.
Veröffentlicht: (2024)
von: Gorti, Atmika, et al.
Veröffentlicht: (2024)
Guiding Vision-Language Model Selection for Visual Question-Answering Across Tasks, Domains, and Knowledge Types
von: Sinha, Neelabh, et al.
Veröffentlicht: (2024)
von: Sinha, Neelabh, et al.
Veröffentlicht: (2024)
SOLO: A Single Transformer for Scalable Vision-Language Modeling
von: Chen, Yangyi, et al.
Veröffentlicht: (2024)
von: Chen, Yangyi, et al.
Veröffentlicht: (2024)
Simulating Meaning, Nevermore! Introducing ICR: A Semiotic-Hermeneutic Metric for Evaluating Meaning in LLM Text Summaries
von: Perez, Natalie, et al.
Veröffentlicht: (2026)
von: Perez, Natalie, et al.
Veröffentlicht: (2026)
Born With a Silver Spoon? Investigating Socioeconomic Bias in Large Language Models
von: Singh, Smriti, et al.
Veröffentlicht: (2024)
von: Singh, Smriti, et al.
Veröffentlicht: (2024)
Mental Health Equity in LLMs: Leveraging Multi-Hop Question Answering to Detect Amplified and Silenced Perspectives
von: Haider, Batool, et al.
Veröffentlicht: (2025)
von: Haider, Batool, et al.
Veröffentlicht: (2025)
Density Adaptive Attention is All You Need: Robust Parameter-Efficient Fine-Tuning Across Multiple Modalities
von: Ioannides, Georgios, et al.
Veröffentlicht: (2024)
von: Ioannides, Georgios, et al.
Veröffentlicht: (2024)
The Reasoning Trap -- Logical Reasoning as a Mechanistic Pathway to Situational Awareness
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
Are Small Language Models Ready to Compete with Large Language Models for Practical Applications?
von: Sinha, Neelabh, et al.
Veröffentlicht: (2024)
von: Sinha, Neelabh, et al.
Veröffentlicht: (2024)
Must Read: A Comprehensive Survey of Computational Persuasion
von: Bozdag, Nimet Beyza, et al.
Veröffentlicht: (2025)
von: Bozdag, Nimet Beyza, et al.
Veröffentlicht: (2025)
Ink Spiral: Symbolic Transformation from The Thinker to the Four Gentlemen
von: Peng, Lingyu, et al.
Veröffentlicht: (2026)
von: Peng, Lingyu, et al.
Veröffentlicht: (2026)
Transformation of Biological Networks into Images via Semantic Cartography for Visual Interpretation and Scalable Deep Analysis
von: Mostafa, Sakib, et al.
Veröffentlicht: (2025)
von: Mostafa, Sakib, et al.
Veröffentlicht: (2025)
Dial E for Ethical Enforcement: institutional VETO power as a governance primitive
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
How Well Do LLMs Represent Values Across Cultures? Empirical Analysis of LLM Responses Based on Hofstede Cultural Dimensions
von: Kharchenko, Julia, et al.
Veröffentlicht: (2024)
von: Kharchenko, Julia, et al.
Veröffentlicht: (2024)
I Think, Therefore I Am Under-Qualified? A Benchmark for Evaluating Linguistic Shibboleth Detection in LLM Hiring Evaluations
von: Kharchenko, Julia, et al.
Veröffentlicht: (2025)
von: Kharchenko, Julia, et al.
Veröffentlicht: (2025)
Constructive Distortion: Improving MLLMs with Attention-Guided Image Warping
von: Dalal, Dwip, et al.
Veröffentlicht: (2025)
von: Dalal, Dwip, et al.
Veröffentlicht: (2025)
DecisionFlow: Advancing Large Language Model as Principled Decision Maker
von: Chen, Xiusi, et al.
Veröffentlicht: (2025)
von: Chen, Xiusi, et al.
Veröffentlicht: (2025)
BEAT: Visual Backdoor Attacks on VLM-based Embodied Agents via Contrastive Trigger Learning
von: Zhan, Qiusi, et al.
Veröffentlicht: (2025)
von: Zhan, Qiusi, et al.
Veröffentlicht: (2025)
Equilibrium Matching: Generative Modeling with Implicit Energy-Based Models
von: Wang, Runqian, et al.
Veröffentlicht: (2025)
von: Wang, Runqian, et al.
Veröffentlicht: (2025)
NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding
von: Ha, Hyeonjeong, et al.
Veröffentlicht: (2026)
von: Ha, Hyeonjeong, et al.
Veröffentlicht: (2026)
Perception-Aware Policy Optimization for Multimodal Reasoning
von: Wang, Zhenhailong, et al.
Veröffentlicht: (2025)
von: Wang, Zhenhailong, et al.
Veröffentlicht: (2025)
Siamese Vision Transformers are Scalable Audio-visual Learners
von: Lin, Yan-Bo, et al.
Veröffentlicht: (2024)
von: Lin, Yan-Bo, et al.
Veröffentlicht: (2024)
The What, Why, and How of Context Length Extension Techniques in Large Language Models -- A Detailed Survey
von: Pawar, Saurav, et al.
Veröffentlicht: (2024)
von: Pawar, Saurav, et al.
Veröffentlicht: (2024)
SPIE: Semantic and Structural Post-Training of Image Editing Diffusion Models with AI feedback
von: Benarous, Elior, et al.
Veröffentlicht: (2025)
von: Benarous, Elior, et al.
Veröffentlicht: (2025)
Human-Readable Adversarial Prompts: An Investigation into LLM Vulnerabilities Using Situational Context
von: Das, Nilanjana, et al.
Veröffentlicht: (2024)
von: Das, Nilanjana, et al.
Veröffentlicht: (2024)
I Can't Believe It's Not Robust: Catastrophic Collapse of Safety Classifiers under Embedding Drift
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
How Culturally Aware are Vision-Language Models?
von: Burda-Lassen, Olena, et al.
Veröffentlicht: (2024)
von: Burda-Lassen, Olena, et al.
Veröffentlicht: (2024)
DocSplit: A Comprehensive Benchmark Dataset and Evaluation Approach for Document Packet Recognition and Splitting
von: Islam, Md Mofijul, et al.
Veröffentlicht: (2026)
von: Islam, Md Mofijul, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Cognitively Inspired Energy-Based World Models
von: Gladstone, Alexi, et al.
Veröffentlicht: (2024) -
Embodied Referring Expression Comprehension in Human-Robot Interaction
von: Islam, Md Mofijul, et al.
Veröffentlicht: (2025) -
Assessing the Creativity of Large Language Models: Testing, Limits, and New Frontiers
von: Schapiro, Samuel, et al.
Veröffentlicht: (2026) -
DM-Codec: Distilling Multimodal Representations for Speech Tokenization
von: Ahasan, Md Mubtasim, et al.
Veröffentlicht: (2024) -
FuseCodec: Semantic-Contextual Fusion and Supervision for Neural Codecs
von: Ahasan, Md Mubtasim, et al.
Veröffentlicht: (2025)