On Explaining with Attention Matrices
Fuente:
arXiv
Guardado en:
| Autores principales: | Naim, Omar, Asher, Nicholas |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Neural Machine Translation for Malayalam Paraphrase Generation
por: Varghese, Christeena, et al.
Publicado: (2024)
por: Varghese, Christeena, et al.
Publicado: (2024)
Is Our Chatbot Telling Lies? Assessing Correctness of an LLM-based Dutch Support Chatbot
por: Lassche, Herman, et al.
Publicado: (2024)
por: Lassche, Herman, et al.
Publicado: (2024)
Predictive Simultaneous Interpretation: Harnessing Large Language Models for Democratizing Real-Time Multilingual Communication
por: Iida, Kurando, et al.
Publicado: (2024)
por: Iida, Kurando, et al.
Publicado: (2024)
UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning
por: Ovcharov, Volodymyr
Publicado: (2026)
por: Ovcharov, Volodymyr
Publicado: (2026)
Semantic Decomposition and Selective Context Filtering -- Text Processing Techniques for Context-Aware NLP-Based Systems
por: Villardar, Karl John
Publicado: (2025)
por: Villardar, Karl John
Publicado: (2025)
Breaking the HISCO Barrier: Automatic Occupational Standardization with OccCANINE
por: Dahl, Christian Møller, et al.
Publicado: (2024)
por: Dahl, Christian Møller, et al.
Publicado: (2024)
Classification of descriptions and summary using multiple passes of statistical and natural language toolkits
por: Banthia, Saumya, et al.
Publicado: (2020)
por: Banthia, Saumya, et al.
Publicado: (2020)
Benchmarking quantized LLaMa-based models on the Brazilian Secondary School Exam
por: Santos, Matheus L. O., et al.
Publicado: (2023)
por: Santos, Matheus L. O., et al.
Publicado: (2023)
Strong hallucinations from negation and how to fix them
por: Asher, Nicholas, et al.
Publicado: (2024)
por: Asher, Nicholas, et al.
Publicado: (2024)
Quo Vadis ChatGPT? From Large Language Models to Large Knowledge Models
por: Venkatasubramanian, Venkat, et al.
Publicado: (2024)
por: Venkatasubramanian, Venkat, et al.
Publicado: (2024)
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization
por: Khanna, Danush, et al.
Publicado: (2025)
por: Khanna, Danush, et al.
Publicado: (2025)
OpenAI Cribbed Our Tax Example, But Can GPT-4 Really Do Tax?
por: Blair-Stanek, Andrew, et al.
Publicado: (2023)
por: Blair-Stanek, Andrew, et al.
Publicado: (2023)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
por: Saji, Alan, et al.
Publicado: (2025)
por: Saji, Alan, et al.
Publicado: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
por: Peters, Sydney, et al.
Publicado: (2025)
por: Peters, Sydney, et al.
Publicado: (2025)
KisMATH: Do LLMs Have Knowledge of Implicit Structures in Mathematical Reasoning?
por: Saha, Soumadeep, et al.
Publicado: (2025)
por: Saha, Soumadeep, et al.
Publicado: (2025)
Semantic Commit: Helping Users Update Intent Specifications for AI Memory at Scale
por: Vaithilingam, Priyan, et al.
Publicado: (2025)
por: Vaithilingam, Priyan, et al.
Publicado: (2025)
Text2Model: Modeling Copilots for Text-to-Model Translation
por: Kadioglu, Serdar, et al.
Publicado: (2026)
por: Kadioglu, Serdar, et al.
Publicado: (2026)
Low-Resource Court Judgment Summarization for Common Law Systems
por: Liu, Shuaiqi, et al.
Publicado: (2024)
por: Liu, Shuaiqi, et al.
Publicado: (2024)
Synergy of Large Language Model and Model Driven Engineering for Automated Development of Centralized Vehicular Systems
por: Petrovic, Nenad, et al.
Publicado: (2024)
por: Petrovic, Nenad, et al.
Publicado: (2024)
Towards Single-System Illusion in Software-Defined Vehicles -- Automated, AI-Powered Workflow
por: Lebioda, Krzysztof, et al.
Publicado: (2024)
por: Lebioda, Krzysztof, et al.
Publicado: (2024)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
por: Oketunji, Abiodun Finbarrs
Publicado: (2023)
por: Oketunji, Abiodun Finbarrs
Publicado: (2023)
Word Importance Explains How Prompts Affect Language Model Outputs
por: Hackmann, Stefan, et al.
Publicado: (2024)
por: Hackmann, Stefan, et al.
Publicado: (2024)
Cost-Effective Attention Mechanisms for Low Resource Settings: Necessity & Sufficiency of Linear Transformations
por: Hosseini, Peyman, et al.
Publicado: (2024)
por: Hosseini, Peyman, et al.
Publicado: (2024)
Layer-Wise Quantization: A Pragmatic and Effective Method for Quantizing LLMs Beyond Integer Bit-Levels
por: Dumitru, Razvan-Gabriel, et al.
Publicado: (2024)
por: Dumitru, Razvan-Gabriel, et al.
Publicado: (2024)
Change Is the Only Constant: Dynamic LLM Slicing based on Layer Redundancy
por: Dumitru, Razvan-Gabriel, et al.
Publicado: (2024)
por: Dumitru, Razvan-Gabriel, et al.
Publicado: (2024)
ALISON: Fast and Effective Stylometric Authorship Obfuscation
por: Xing, Eric, et al.
Publicado: (2024)
por: Xing, Eric, et al.
Publicado: (2024)
ACCORD: Closing the Commonsense Measurability Gap
por: Roewer-Després, François, et al.
Publicado: (2024)
por: Roewer-Després, François, et al.
Publicado: (2024)
Enhancing Transformer RNNs with Multiple Temporal Perspectives
por: Dumitru, Razvan-Gabriel, et al.
Publicado: (2024)
por: Dumitru, Razvan-Gabriel, et al.
Publicado: (2024)
ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models
por: Dumitru, Razvan-Gabriel, et al.
Publicado: (2025)
por: Dumitru, Razvan-Gabriel, et al.
Publicado: (2025)
Key-Value Means: Transformers with Expandable Block-Recurrent Compressed Memory
por: Goldstein, Daniel, et al.
Publicado: (2026)
por: Goldstein, Daniel, et al.
Publicado: (2026)
Chatbots put to the test in math and logic problems: A preliminary comparison and assessment of ChatGPT-3.5, ChatGPT-4, and Google Bard
por: Plevris, Vagelis, et al.
Publicado: (2023)
por: Plevris, Vagelis, et al.
Publicado: (2023)
Prompt Tuned Embedding Classification for Multi-Label Industry Sector Allocation
por: Buchner, Valentin Leonhard, et al.
Publicado: (2023)
por: Buchner, Valentin Leonhard, et al.
Publicado: (2023)
NRR-Core: Non-Resolution Reasoning as a Computational Framework for Contextual Identity and Ambiguity Preservation
por: Saito, Kei
Publicado: (2025)
por: Saito, Kei
Publicado: (2025)
Amazon Nova AI Challenge -- Trusted AI: Advancing secure, AI-assisted software development
por: Sahai, Sattvik, et al.
Publicado: (2025)
por: Sahai, Sattvik, et al.
Publicado: (2025)
CopySpec: Accelerating LLMs with Speculative Copy-and-Paste Without Compromising Quality
por: Dumitru, Razvan-Gabriel, et al.
Publicado: (2025)
por: Dumitru, Razvan-Gabriel, et al.
Publicado: (2025)
NRR-Phi: Text-to-State Mapping for Ambiguity Preservation in LLM Inference
por: Saito, Kei
Publicado: (2026)
por: Saito, Kei
Publicado: (2026)
Next Token Prediction Is a Dead End for Creativity
por: Olatunji, Ibukun, et al.
Publicado: (2025)
por: Olatunji, Ibukun, et al.
Publicado: (2025)
RWKV-7 "Goose" with Expressive Dynamic State Evolution
por: Peng, Bo, et al.
Publicado: (2025)
por: Peng, Bo, et al.
Publicado: (2025)
BabyReasoningBench: Generating Developmentally-Inspired Reasoning Tasks for Evaluating Baby Language Models
por: Dhole, Kaustubh D.
Publicado: (2026)
por: Dhole, Kaustubh D.
Publicado: (2026)
Robustness of Large Language Models to Perturbations in Text
por: Singh, Ayush, et al.
Publicado: (2024)
por: Singh, Ayush, et al.
Publicado: (2024)
Ejemplares similares
-
Neural Machine Translation for Malayalam Paraphrase Generation
por: Varghese, Christeena, et al.
Publicado: (2024) -
Is Our Chatbot Telling Lies? Assessing Correctness of an LLM-based Dutch Support Chatbot
por: Lassche, Herman, et al.
Publicado: (2024) -
Predictive Simultaneous Interpretation: Harnessing Large Language Models for Democratizing Real-Time Multilingual Communication
por: Iida, Kurando, et al.
Publicado: (2024) -
UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning
por: Ovcharov, Volodymyr
Publicado: (2026) -
Semantic Decomposition and Selective Context Filtering -- Text Processing Techniques for Context-Aware NLP-Based Systems
por: Villardar, Karl John
Publicado: (2025)