ACCORD: Closing the Commonsense Measurability Gap
Fuente:
arXiv
Guardado en:
| Autores principales: | Roewer-Després, François, Feng, Jinyue, Zhu, Zining, Rudzicz, Frank |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Graph Language Models
por: Plenz, Moritz, et al.
Publicado: (2024)
por: Plenz, Moritz, et al.
Publicado: (2024)
The Drill-Down and Fabricate Test (DDFT): A Protocol for Measuring Epistemic Robustness in Language Models
por: Baxi, Rahul
Publicado: (2025)
por: Baxi, Rahul
Publicado: (2025)
ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models
por: Dumitru, Razvan-Gabriel, et al.
Publicado: (2025)
por: Dumitru, Razvan-Gabriel, et al.
Publicado: (2025)
Layer-Wise Quantization: A Pragmatic and Effective Method for Quantizing LLMs Beyond Integer Bit-Levels
por: Dumitru, Razvan-Gabriel, et al.
Publicado: (2024)
por: Dumitru, Razvan-Gabriel, et al.
Publicado: (2024)
Key-Value Means: Transformers with Expandable Block-Recurrent Compressed Memory
por: Goldstein, Daniel, et al.
Publicado: (2026)
por: Goldstein, Daniel, et al.
Publicado: (2026)
Change Is the Only Constant: Dynamic LLM Slicing based on Layer Redundancy
por: Dumitru, Razvan-Gabriel, et al.
Publicado: (2024)
por: Dumitru, Razvan-Gabriel, et al.
Publicado: (2024)
NRR-Core: Non-Resolution Reasoning as a Computational Framework for Contextual Identity and Ambiguity Preservation
por: Saito, Kei
Publicado: (2025)
por: Saito, Kei
Publicado: (2025)
ALISON: Fast and Effective Stylometric Authorship Obfuscation
por: Xing, Eric, et al.
Publicado: (2024)
por: Xing, Eric, et al.
Publicado: (2024)
CopySpec: Accelerating LLMs with Speculative Copy-and-Paste Without Compromising Quality
por: Dumitru, Razvan-Gabriel, et al.
Publicado: (2025)
por: Dumitru, Razvan-Gabriel, et al.
Publicado: (2025)
NRR-Phi: Text-to-State Mapping for Ambiguity Preservation in LLM Inference
por: Saito, Kei
Publicado: (2026)
por: Saito, Kei
Publicado: (2026)
RWKV-7 "Goose" with Expressive Dynamic State Evolution
por: Peng, Bo, et al.
Publicado: (2025)
por: Peng, Bo, et al.
Publicado: (2025)
Enhancing Transformer RNNs with Multiple Temporal Perspectives
por: Dumitru, Razvan-Gabriel, et al.
Publicado: (2024)
por: Dumitru, Razvan-Gabriel, et al.
Publicado: (2024)
Social Learning through Interactions with Other Agents: A Survey
por: Hillier, Dylan, et al.
Publicado: (2024)
por: Hillier, Dylan, et al.
Publicado: (2024)
Perturbation Dose Responses in Recursive LLM Loops: Raw Switching, Stochastic Floors, and Persistent Escape under Append, Replace, and Dialog Updates
por: Kaplanski, Pawel
Publicado: (2026)
por: Kaplanski, Pawel
Publicado: (2026)
Incentives or Ontology? A Structural Rebuttal to OpenAI's Hallucination Thesis
por: Ackermann, Richard, et al.
Publicado: (2025)
por: Ackermann, Richard, et al.
Publicado: (2025)
Behavioural vs. Representational Systematicity in End-to-End Models: An Opinionated Survey
por: Vegner, Ivan, et al.
Publicado: (2025)
por: Vegner, Ivan, et al.
Publicado: (2025)
Quo Vadis ChatGPT? From Large Language Models to Large Knowledge Models
por: Venkatasubramanian, Venkat, et al.
Publicado: (2024)
por: Venkatasubramanian, Venkat, et al.
Publicado: (2024)
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization
por: Khanna, Danush, et al.
Publicado: (2025)
por: Khanna, Danush, et al.
Publicado: (2025)
OpenAI Cribbed Our Tax Example, But Can GPT-4 Really Do Tax?
por: Blair-Stanek, Andrew, et al.
Publicado: (2023)
por: Blair-Stanek, Andrew, et al.
Publicado: (2023)
Diverse LLMs or Diverse Question Interpretations? That is the Ensembling Question
por: Rosales, Rafael, et al.
Publicado: (2025)
por: Rosales, Rafael, et al.
Publicado: (2025)
Holistic Audit Dataset Generation for LLM Unlearning via Knowledge Graph Traversal and Redundancy Removal
por: Jiang, Weipeng, et al.
Publicado: (2025)
por: Jiang, Weipeng, et al.
Publicado: (2025)
Pareto-Optimized Open-Source LLMs for Healthcare via Context Retrieval
por: Bayarri-Planas, Jordi, et al.
Publicado: (2024)
por: Bayarri-Planas, Jordi, et al.
Publicado: (2024)
Aspect-Based Sentiment Analysis for Future Tourism Experiences: A BERT-MoE Framework for Persian User Reviews
por: Taskooh, Hamidreza Kazemi, et al.
Publicado: (2026)
por: Taskooh, Hamidreza Kazemi, et al.
Publicado: (2026)
Causal Dimensionality of Transformer Representations: Measurement, Scaling, and Layer Structure
por: Sarkar, Nilesh, et al.
Publicado: (2026)
por: Sarkar, Nilesh, et al.
Publicado: (2026)
Deciphering Digital Detectives: Understanding LLM Behaviors and Capabilities in Multi-Agent Mystery Games
por: Wu, Dekun, et al.
Publicado: (2023)
por: Wu, Dekun, et al.
Publicado: (2023)
Chatbots put to the test in math and logic problems: A preliminary comparison and assessment of ChatGPT-3.5, ChatGPT-4, and Google Bard
por: Plevris, Vagelis, et al.
Publicado: (2023)
por: Plevris, Vagelis, et al.
Publicado: (2023)
Prompt Tuned Embedding Classification for Multi-Label Industry Sector Allocation
por: Buchner, Valentin Leonhard, et al.
Publicado: (2023)
por: Buchner, Valentin Leonhard, et al.
Publicado: (2023)
Benchmarking quantized LLaMa-based models on the Brazilian Secondary School Exam
por: Santos, Matheus L. O., et al.
Publicado: (2023)
por: Santos, Matheus L. O., et al.
Publicado: (2023)
Next Token Prediction Is a Dead End for Creativity
por: Olatunji, Ibukun, et al.
Publicado: (2025)
por: Olatunji, Ibukun, et al.
Publicado: (2025)
BabyReasoningBench: Generating Developmentally-Inspired Reasoning Tasks for Evaluating Baby Language Models
por: Dhole, Kaustubh D.
Publicado: (2026)
por: Dhole, Kaustubh D.
Publicado: (2026)
Machines of Meaning
por: Nunes, Davide, et al.
Publicado: (2024)
por: Nunes, Davide, et al.
Publicado: (2024)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
por: Fadli, Samih
Publicado: (2025)
por: Fadli, Samih
Publicado: (2025)
ReFoRCE: A Text-to-SQL Agent with Self-Refinement, Consensus Enforcement, and Column Exploration
por: Deng, Minghang, et al.
Publicado: (2025)
por: Deng, Minghang, et al.
Publicado: (2025)
Correcting Gradient-Based Circuit Localization via Interaction-Aware Backpropagation
por: Edin, Joakim, et al.
Publicado: (2025)
por: Edin, Joakim, et al.
Publicado: (2025)
Prompt Readiness Levels (PRL): a maturity scale and scoring framework for production grade prompt assets
por: Guinard, Sebastien
Publicado: (2026)
por: Guinard, Sebastien
Publicado: (2026)
Syntactic Blind Spots: How Misalignment Leads to LLMs Mathematical Errors
por: Williamson, Dane, et al.
Publicado: (2025)
por: Williamson, Dane, et al.
Publicado: (2025)
LLMs and the Human Condition
por: Wallis, Peter
Publicado: (2024)
por: Wallis, Peter
Publicado: (2024)
KnowThyself: An Agentic Assistant for LLM Interpretability
por: Prasai, Suraj, et al.
Publicado: (2025)
por: Prasai, Suraj, et al.
Publicado: (2025)
TwinVoice: A Multi-dimensional Benchmark Towards Digital Twins via LLM Persona Simulation
por: Du, Bangde, et al.
Publicado: (2025)
por: Du, Bangde, et al.
Publicado: (2025)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
por: Oketunji, Abiodun Finbarrs
Publicado: (2023)
por: Oketunji, Abiodun Finbarrs
Publicado: (2023)
Ejemplares similares
-
Graph Language Models
por: Plenz, Moritz, et al.
Publicado: (2024) -
The Drill-Down and Fabricate Test (DDFT): A Protocol for Measuring Epistemic Robustness in Language Models
por: Baxi, Rahul
Publicado: (2025) -
ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models
por: Dumitru, Razvan-Gabriel, et al.
Publicado: (2025) -
Layer-Wise Quantization: A Pragmatic and Effective Method for Quantizing LLMs Beyond Integer Bit-Levels
por: Dumitru, Razvan-Gabriel, et al.
Publicado: (2024) -
Key-Value Means: Transformers with Expandable Block-Recurrent Compressed Memory
por: Goldstein, Daniel, et al.
Publicado: (2026)