A surprisal oracle for when every layer counts
Fuente:
arXiv
Salvato in:
| Autori principali: | Hong, Xudong, Loáiciga, Sharid, Sayeed, Asad |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Coreference as an indicator of context scope in multimodal narrative
di: Ilinykh, Nikolai, et al.
Pubblicazione: (2025)
di: Ilinykh, Nikolai, et al.
Pubblicazione: (2025)
Humans vs Vision-Language Models: A Unified Measure of Narrative Coherence
di: Ilinykh, Nikolai, et al.
Pubblicazione: (2026)
di: Ilinykh, Nikolai, et al.
Pubblicazione: (2026)
Understanding and Analyzing Model Robustness and Knowledge-Transfer in Multilingual Neural Machine Translation using TX-Ray
di: Saxena, Vageesh, et al.
Pubblicazione: (2024)
di: Saxena, Vageesh, et al.
Pubblicazione: (2024)
Predicting Sentence Acceptability Judgments in Multimodal Contexts
di: Jang, Hyewon, et al.
Pubblicazione: (2026)
di: Jang, Hyewon, et al.
Pubblicazione: (2026)
Diacritic Restoration for Low-Resource Indigenous Languages: Case Study with Bribri and Cook Islands Māori
di: Coto-Solano, Rolando, et al.
Pubblicazione: (2025)
di: Coto-Solano, Rolando, et al.
Pubblicazione: (2025)
Subword models struggle with word learning, but surprisal hides it
di: Bunzeck, Bastian, et al.
Pubblicazione: (2025)
di: Bunzeck, Bastian, et al.
Pubblicazione: (2025)
Spoken Word2Vec: Learning Skipgram Embeddings from Speech
di: Sayeed, Mohammad Amaan, et al.
Pubblicazione: (2023)
di: Sayeed, Mohammad Amaan, et al.
Pubblicazione: (2023)
Deconstructing sentence disambiguation by joint latent modeling of reading paradigms: LLM surprisal is not enough
di: Paape, Dario, et al.
Pubblicazione: (2026)
di: Paape, Dario, et al.
Pubblicazione: (2026)
Position: Avoid Overstretching LLMs for every Enterprise Task
di: Singh, Kuldeep, et al.
Pubblicazione: (2026)
di: Singh, Kuldeep, et al.
Pubblicazione: (2026)
Decomposition of surprisal: Unified computational model of ERP components in language processing
di: Li, Jiaxuan, et al.
Pubblicazione: (2024)
di: Li, Jiaxuan, et al.
Pubblicazione: (2024)
Why are language models less surprised than humans? Testing the Parse Multiplicity Mismatch Hypothesis
di: Timkey, William, et al.
Pubblicazione: (2026)
di: Timkey, William, et al.
Pubblicazione: (2026)
Prompt-Based LLMs for Position Bias-Aware Reranking in Personalized Recommendations
di: Islam, Md Aminul, et al.
Pubblicazione: (2025)
di: Islam, Md Aminul, et al.
Pubblicazione: (2025)
Learning From Failure: Integrating Negative Examples when Fine-tuning Large Language Models as Agents
di: Wang, Renxi, et al.
Pubblicazione: (2024)
di: Wang, Renxi, et al.
Pubblicazione: (2024)
BLEUBERI: BLEU is a surprisingly effective reward for instruction following
di: Chang, Yapei, et al.
Pubblicazione: (2025)
di: Chang, Yapei, et al.
Pubblicazione: (2025)
MLPs Compass: What is learned when MLPs are combined with PLMs?
di: Zhou, Li, et al.
Pubblicazione: (2024)
di: Zhou, Li, et al.
Pubblicazione: (2024)
Large language models as oracles for instantiating ontologies with domain-specific knowledge
di: Ciatto, Giovanni, et al.
Pubblicazione: (2024)
di: Ciatto, Giovanni, et al.
Pubblicazione: (2024)
Contextual Breach: Assessing the Robustness of Transformer-based QA Models
di: Saadat, Asir, et al.
Pubblicazione: (2024)
di: Saadat, Asir, et al.
Pubblicazione: (2024)
Temperature-scaling surprisal estimates improve fit to human reading times -- but does it do so for the "right reasons"?
di: Liu, Tong, et al.
Pubblicazione: (2023)
di: Liu, Tong, et al.
Pubblicazione: (2023)
Too Late to Train, Too Early To Use? A Study on Necessity and Viability of Low-Resource Bengali LLMs
di: Mahfuz, Tamzeed, et al.
Pubblicazione: (2024)
di: Mahfuz, Tamzeed, et al.
Pubblicazione: (2024)
From RAG to Agentic: Validating Islamic-Medicine Responses with LLM Agents
di: Sayeed, Mohammad Amaan, et al.
Pubblicazione: (2025)
di: Sayeed, Mohammad Amaan, et al.
Pubblicazione: (2025)
Theoretical Proof that Auto-regressive Language Models Collapse when Real-world Data is a Finite Set
di: Wang, Lecheng, et al.
Pubblicazione: (2024)
di: Wang, Lecheng, et al.
Pubblicazione: (2024)
When does word order matter and when doesn't it?
di: Chen, Xuanda, et al.
Pubblicazione: (2024)
di: Chen, Xuanda, et al.
Pubblicazione: (2024)
RedDebate: Safer Responses Through Multi-Agent Red Teaming Debates
di: Asad, Ali, et al.
Pubblicazione: (2025)
di: Asad, Ali, et al.
Pubblicazione: (2025)
Accurate and Data-Efficient Toxicity Prediction when Annotators Disagree
di: Jaggi, Harbani, et al.
Pubblicazione: (2024)
di: Jaggi, Harbani, et al.
Pubblicazione: (2024)
Is Information Density Uniform when Utterances are Grounded on Perception and Discourse?
di: Gay, Matteo, et al.
Pubblicazione: (2026)
di: Gay, Matteo, et al.
Pubblicazione: (2026)
As easy as PIE: understanding when pruning causes language models to disagree
di: Tropeano, Pietro, et al.
Pubblicazione: (2025)
di: Tropeano, Pietro, et al.
Pubblicazione: (2025)
LLMs Struggle to Reject False Presuppositions when Misinformation Stakes are High
di: Sieker, Judith, et al.
Pubblicazione: (2025)
di: Sieker, Judith, et al.
Pubblicazione: (2025)
IllusionVQA: A Challenging Optical Illusion Dataset for Vision Language Models
di: Shahgir, Haz Sameen, et al.
Pubblicazione: (2024)
di: Shahgir, Haz Sameen, et al.
Pubblicazione: (2024)
Revisiting Greedy Decoding for Visual Question Answering: A Calibration Perspective
di: Chen, Boqi, et al.
Pubblicazione: (2026)
di: Chen, Boqi, et al.
Pubblicazione: (2026)
OpenDeception: Learning Deception and Trust in Human-AI Interaction via Multi-Agent Simulation
di: Wu, Yichen, et al.
Pubblicazione: (2025)
di: Wu, Yichen, et al.
Pubblicazione: (2025)
ReasoningShield: Safety Detection over Reasoning Traces of Large Reasoning Models
di: Li, Changyi, et al.
Pubblicazione: (2025)
di: Li, Changyi, et al.
Pubblicazione: (2025)
CogRAG+: Cognitive-Level Guided Diagnosis and Remediation of Memory and Reasoning Deficiencies in Professional Exam QA
di: Wang, Xudong, et al.
Pubblicazione: (2026)
di: Wang, Xudong, et al.
Pubblicazione: (2026)
Strengthening False Information Propagation Detection: Leveraging SVM and Sophisticated Text Vectorization Techniques in comparison to BERT
di: Karim, Ahmed Akib Jawad, et al.
Pubblicazione: (2024)
di: Karim, Ahmed Akib Jawad, et al.
Pubblicazione: (2024)
Who Benefits From Sinus Surgery? Comparing Generative AI and Supervised Machine Learning for Predicting Surgical Outcomes in Chronic Rhinosinusitis
di: Chowdhury, Sayeed Shafayet, et al.
Pubblicazione: (2026)
di: Chowdhury, Sayeed Shafayet, et al.
Pubblicazione: (2026)
Modeling the Sacred: Considerations when Using Religious Texts in Natural Language Processing
di: Hutchinson, Ben
Pubblicazione: (2024)
di: Hutchinson, Ben
Pubblicazione: (2024)
Misgendering and Assuming Gender in Machine Translation when Working with Low-Resource Languages
di: Ghosh, Sourojit, et al.
Pubblicazione: (2024)
di: Ghosh, Sourojit, et al.
Pubblicazione: (2024)
Building Better: Avoiding Pitfalls in Developing Language Resources when Data is Scarce
di: Ousidhoum, Nedjma, et al.
Pubblicazione: (2024)
di: Ousidhoum, Nedjma, et al.
Pubblicazione: (2024)
WaterJudge: Quality-Detection Trade-off when Watermarking Large Language Models
di: Molenda, Piotr, et al.
Pubblicazione: (2024)
di: Molenda, Piotr, et al.
Pubblicazione: (2024)
Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?
di: Singh, Aaditya K., et al.
Pubblicazione: (2024)
di: Singh, Aaditya K., et al.
Pubblicazione: (2024)
Granularity is crucial when applying differential privacy to text: An investigation for neural machine translation
di: Vu, Doan Nam Long, et al.
Pubblicazione: (2024)
di: Vu, Doan Nam Long, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Coreference as an indicator of context scope in multimodal narrative
di: Ilinykh, Nikolai, et al.
Pubblicazione: (2025) -
Humans vs Vision-Language Models: A Unified Measure of Narrative Coherence
di: Ilinykh, Nikolai, et al.
Pubblicazione: (2026) -
Understanding and Analyzing Model Robustness and Knowledge-Transfer in Multilingual Neural Machine Translation using TX-Ray
di: Saxena, Vageesh, et al.
Pubblicazione: (2024) -
Predicting Sentence Acceptability Judgments in Multimodal Contexts
di: Jang, Hyewon, et al.
Pubblicazione: (2026) -
Diacritic Restoration for Low-Resource Indigenous Languages: Case Study with Bribri and Cook Islands Māori
di: Coto-Solano, Rolando, et al.
Pubblicazione: (2025)