BooookScore: A systematic exploration of book-length summarization in the era of LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Chang, Yapei, Lo, Kyle, Goyal, Tanya, Iyyer, Mohit |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
FABLES: Evaluating faithfulness and content selection in book-length summarization
di: Kim, Yekyung, et al.
Pubblicazione: (2024)
di: Kim, Yekyung, et al.
Pubblicazione: (2024)
PostMark: A Robust Blackbox Watermark for Large Language Models
di: Chang, Yapei, et al.
Pubblicazione: (2024)
di: Chang, Yapei, et al.
Pubblicazione: (2024)
One Thousand and One Pairs: A "novel" challenge for long-context language models
di: Karpinska, Marzena, et al.
Pubblicazione: (2024)
di: Karpinska, Marzena, et al.
Pubblicazione: (2024)
BEARCUBS: A benchmark for computer-using web agents
di: Song, Yixiao, et al.
Pubblicazione: (2025)
di: Song, Yixiao, et al.
Pubblicazione: (2025)
BLEUBERI: BLEU is a surprisingly effective reward for instruction following
di: Chang, Yapei, et al.
Pubblicazione: (2025)
di: Chang, Yapei, et al.
Pubblicazione: (2025)
Argument Collapse: LLMs Flatten Long-Form Public Debate
di: Kim, Yekyung, et al.
Pubblicazione: (2026)
di: Kim, Yekyung, et al.
Pubblicazione: (2026)
How2Everything: Mining the Web for How-To Procedures to Evaluate and Improve LLMs
di: Chang, Yapei, et al.
Pubblicazione: (2026)
di: Chang, Yapei, et al.
Pubblicazione: (2026)
Memorization vs. Reasoning: Updating LLMs with New Knowledge
di: Li, Aochong Oliver, et al.
Pubblicazione: (2025)
di: Li, Aochong Oliver, et al.
Pubblicazione: (2025)
DCRM: A Heuristic to Measure Response Pair Quality in Preference Optimization
di: Huang, Chengyu, et al.
Pubblicazione: (2025)
di: Huang, Chengyu, et al.
Pubblicazione: (2025)
Recovering Diversity Without Losing Alignment: A DPO Recipe for Post-Trained LLMs
di: Samuel, Vinay, et al.
Pubblicazione: (2026)
di: Samuel, Vinay, et al.
Pubblicazione: (2026)
CaLMQA: Exploring culturally specific long-form question answering across 23 languages
di: Arora, Shane, et al.
Pubblicazione: (2024)
di: Arora, Shane, et al.
Pubblicazione: (2024)
A dataset and benchmark for hospital course summarization with adapted large language models
di: Aali, Asad, et al.
Pubblicazione: (2024)
di: Aali, Asad, et al.
Pubblicazione: (2024)
CLIPPER: Compression enables long-context synthetic data generation
di: Pham, Chau Minh, et al.
Pubblicazione: (2025)
di: Pham, Chau Minh, et al.
Pubblicazione: (2025)
ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs
di: Chen, Justin Chih-Yao, et al.
Pubblicazione: (2023)
di: Chen, Justin Chih-Yao, et al.
Pubblicazione: (2023)
People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text
di: Russell, Jenna, et al.
Pubblicazione: (2025)
di: Russell, Jenna, et al.
Pubblicazione: (2025)
EnvGen: Generating and Adapting Environments via LLMs for Training Embodied Agents
di: Zala, Abhay, et al.
Pubblicazione: (2024)
di: Zala, Abhay, et al.
Pubblicazione: (2024)
RSQ: Learning from Important Tokens Leads to Better Quantized LLMs
di: Sung, Yi-Lin, et al.
Pubblicazione: (2025)
di: Sung, Yi-Lin, et al.
Pubblicazione: (2025)
A comprehensive study of on-device NLP applications -- VQA, automated Form filling, Smart Replies for Linguistic Codeswitching
di: Goyal, Naman
Pubblicazione: (2024)
di: Goyal, Naman
Pubblicazione: (2024)
Extracting books from production language models
di: Ahmed, Ahmed, et al.
Pubblicazione: (2026)
di: Ahmed, Ahmed, et al.
Pubblicazione: (2026)
Learning Beyond Pattern Matching? Assaying Mathematical Understanding in LLMs
di: Guo, Siyuan, et al.
Pubblicazione: (2024)
di: Guo, Siyuan, et al.
Pubblicazione: (2024)
Keeping LLMs Aligned After Fine-tuning: The Crucial Role of Prompt Templates
di: Lyu, Kaifeng, et al.
Pubblicazione: (2024)
di: Lyu, Kaifeng, et al.
Pubblicazione: (2024)
Retrieval Enhanced Feedback via In-context Neural Error-book
di: Hyun, Jongyeop, et al.
Pubblicazione: (2025)
di: Hyun, Jongyeop, et al.
Pubblicazione: (2025)
Text2Freq: Learning Series Patterns from Text via Frequency Domain
di: Lo, Ming-Chih, et al.
Pubblicazione: (2024)
di: Lo, Ming-Chih, et al.
Pubblicazione: (2024)
Addressing LLM Diversity by Infusing Random Concepts
di: Agrawal, Pulin, et al.
Pubblicazione: (2026)
di: Agrawal, Pulin, et al.
Pubblicazione: (2026)
LLMs can hide text in other text of the same length
di: Norelli, Antonio, et al.
Pubblicazione: (2025)
di: Norelli, Antonio, et al.
Pubblicazione: (2025)
GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations
di: Duan, Jinhao, et al.
Pubblicazione: (2024)
di: Duan, Jinhao, et al.
Pubblicazione: (2024)
Reverse Thinking Makes LLMs Stronger Reasoners
di: Chen, Justin Chih-Yao, et al.
Pubblicazione: (2024)
di: Chen, Justin Chih-Yao, et al.
Pubblicazione: (2024)
ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
di: Lin, Bill Yuchen, et al.
Pubblicazione: (2025)
di: Lin, Bill Yuchen, et al.
Pubblicazione: (2025)
Uncertainty Quantification for Language Models: A Suite of Black-Box, White-Box, LLM Judge, and Ensemble Scorers
di: Bouchard, Dylan, et al.
Pubblicazione: (2025)
di: Bouchard, Dylan, et al.
Pubblicazione: (2025)
Generative Image as Action Models
di: Shridhar, Mohit, et al.
Pubblicazione: (2024)
di: Shridhar, Mohit, et al.
Pubblicazione: (2024)
Steering Safely or Off a Cliff? Rethinking Specificity and Robustness in Inference-Time Interventions
di: Goyal, Navita, et al.
Pubblicazione: (2026)
di: Goyal, Navita, et al.
Pubblicazione: (2026)
Can large language models explore in-context?
di: Krishnamurthy, Akshay, et al.
Pubblicazione: (2024)
di: Krishnamurthy, Akshay, et al.
Pubblicazione: (2024)
Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language Models
di: Wang, Fei, et al.
Pubblicazione: (2024)
di: Wang, Fei, et al.
Pubblicazione: (2024)
Inducing Systematicity in Transformers by Attending to Structurally Quantized Embeddings
di: Jiang, Yichen, et al.
Pubblicazione: (2024)
di: Jiang, Yichen, et al.
Pubblicazione: (2024)
Revisiting the Superficial Alignment Hypothesis
di: Raghavendra, Mohit, et al.
Pubblicazione: (2024)
di: Raghavendra, Mohit, et al.
Pubblicazione: (2024)
MAP's not dead yet: Uncovering true language model modes by conditioning away degeneracy
di: Yoshida, Davis, et al.
Pubblicazione: (2023)
di: Yoshida, Davis, et al.
Pubblicazione: (2023)
Debate Helps Weak Judges Reward Stronger Models
di: Elasky, Ethan, et al.
Pubblicazione: (2026)
di: Elasky, Ethan, et al.
Pubblicazione: (2026)
Pair2Score: Pairwise-to-Absolute Transfer for LLM-Based Essay Scoring
di: Hallaç, İbrahim Rıza, et al.
Pubblicazione: (2026)
di: Hallaç, İbrahim Rıza, et al.
Pubblicazione: (2026)
Olmix: A Framework for Data Mixing Throughout LM Development
di: Chen, Mayee F., et al.
Pubblicazione: (2026)
di: Chen, Mayee F., et al.
Pubblicazione: (2026)
UPCORE: Utility-Preserving Coreset Selection for Balanced Unlearning
di: Patil, Vaidehi, et al.
Pubblicazione: (2025)
di: Patil, Vaidehi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
FABLES: Evaluating faithfulness and content selection in book-length summarization
di: Kim, Yekyung, et al.
Pubblicazione: (2024) -
PostMark: A Robust Blackbox Watermark for Large Language Models
di: Chang, Yapei, et al.
Pubblicazione: (2024) -
One Thousand and One Pairs: A "novel" challenge for long-context language models
di: Karpinska, Marzena, et al.
Pubblicazione: (2024) -
BEARCUBS: A benchmark for computer-using web agents
di: Song, Yixiao, et al.
Pubblicazione: (2025) -
BLEUBERI: BLEU is a surprisingly effective reward for instruction following
di: Chang, Yapei, et al.
Pubblicazione: (2025)