POEMetric: The Last Stanza of Humanity
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Bingru, Wang, Han, Wilkinson, Hazel |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
TACOMORE: Leveraging the Potential of LLMs in Corpus-based Discourse Analysis with Prompt Engineering
di: Li, Bingru, et al.
Pubblicazione: (2024)
di: Li, Bingru, et al.
Pubblicazione: (2024)
LinguistAgent: A Reflective Multi-Model Platform for Automated Linguistic Annotation
di: Li, Bingru
Pubblicazione: (2026)
di: Li, Bingru
Pubblicazione: (2026)
HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam
di: Zhai, Weiqi, et al.
Pubblicazione: (2026)
di: Zhai, Weiqi, et al.
Pubblicazione: (2026)
Humanity's Last Exam
di: Phan, Long, et al.
Pubblicazione: (2025)
di: Phan, Long, et al.
Pubblicazione: (2025)
Humanity's Last Code Exam: Can Advanced LLMs Conquer Human's Hardest Code Competition?
di: Li, Xiangyang, et al.
Pubblicazione: (2025)
di: Li, Xiangyang, et al.
Pubblicazione: (2025)
Vision as LoRA
di: Wang, Han, et al.
Pubblicazione: (2025)
di: Wang, Han, et al.
Pubblicazione: (2025)
Single LLM Debate, MoLaCE: Mixture of Latent Concept Experts Against Confirmation Bias
di: Kim, Hazel, et al.
Pubblicazione: (2025)
di: Kim, Hazel, et al.
Pubblicazione: (2025)
How Ambiguous Are the Rationales for Natural Language Reasoning? A Simple Approach to Handling Rationale Uncertainty
di: Kim, Hazel H.
Pubblicazione: (2024)
di: Kim, Hazel H.
Pubblicazione: (2024)
The Moltbook Files: A Harmless Slopocalypse or Humanity's Last Experiment
di: Brach, William, et al.
Pubblicazione: (2026)
di: Brach, William, et al.
Pubblicazione: (2026)
Magnetic Preference Optimization: Achieving Last-iterate Convergence for Language Model Alignment
di: Wang, Mingzhi, et al.
Pubblicazione: (2024)
di: Wang, Mingzhi, et al.
Pubblicazione: (2024)
Query Disambiguation via Answer-Free Context: Doubling Performance on Humanity's Last Exam
di: Majurski, Michael, et al.
Pubblicazione: (2026)
di: Majurski, Michael, et al.
Pubblicazione: (2026)
Reinforcement Learning without Human Feedback for Last Mile Fine-Tuning of Large Language Models
di: Solway, Alec
Pubblicazione: (2024)
di: Solway, Alec
Pubblicazione: (2024)
HLL: Can Agents Cross Humanity's Last Line of Verification?
di: Song, Xinhao, et al.
Pubblicazione: (2026)
di: Song, Xinhao, et al.
Pubblicazione: (2026)
jina-reranker-v3: Last but Not Late Interaction for Listwise Document Reranking
di: Wang, Feng, et al.
Pubblicazione: (2025)
di: Wang, Feng, et al.
Pubblicazione: (2025)
LastingBench: Defend Benchmarks Against Knowledge Leakage
di: Fang, Yixiong, et al.
Pubblicazione: (2025)
di: Fang, Yixiong, et al.
Pubblicazione: (2025)
From Code to Correctness: Closing the Last Mile of Code Generation with Hierarchical Debugging
di: Shi, Yuling, et al.
Pubblicazione: (2024)
di: Shi, Yuling, et al.
Pubblicazione: (2024)
Last Layer Logits to Logic: Empowering LLMs with Logic-Consistent Structured Knowledge Reasoning
di: Li, Songze, et al.
Pubblicazione: (2025)
di: Li, Songze, et al.
Pubblicazione: (2025)
SciMaster: Towards General-Purpose Scientific AI Agents, Part I. X-Master as Foundation: Can We Lead on Humanity's Last Exam?
di: Chai, Jingyi, et al.
Pubblicazione: (2025)
di: Chai, Jingyi, et al.
Pubblicazione: (2025)
Beyond Coarse-Grained Matching in Video-Text Retrieval
di: Chen, Aozhu, et al.
Pubblicazione: (2024)
di: Chen, Aozhu, et al.
Pubblicazione: (2024)
Stanzas of Woe, Stanzas of Anger: Ann Yearsley and Poetry as Complaint and Instruction
di: Catherine Keohane
Pubblicazione: (2026)
di: Catherine Keohane
Pubblicazione: (2026)
LIFT: Last-Mile Fine-Tuning for Table Explicitation
di: Khaitan, Divij, et al.
Pubblicazione: (2026)
di: Khaitan, Divij, et al.
Pubblicazione: (2026)
Hallucinate at the Last in Long Response Generation: A Case Study on Long Document Summarization
di: Yang, Joonho, et al.
Pubblicazione: (2025)
di: Yang, Joonho, et al.
Pubblicazione: (2025)
AlignSum: Data Pyramid Hierarchical Fine-tuning for Aligning with Human Summarization Preference
di: Han, Yang, et al.
Pubblicazione: (2024)
di: Han, Yang, et al.
Pubblicazione: (2024)
The Position Curse: LLMs Struggle to Locate the Last Few Items in a List
di: Zhang, Zhanqi, et al.
Pubblicazione: (2026)
di: Zhang, Zhanqi, et al.
Pubblicazione: (2026)
The Last Fingerprint: How Markdown Training Shapes LLM Prose
di: Freeburg, E. M.
Pubblicazione: (2026)
di: Freeburg, E. M.
Pubblicazione: (2026)
LocoMotion: Learning Motion-Focused Video-Language Representations
di: Doughty, Hazel, et al.
Pubblicazione: (2024)
di: Doughty, Hazel, et al.
Pubblicazione: (2024)
Why teaching resists automation in an AI-inundated era: Human judgment, non-modular work, and the limits of delegation
di: Han, Songhee
Pubblicazione: (2026)
di: Han, Songhee
Pubblicazione: (2026)
BatchEval: Towards Human-like Text Evaluation
di: Yuan, Peiwen, et al.
Pubblicazione: (2023)
di: Yuan, Peiwen, et al.
Pubblicazione: (2023)
Improving Next Tokens via Second-to-Last Predictions with Generate and Refine
di: Schneider, Johannes
Pubblicazione: (2024)
di: Schneider, Johannes
Pubblicazione: (2024)
EvoGrad: A Dynamic Take on the Winograd Schema Challenge with Human Adversaries
di: Sun, Jing Han, et al.
Pubblicazione: (2024)
di: Sun, Jing Han, et al.
Pubblicazione: (2024)
Probing Large Language Models from A Human Behavioral Perspective
di: Wang, Xintong, et al.
Pubblicazione: (2023)
di: Wang, Xintong, et al.
Pubblicazione: (2023)
IQA-EVAL: Automatic Evaluation of Human-Model Interactive Question Answering
di: Li, Ruosen, et al.
Pubblicazione: (2024)
di: Li, Ruosen, et al.
Pubblicazione: (2024)
MANBench: Is Your Multimodal Model Smarter than Human?
di: Zhou, Han, et al.
Pubblicazione: (2025)
di: Zhou, Han, et al.
Pubblicazione: (2025)
Aligning Language Models with Human Preferences via a Bayesian Approach
di: Wang, Jiashuo, et al.
Pubblicazione: (2023)
di: Wang, Jiashuo, et al.
Pubblicazione: (2023)
On the Interplay between Human Label Variation and Model Fairness
di: Kurniawan, Kemal, et al.
Pubblicazione: (2025)
di: Kurniawan, Kemal, et al.
Pubblicazione: (2025)
Let the Expert Stick to His Last: Expert-Specialized Fine-Tuning for Sparse Architectural Large Language Models
di: Wang, Zihan, et al.
Pubblicazione: (2024)
di: Wang, Zihan, et al.
Pubblicazione: (2024)
LVLMs are Bad at Overhearing Human Referential Communication
di: Wang, Zhengxiang, et al.
Pubblicazione: (2025)
di: Wang, Zhengxiang, et al.
Pubblicazione: (2025)
Reward Modeling from Natural Language Human Feedback
di: Wang, Zongqi, et al.
Pubblicazione: (2026)
di: Wang, Zongqi, et al.
Pubblicazione: (2026)
Aligning MLLM Benchmark With Human Preferences via Structural Equation Modeling
di: Xiong, Shengwu., et al.
Pubblicazione: (2025)
di: Xiong, Shengwu., et al.
Pubblicazione: (2025)
Ara-HOPE: Human-Centric Post-Editing Evaluation for Dialectal Arabic to Modern Standard Arabic Translation
di: Alabdullah, Abdullah, et al.
Pubblicazione: (2025)
di: Alabdullah, Abdullah, et al.
Pubblicazione: (2025)
Documenti analoghi
-
TACOMORE: Leveraging the Potential of LLMs in Corpus-based Discourse Analysis with Prompt Engineering
di: Li, Bingru, et al.
Pubblicazione: (2024) -
LinguistAgent: A Reflective Multi-Model Platform for Automated Linguistic Annotation
di: Li, Bingru
Pubblicazione: (2026) -
HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam
di: Zhai, Weiqi, et al.
Pubblicazione: (2026) -
Humanity's Last Exam
di: Phan, Long, et al.
Pubblicazione: (2025) -
Humanity's Last Code Exam: Can Advanced LLMs Conquer Human's Hardest Code Competition?
di: Li, Xiangyang, et al.
Pubblicazione: (2025)