POEMetric: The Last Stanza of Humanity
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Bingru, Wang, Han, Wilkinson, Hazel |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
TACOMORE: Leveraging the Potential of LLMs in Corpus-based Discourse Analysis with Prompt Engineering
por: Li, Bingru, et al.
Publicado: (2024)
por: Li, Bingru, et al.
Publicado: (2024)
LinguistAgent: A Reflective Multi-Model Platform for Automated Linguistic Annotation
por: Li, Bingru
Publicado: (2026)
por: Li, Bingru
Publicado: (2026)
HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam
por: Zhai, Weiqi, et al.
Publicado: (2026)
por: Zhai, Weiqi, et al.
Publicado: (2026)
Humanity's Last Exam
por: Phan, Long, et al.
Publicado: (2025)
por: Phan, Long, et al.
Publicado: (2025)
Humanity's Last Code Exam: Can Advanced LLMs Conquer Human's Hardest Code Competition?
por: Li, Xiangyang, et al.
Publicado: (2025)
por: Li, Xiangyang, et al.
Publicado: (2025)
Vision as LoRA
por: Wang, Han, et al.
Publicado: (2025)
por: Wang, Han, et al.
Publicado: (2025)
Single LLM Debate, MoLaCE: Mixture of Latent Concept Experts Against Confirmation Bias
por: Kim, Hazel, et al.
Publicado: (2025)
por: Kim, Hazel, et al.
Publicado: (2025)
How Ambiguous Are the Rationales for Natural Language Reasoning? A Simple Approach to Handling Rationale Uncertainty
por: Kim, Hazel H.
Publicado: (2024)
por: Kim, Hazel H.
Publicado: (2024)
The Moltbook Files: A Harmless Slopocalypse or Humanity's Last Experiment
por: Brach, William, et al.
Publicado: (2026)
por: Brach, William, et al.
Publicado: (2026)
Magnetic Preference Optimization: Achieving Last-iterate Convergence for Language Model Alignment
por: Wang, Mingzhi, et al.
Publicado: (2024)
por: Wang, Mingzhi, et al.
Publicado: (2024)
Query Disambiguation via Answer-Free Context: Doubling Performance on Humanity's Last Exam
por: Majurski, Michael, et al.
Publicado: (2026)
por: Majurski, Michael, et al.
Publicado: (2026)
Reinforcement Learning without Human Feedback for Last Mile Fine-Tuning of Large Language Models
por: Solway, Alec
Publicado: (2024)
por: Solway, Alec
Publicado: (2024)
HLL: Can Agents Cross Humanity's Last Line of Verification?
por: Song, Xinhao, et al.
Publicado: (2026)
por: Song, Xinhao, et al.
Publicado: (2026)
jina-reranker-v3: Last but Not Late Interaction for Listwise Document Reranking
por: Wang, Feng, et al.
Publicado: (2025)
por: Wang, Feng, et al.
Publicado: (2025)
LastingBench: Defend Benchmarks Against Knowledge Leakage
por: Fang, Yixiong, et al.
Publicado: (2025)
por: Fang, Yixiong, et al.
Publicado: (2025)
From Code to Correctness: Closing the Last Mile of Code Generation with Hierarchical Debugging
por: Shi, Yuling, et al.
Publicado: (2024)
por: Shi, Yuling, et al.
Publicado: (2024)
Last Layer Logits to Logic: Empowering LLMs with Logic-Consistent Structured Knowledge Reasoning
por: Li, Songze, et al.
Publicado: (2025)
por: Li, Songze, et al.
Publicado: (2025)
SciMaster: Towards General-Purpose Scientific AI Agents, Part I. X-Master as Foundation: Can We Lead on Humanity's Last Exam?
por: Chai, Jingyi, et al.
Publicado: (2025)
por: Chai, Jingyi, et al.
Publicado: (2025)
Beyond Coarse-Grained Matching in Video-Text Retrieval
por: Chen, Aozhu, et al.
Publicado: (2024)
por: Chen, Aozhu, et al.
Publicado: (2024)
Stanzas of Woe, Stanzas of Anger: Ann Yearsley and Poetry as Complaint and Instruction
por: Catherine Keohane
Publicado: (2026)
por: Catherine Keohane
Publicado: (2026)
LIFT: Last-Mile Fine-Tuning for Table Explicitation
por: Khaitan, Divij, et al.
Publicado: (2026)
por: Khaitan, Divij, et al.
Publicado: (2026)
Hallucinate at the Last in Long Response Generation: A Case Study on Long Document Summarization
por: Yang, Joonho, et al.
Publicado: (2025)
por: Yang, Joonho, et al.
Publicado: (2025)
AlignSum: Data Pyramid Hierarchical Fine-tuning for Aligning with Human Summarization Preference
por: Han, Yang, et al.
Publicado: (2024)
por: Han, Yang, et al.
Publicado: (2024)
The Position Curse: LLMs Struggle to Locate the Last Few Items in a List
por: Zhang, Zhanqi, et al.
Publicado: (2026)
por: Zhang, Zhanqi, et al.
Publicado: (2026)
The Last Fingerprint: How Markdown Training Shapes LLM Prose
por: Freeburg, E. M.
Publicado: (2026)
por: Freeburg, E. M.
Publicado: (2026)
LocoMotion: Learning Motion-Focused Video-Language Representations
por: Doughty, Hazel, et al.
Publicado: (2024)
por: Doughty, Hazel, et al.
Publicado: (2024)
Why teaching resists automation in an AI-inundated era: Human judgment, non-modular work, and the limits of delegation
por: Han, Songhee
Publicado: (2026)
por: Han, Songhee
Publicado: (2026)
BatchEval: Towards Human-like Text Evaluation
por: Yuan, Peiwen, et al.
Publicado: (2023)
por: Yuan, Peiwen, et al.
Publicado: (2023)
Improving Next Tokens via Second-to-Last Predictions with Generate and Refine
por: Schneider, Johannes
Publicado: (2024)
por: Schneider, Johannes
Publicado: (2024)
EvoGrad: A Dynamic Take on the Winograd Schema Challenge with Human Adversaries
por: Sun, Jing Han, et al.
Publicado: (2024)
por: Sun, Jing Han, et al.
Publicado: (2024)
Probing Large Language Models from A Human Behavioral Perspective
por: Wang, Xintong, et al.
Publicado: (2023)
por: Wang, Xintong, et al.
Publicado: (2023)
IQA-EVAL: Automatic Evaluation of Human-Model Interactive Question Answering
por: Li, Ruosen, et al.
Publicado: (2024)
por: Li, Ruosen, et al.
Publicado: (2024)
MANBench: Is Your Multimodal Model Smarter than Human?
por: Zhou, Han, et al.
Publicado: (2025)
por: Zhou, Han, et al.
Publicado: (2025)
Aligning Language Models with Human Preferences via a Bayesian Approach
por: Wang, Jiashuo, et al.
Publicado: (2023)
por: Wang, Jiashuo, et al.
Publicado: (2023)
On the Interplay between Human Label Variation and Model Fairness
por: Kurniawan, Kemal, et al.
Publicado: (2025)
por: Kurniawan, Kemal, et al.
Publicado: (2025)
Let the Expert Stick to His Last: Expert-Specialized Fine-Tuning for Sparse Architectural Large Language Models
por: Wang, Zihan, et al.
Publicado: (2024)
por: Wang, Zihan, et al.
Publicado: (2024)
LVLMs are Bad at Overhearing Human Referential Communication
por: Wang, Zhengxiang, et al.
Publicado: (2025)
por: Wang, Zhengxiang, et al.
Publicado: (2025)
Reward Modeling from Natural Language Human Feedback
por: Wang, Zongqi, et al.
Publicado: (2026)
por: Wang, Zongqi, et al.
Publicado: (2026)
Aligning MLLM Benchmark With Human Preferences via Structural Equation Modeling
por: Xiong, Shengwu., et al.
Publicado: (2025)
por: Xiong, Shengwu., et al.
Publicado: (2025)
Ara-HOPE: Human-Centric Post-Editing Evaluation for Dialectal Arabic to Modern Standard Arabic Translation
por: Alabdullah, Abdullah, et al.
Publicado: (2025)
por: Alabdullah, Abdullah, et al.
Publicado: (2025)
Ejemplares similares
-
TACOMORE: Leveraging the Potential of LLMs in Corpus-based Discourse Analysis with Prompt Engineering
por: Li, Bingru, et al.
Publicado: (2024) -
LinguistAgent: A Reflective Multi-Model Platform for Automated Linguistic Annotation
por: Li, Bingru
Publicado: (2026) -
HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam
por: Zhai, Weiqi, et al.
Publicado: (2026) -
Humanity's Last Exam
por: Phan, Long, et al.
Publicado: (2025) -
Humanity's Last Code Exam: Can Advanced LLMs Conquer Human's Hardest Code Competition?
por: Li, Xiangyang, et al.
Publicado: (2025)