Instructional Text Across Disciplines: A Survey of Representations, Downstream Tasks, and Open Challenges Toward Capable AI Agents
Fuente:
arXiv
Salvato in:
| Autori principali: | Safa, Abdulfattah, Kapanadze, Tamta, Uzunoğlu, Arda, Şahin, Gözde Gül |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Zero-Shot Open-Vocabulary Pipeline for Dialogue Understanding
di: Safa, Abdulfattah, et al.
Pubblicazione: (2024)
di: Safa, Abdulfattah, et al.
Pubblicazione: (2024)
Benchmarking Procedural Language Understanding for Low-Resource Languages: A Case Study on Turkish
di: Uzunoglu, Arda, et al.
Pubblicazione: (2023)
di: Uzunoglu, Arda, et al.
Pubblicazione: (2023)
PARADISE: Evaluating Implicit Planning Skills of Language Models with Procedural Warnings and Tips Dataset
di: Uzunoglu, Arda, et al.
Pubblicazione: (2024)
di: Uzunoglu, Arda, et al.
Pubblicazione: (2024)
Are Non-English Papers Reviewed Fairly? Language-of-Study Bias in NLP Peer Reviews
di: Barkhordar, Ehsan, et al.
Pubblicazione: (2026)
di: Barkhordar, Ehsan, et al.
Pubblicazione: (2026)
Linguistically-Informed Multilingual Instruction Tuning: Is There an Optimal Set of Languages to Tune?
di: Soykan, Gürkan, et al.
Pubblicazione: (2024)
di: Soykan, Gürkan, et al.
Pubblicazione: (2024)
GECTurk WEB: An Explainable Online Platform for Turkish Grammatical Error Detection and Correction
di: Gebeşçe, Ali, et al.
Pubblicazione: (2024)
di: Gebeşçe, Ali, et al.
Pubblicazione: (2024)
Trust Functions: Near-Lossless Weak-to-Strong Generalization by Learning When to Trust the Weak Teacher
di: Uzunoglu, Arda, et al.
Pubblicazione: (2026)
di: Uzunoglu, Arda, et al.
Pubblicazione: (2026)
The Flaw of Averages: Quantifying Uniformity of Performance on Benchmarks
di: Uzunoglu, Arda, et al.
Pubblicazione: (2025)
di: Uzunoglu, Arda, et al.
Pubblicazione: (2025)
WorldAPIs: The World Is Worth How Many APIs? A Thought Experiment
di: Ou, Jiefu, et al.
Pubblicazione: (2024)
di: Ou, Jiefu, et al.
Pubblicazione: (2024)
Capability Salience Vector: Fine-grained Alignment of Loss and Capabilities for Downstream Task Scaling Law
di: Ge, Qiming, et al.
Pubblicazione: (2025)
di: Ge, Qiming, et al.
Pubblicazione: (2025)
Cetvel: A Unified Benchmark for Evaluating Language Understanding, Generation and Cultural Capacity of LLMs for Turkish
di: Er, Yakup Abrek, et al.
Pubblicazione: (2025)
di: Er, Yakup Abrek, et al.
Pubblicazione: (2025)
Applying Intrinsic Debiasing on Downstream Tasks: Challenges and Considerations for Machine Translation
di: Iluz, Bar, et al.
Pubblicazione: (2024)
di: Iluz, Bar, et al.
Pubblicazione: (2024)
Instruction Embedding: Latent Representations of Instructions Towards Task Identification
di: Li, Yiwei, et al.
Pubblicazione: (2024)
di: Li, Yiwei, et al.
Pubblicazione: (2024)
A Survey of Large Language Models in Discipline-specific Research: Challenges, Methods and Opportunities
di: Xiang, Lu, et al.
Pubblicazione: (2025)
di: Xiang, Lu, et al.
Pubblicazione: (2025)
Task-Informed Anti-Curriculum by Masking Improves Downstream Performance on Text
di: Jarca, Andrei, et al.
Pubblicazione: (2025)
di: Jarca, Andrei, et al.
Pubblicazione: (2025)
SuperValid: Capability-Aligned OOD Validation for Generalizable Downstream Scaling
di: Sun, Quanen, et al.
Pubblicazione: (2026)
di: Sun, Quanen, et al.
Pubblicazione: (2026)
SciReasoner: Laying the Scientific Reasoning Ground Across Disciplines
di: Wang, Yizhou, et al.
Pubblicazione: (2025)
di: Wang, Yizhou, et al.
Pubblicazione: (2025)
Instability in Downstream Task Performance During LLM Pretraining
di: Nishida, Yuto, et al.
Pubblicazione: (2025)
di: Nishida, Yuto, et al.
Pubblicazione: (2025)
Why Has Predicting Downstream Capabilities of Frontier AI Models with Scale Remained Elusive?
di: Schaeffer, Rylan, et al.
Pubblicazione: (2024)
di: Schaeffer, Rylan, et al.
Pubblicazione: (2024)
LLM-Based Data Science Agents: A Survey of Capabilities, Challenges, and Future Directions
di: Rahman, Mizanur, et al.
Pubblicazione: (2025)
di: Rahman, Mizanur, et al.
Pubblicazione: (2025)
Survey-to-Behavior: Downstream Alignment of Human Values in LLMs via Survey Questions
di: Nie, Shangrui, et al.
Pubblicazione: (2025)
di: Nie, Shangrui, et al.
Pubblicazione: (2025)
AgentIF-OneDay: A Task-level Instruction-Following Benchmark for General AI Agents in Daily Scenarios
di: Chen, Kaiyuan, et al.
Pubblicazione: (2026)
di: Chen, Kaiyuan, et al.
Pubblicazione: (2026)
An Evaluation of Sindhi Word Embedding in Semantic Analogies and Downstream Tasks
di: Ali, Wazir, et al.
Pubblicazione: (2024)
di: Ali, Wazir, et al.
Pubblicazione: (2024)
Edit Distances and Their Applications to Downstream Tasks in Research and Commercial Contexts
di: Carmo, Félix do, et al.
Pubblicazione: (2024)
di: Carmo, Félix do, et al.
Pubblicazione: (2024)
Measuring the Effect of Transcription Noise on Downstream Language Understanding Tasks
di: Shapira, Ori, et al.
Pubblicazione: (2025)
di: Shapira, Ori, et al.
Pubblicazione: (2025)
LLM-Detector: Improving AI-Generated Chinese Text Detection with Open-Source LLM Instruction Tuning
di: Wang, Rongsheng, et al.
Pubblicazione: (2024)
di: Wang, Rongsheng, et al.
Pubblicazione: (2024)
SurveyLens: A Research Discipline-Aware Benchmark for Automatic Survey Generation
di: Guo, Beichen, et al.
Pubblicazione: (2026)
di: Guo, Beichen, et al.
Pubblicazione: (2026)
Instruction Tuning with and without Context: Behavioral Shifts and Downstream Impact
di: Lee, Hyunji, et al.
Pubblicazione: (2025)
di: Lee, Hyunji, et al.
Pubblicazione: (2025)
CDT: A Comprehensive Capability Framework for Large Language Models Across Cognition, Domain, and Task
di: Mo, Haosi, et al.
Pubblicazione: (2025)
di: Mo, Haosi, et al.
Pubblicazione: (2025)
Towards Outcome-Oriented, Task-Agnostic Evaluation of AI Agents
di: AlShikh, Waseem, et al.
Pubblicazione: (2025)
di: AlShikh, Waseem, et al.
Pubblicazione: (2025)
LLMs4All: A Review of Large Language Models Across Academic Disciplines
di: Ye, Yanfang, et al.
Pubblicazione: (2025)
di: Ye, Yanfang, et al.
Pubblicazione: (2025)
It's Not the Capability: Harness Sensitivity Is Non-Monotone Across LLM Agent Tiers
di: Cho, Yong-eun
Pubblicazione: (2026)
di: Cho, Yong-eun
Pubblicazione: (2026)
Scaling Laws Are Unreliable for Downstream Tasks: A Reality Check
di: Lourie, Nicholas, et al.
Pubblicazione: (2025)
di: Lourie, Nicholas, et al.
Pubblicazione: (2025)
Scaling Laws for Downstream Task Performance of Large Language Models
di: Isik, Berivan, et al.
Pubblicazione: (2024)
di: Isik, Berivan, et al.
Pubblicazione: (2024)
Advancing Social Intelligence in AI Agents: Technical Challenges and Open Questions
di: Mathur, Leena, et al.
Pubblicazione: (2024)
di: Mathur, Leena, et al.
Pubblicazione: (2024)
Large Language Model Instruction Following: A Survey of Progresses and Challenges
di: Lou, Renze, et al.
Pubblicazione: (2023)
di: Lou, Renze, et al.
Pubblicazione: (2023)
CHANCERY: Evaluating Corporate Governance Reasoning Capabilities in Language Models
di: Irwin, Lucas, et al.
Pubblicazione: (2025)
di: Irwin, Lucas, et al.
Pubblicazione: (2025)
Large Scale Generative AI Text Applied to Sports and Music
di: Baughman, Aaron, et al.
Pubblicazione: (2024)
di: Baughman, Aaron, et al.
Pubblicazione: (2024)
Adapting Decoder-Based Language Models for Diverse Encoder Downstream Tasks
di: Suganthan, Paul, et al.
Pubblicazione: (2025)
di: Suganthan, Paul, et al.
Pubblicazione: (2025)
Optimising Language Models for Downstream Tasks: A Post-Training Perspective
di: Shi, Zhengyan
Pubblicazione: (2025)
di: Shi, Zhengyan
Pubblicazione: (2025)
Documenti analoghi
-
A Zero-Shot Open-Vocabulary Pipeline for Dialogue Understanding
di: Safa, Abdulfattah, et al.
Pubblicazione: (2024) -
Benchmarking Procedural Language Understanding for Low-Resource Languages: A Case Study on Turkish
di: Uzunoglu, Arda, et al.
Pubblicazione: (2023) -
PARADISE: Evaluating Implicit Planning Skills of Language Models with Procedural Warnings and Tips Dataset
di: Uzunoglu, Arda, et al.
Pubblicazione: (2024) -
Are Non-English Papers Reviewed Fairly? Language-of-Study Bias in NLP Peer Reviews
di: Barkhordar, Ehsan, et al.
Pubblicazione: (2026) -
Linguistically-Informed Multilingual Instruction Tuning: Is There an Optimal Set of Languages to Tune?
di: Soykan, Gürkan, et al.
Pubblicazione: (2024)