Executing Natural Language-Described Algorithms with Large Language Models: An Investigation
Fuente:
arXiv
Salvato in:
| Autori principali: | Zheng, Xin, Zhu, Qiming, Lin, Hongyu, Lu, Yaojie, Han, Xianpei, Sun, Le |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Meta-Cognitive Analysis: Evaluating Declarative and Procedural Knowledge in Datasets and Large Language Models
di: Li, Zhuoqun, et al.
Pubblicazione: (2024)
di: Li, Zhuoqun, et al.
Pubblicazione: (2024)
Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models
di: Zheng, Jiasheng, et al.
Pubblicazione: (2024)
di: Zheng, Jiasheng, et al.
Pubblicazione: (2024)
DeepRAG: Thinking to Retrieve Step by Step for Large Language Models
di: Guan, Xinyan, et al.
Pubblicazione: (2025)
di: Guan, Xinyan, et al.
Pubblicazione: (2025)
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models
di: Liu, Yanjiang, et al.
Pubblicazione: (2025)
di: Liu, Yanjiang, et al.
Pubblicazione: (2025)
Coupled Variational Reinforcement Learning for Language Model General Reasoning
di: Wen, Xueru, et al.
Pubblicazione: (2025)
di: Wen, Xueru, et al.
Pubblicazione: (2025)
When Models Outthink Their Safety: Unveiling and Mitigating Self-Jailbreak in Large Reasoning Models
di: Mao, Yingzhi, et al.
Pubblicazione: (2025)
di: Mao, Yingzhi, et al.
Pubblicazione: (2025)
Self-Retrieval: End-to-End Information Retrieval with One Large Language Model
di: Tang, Qiaoyu, et al.
Pubblicazione: (2024)
di: Tang, Qiaoyu, et al.
Pubblicazione: (2024)
StructEval: Deepen and Broaden Large Language Model Assessment via Structured Evaluation
di: Cao, Boxi, et al.
Pubblicazione: (2024)
di: Cao, Boxi, et al.
Pubblicazione: (2024)
Rule or Story, Which is a Better Commonsense Expression for Talking with Large Language Models?
di: Bian, Ning, et al.
Pubblicazione: (2024)
di: Bian, Ning, et al.
Pubblicazione: (2024)
PPTAgent: Generating and Evaluating Presentations Beyond Text-to-Slides
di: Zheng, Hao, et al.
Pubblicazione: (2025)
di: Zheng, Hao, et al.
Pubblicazione: (2025)
DeepSolution: Boosting Complex Engineering Solution Design via Tree-based Exploration and Bi-point Thinking
di: Li, Zhuoqun, et al.
Pubblicazione: (2025)
di: Li, Zhuoqun, et al.
Pubblicazione: (2025)
MemSearcher: Training LLMs to Reason, Search and Manage Memory via End-to-End Reinforcement Learning
di: Yuan, Qianhao, et al.
Pubblicazione: (2025)
di: Yuan, Qianhao, et al.
Pubblicazione: (2025)
Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation
di: Yuan, Qianhao, et al.
Pubblicazione: (2026)
di: Yuan, Qianhao, et al.
Pubblicazione: (2026)
ChatGPT is a Knowledgeable but Inexperienced Solver: An Investigation of Commonsense Problem in Large Language Models
di: Bian, Ning, et al.
Pubblicazione: (2023)
di: Bian, Ning, et al.
Pubblicazione: (2023)
ScaleBox: Enabling High-Fidelity and Scalable Code Verification for Large Language Models
di: Zheng, Jiasheng, et al.
Pubblicazione: (2026)
di: Zheng, Jiasheng, et al.
Pubblicazione: (2026)
Towards Real-world Human Behavior Simulation: Benchmarking Large Language Models on Long-horizon, Cross-scenario, Heterogeneous Behavior Traces
di: Chen, Jiawei, et al.
Pubblicazione: (2026)
di: Chen, Jiawei, et al.
Pubblicazione: (2026)
Memorizing is Not Enough: Deep Knowledge Injection Through Reasoning
di: Xu, Ruoxi, et al.
Pubblicazione: (2025)
di: Xu, Ruoxi, et al.
Pubblicazione: (2025)
Retentive or Forgetful? Diving into the Knowledge Memorizing Mechanism of Language Models
di: Cao, Boxi, et al.
Pubblicazione: (2023)
di: Cao, Boxi, et al.
Pubblicazione: (2023)
Beyond Text-Dominance: Understanding Modality Preference of Omni-modal Large Language Models
di: Yan, Xinru, et al.
Pubblicazione: (2026)
di: Yan, Xinru, et al.
Pubblicazione: (2026)
Rethinking Reward Model Evaluation: Are We Barking up the Wrong Tree?
di: Wen, Xueru, et al.
Pubblicazione: (2024)
di: Wen, Xueru, et al.
Pubblicazione: (2024)
URL: Universal Referential Knowledge Linking via Task-instructed Representation Compression
di: Li, Zhuoqun, et al.
Pubblicazione: (2024)
di: Li, Zhuoqun, et al.
Pubblicazione: (2024)
StructRAG: Boosting Knowledge Intensive Reasoning of LLMs via Inference-time Hybrid Information Structurization
di: Li, Zhuoqun, et al.
Pubblicazione: (2024)
di: Li, Zhuoqun, et al.
Pubblicazione: (2024)
LiveMCPBench: Can Agents Navigate an Ocean of MCP Tools?
di: Mo, Guozhao, et al.
Pubblicazione: (2025)
di: Mo, Guozhao, et al.
Pubblicazione: (2025)
Your Teacher Can't Help You Here: Combating Supervision Fidelity Decay in On-Policy Distillation
di: Liu, Yanjiang, et al.
Pubblicazione: (2026)
di: Liu, Yanjiang, et al.
Pubblicazione: (2026)
The Devil Is in the Details: Tackling Unimodal Spurious Correlations for Generalizable Multimodal Reward Models
di: Li, Zichao, et al.
Pubblicazione: (2025)
di: Li, Zichao, et al.
Pubblicazione: (2025)
Translating Regulatory Clauses into Executable Codes for Building Design Checking via Large Language Model Driven Function Matching and Composing
di: Zheng, Zhe, et al.
Pubblicazione: (2023)
di: Zheng, Zhe, et al.
Pubblicazione: (2023)
Decoupling Reasoning and Confidence: Resurrecting Calibration in Reinforcement Learning from Verifiable Rewards
di: Ma, Zhengzhao, et al.
Pubblicazione: (2026)
di: Ma, Zhengzhao, et al.
Pubblicazione: (2026)
The Rise and Down of Babel Tower: Investigating the Evolution Process of Multilingual Code Large Language Model
di: Chen, Jiawei, et al.
Pubblicazione: (2024)
di: Chen, Jiawei, et al.
Pubblicazione: (2024)
SAISA: Towards Multimodal Large Language Models with Both Training and Inference Efficiency
di: Yuan, Qianhao, et al.
Pubblicazione: (2025)
di: Yuan, Qianhao, et al.
Pubblicazione: (2025)
Cheems: A Practical Guidance for Building and Evaluating Chinese Reward Models from Scratch
di: Wen, Xueru, et al.
Pubblicazione: (2025)
di: Wen, Xueru, et al.
Pubblicazione: (2025)
Search, Verify and Feedback: Towards Next Generation Post-training Paradigm of Foundation Models via Verifier Engineering
di: Guan, Xinyan, et al.
Pubblicazione: (2024)
di: Guan, Xinyan, et al.
Pubblicazione: (2024)
From Informal to Formal -- Incorporating and Evaluating LLMs on Natural Language Requirements to Verifiable Formal Proofs
di: Cao, Jialun, et al.
Pubblicazione: (2025)
di: Cao, Jialun, et al.
Pubblicazione: (2025)
DOMAINEVAL: An Auto-Constructed Benchmark for Multi-Domain Code Generation
di: Zhu, Qiming, et al.
Pubblicazione: (2024)
di: Zhu, Qiming, et al.
Pubblicazione: (2024)
AI-Salesman: Towards Reliable Large Language Model Driven Telemarketing
di: Zhang, Qingyu, et al.
Pubblicazione: (2025)
di: Zhang, Qingyu, et al.
Pubblicazione: (2025)
Self-Steering Optimization: Autonomous Preference Optimization for Large Language Models
di: Xiang, Hao, et al.
Pubblicazione: (2024)
di: Xiang, Hao, et al.
Pubblicazione: (2024)
ShortV: Efficient Multimodal Large Language Models by Freezing Visual Tokens in Ineffective Layers
di: Yuan, Qianhao, et al.
Pubblicazione: (2025)
di: Yuan, Qianhao, et al.
Pubblicazione: (2025)
Investigating Instruction Tuning Large Language Models on Graphs
di: Zhu, Kerui, et al.
Pubblicazione: (2024)
di: Zhu, Kerui, et al.
Pubblicazione: (2024)
ShortGPT: Layers in Large Language Models are More Redundant Than You Expect
di: Men, Xin, et al.
Pubblicazione: (2024)
di: Men, Xin, et al.
Pubblicazione: (2024)
Expanding the Boundaries of Vision Prior Knowledge in Multi-modal Large Language Models
di: Liang, Qiao, et al.
Pubblicazione: (2025)
di: Liang, Qiao, et al.
Pubblicazione: (2025)
Towards Scalable Automated Alignment of LLMs: A Survey
di: Cao, Boxi, et al.
Pubblicazione: (2024)
di: Cao, Boxi, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Meta-Cognitive Analysis: Evaluating Declarative and Procedural Knowledge in Datasets and Large Language Models
di: Li, Zhuoqun, et al.
Pubblicazione: (2024) -
Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models
di: Zheng, Jiasheng, et al.
Pubblicazione: (2024) -
DeepRAG: Thinking to Retrieve Step by Step for Large Language Models
di: Guan, Xinyan, et al.
Pubblicazione: (2025) -
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models
di: Liu, Yanjiang, et al.
Pubblicazione: (2025) -
Coupled Variational Reinforcement Learning for Language Model General Reasoning
di: Wen, Xueru, et al.
Pubblicazione: (2025)