Guardado en:
| Autores principales: | Yang, Jingpu, Han, Zehua, Xiang, Mengyu, Wang, Helin, Huang, Yuxiao, Fang, Miao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2402.14849 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Enhancing Rare Codes via Probability-Biased Directed Graph Attention for Long-Tail ICD Coding
por: Chen, Tianlei, et al.
Publicado: (2025)
por: Chen, Tianlei, et al.
Publicado: (2025)
Large Language Models Are Bad Dice Players: LLMs Struggle to Generate Random Numbers from Statistical Distributions
por: Zhao, Minda, et al.
Publicado: (2026)
por: Zhao, Minda, et al.
Publicado: (2026)
PromptIntern: Saving Inference Costs by Internalizing Recurrent Prompt during Large Language Model Fine-tuning
por: Zou, Jiaru, et al.
Publicado: (2024)
por: Zou, Jiaru, et al.
Publicado: (2024)
Graph-enhanced Large Language Models in Asynchronous Plan Reasoning
por: Lin, Fangru, et al.
Publicado: (2024)
por: Lin, Fangru, et al.
Publicado: (2024)
Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models
por: Noukhovitch, Michael, et al.
Publicado: (2024)
por: Noukhovitch, Michael, et al.
Publicado: (2024)
Group Representational Position Encoding
por: Zhang, Yifan, et al.
Publicado: (2025)
por: Zhang, Yifan, et al.
Publicado: (2025)
Context-Aware Initialization for Reducing Generative Path Length in Diffusion Language Models
por: Miao, Tongyuan, et al.
Publicado: (2025)
por: Miao, Tongyuan, et al.
Publicado: (2025)
CasualSynth: Generating Structurally Sound Synthetic Data
por: Cheng, Zehua, et al.
Publicado: (2026)
por: Cheng, Zehua, et al.
Publicado: (2026)
Two Stones Hit One Bird: Bilevel Positional Encoding for Better Length Extrapolation
por: He, Zhenyu, et al.
Publicado: (2024)
por: He, Zhenyu, et al.
Publicado: (2024)
Scaling Bidirectional Spans and Span Violations in Attention Mechanism
por: Kim, Jongwook, et al.
Publicado: (2025)
por: Kim, Jongwook, et al.
Publicado: (2025)
Understanding Emergent Abilities of Language Models from the Loss Perspective
por: Du, Zhengxiao, et al.
Publicado: (2024)
por: Du, Zhengxiao, et al.
Publicado: (2024)
Recursive Introspection: Teaching Language Model Agents How to Self-Improve
por: Qu, Yuxiao, et al.
Publicado: (2024)
por: Qu, Yuxiao, et al.
Publicado: (2024)
Extensive Self-Contrast Enables Feedback-Free Language Model Alignment
por: Liu, Xiao, et al.
Publicado: (2024)
por: Liu, Xiao, et al.
Publicado: (2024)
AgenticCache: Cache-Driven Asynchronous Planning for Embodied AI Agents
por: Kim, Hojoon, et al.
Publicado: (2026)
por: Kim, Hojoon, et al.
Publicado: (2026)
AbsTopK: Rethinking Sparse Autoencoders For Bidirectional Features
por: Zhu, Xudong, et al.
Publicado: (2025)
por: Zhu, Xudong, et al.
Publicado: (2025)
Struc-EMB: The Potential of Structure-Aware Encoding in Language Embeddings
por: Liu, Shikun, et al.
Publicado: (2025)
por: Liu, Shikun, et al.
Publicado: (2025)
Parameter-Efficient Fine-Tuning for Foundation Models
por: Zhang, Dan, et al.
Publicado: (2025)
por: Zhang, Dan, et al.
Publicado: (2025)
AutoDetect: Towards a Unified Framework for Automated Weakness Detection in Large Language Models
por: Cheng, Jiale, et al.
Publicado: (2024)
por: Cheng, Jiale, et al.
Publicado: (2024)
Concurrency without Model Changes: Future-based Asynchronous Function Calling for LLMs
por: Feng, Guangyu, et al.
Publicado: (2026)
por: Feng, Guangyu, et al.
Publicado: (2026)
BiSup: Bidirectional Quantization Error Suppression for Large Language Models
por: Zou, Minghui, et al.
Publicado: (2024)
por: Zou, Minghui, et al.
Publicado: (2024)
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models
por: Cheng, Jiale, et al.
Publicado: (2024)
por: Cheng, Jiale, et al.
Publicado: (2024)
CaRT: Teaching LLM Agents to Know When They Know Enough
por: Liu, Grace, et al.
Publicado: (2025)
por: Liu, Grace, et al.
Publicado: (2025)
POPE: Learning to Reason on Hard Problems via Privileged On-Policy Exploration
por: Qu, Yuxiao, et al.
Publicado: (2026)
por: Qu, Yuxiao, et al.
Publicado: (2026)
Understanding Token Probability Encoding in Output Embeddings
por: Cho, Hakaze, et al.
Publicado: (2024)
por: Cho, Hakaze, et al.
Publicado: (2024)
Encoding Agent Trajectories as Representations with Sequence Transformers
por: Tsiligkaridis, Athanasios, et al.
Publicado: (2024)
por: Tsiligkaridis, Athanasios, et al.
Publicado: (2024)
The Bicameral Model: Bidirectional Hidden-State Coupling Between Parallel Language Models
por: Flamant, Cedric, et al.
Publicado: (2026)
por: Flamant, Cedric, et al.
Publicado: (2026)
Do LLMs Encode Functional Importance of Reasoning Tokens?
por: Singh, Janvijay, et al.
Publicado: (2026)
por: Singh, Janvijay, et al.
Publicado: (2026)
CNSight: Evaluation of Clinical Note Segmentation Tools
por: Surana, Risha, et al.
Publicado: (2025)
por: Surana, Risha, et al.
Publicado: (2025)
Maximizing Asynchronicity in Event-based Neural Networks
por: Hao, Haiqing, et al.
Publicado: (2025)
por: Hao, Haiqing, et al.
Publicado: (2025)
SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator
por: Chen, Guoxuan, et al.
Publicado: (2024)
por: Chen, Guoxuan, et al.
Publicado: (2024)
Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
por: Qu, Yuxiao, et al.
Publicado: (2025)
por: Qu, Yuxiao, et al.
Publicado: (2025)
SEE: Sememe Entanglement Encoding for Transformer-bases Models Compression
por: Zhang, Jing, et al.
Publicado: (2024)
por: Zhang, Jing, et al.
Publicado: (2024)
From Pruning to Grafting: Dynamic Knowledge Redistribution via Learnable Layer Fusion
por: Pei, Zehua, et al.
Publicado: (2024)
por: Pei, Zehua, et al.
Publicado: (2024)
LLMs Encode Their Failures: Predicting Success from Pre-Generation Activations
por: Lugoloobi, William, et al.
Publicado: (2026)
por: Lugoloobi, William, et al.
Publicado: (2026)
HoPE: A Novel Positional Encoding Without Long-Term Decay for Enhanced Context Awareness and Extrapolation
por: Chen, Yuhan, et al.
Publicado: (2024)
por: Chen, Yuhan, et al.
Publicado: (2024)
RLAD: Training LLMs to Discover Abstractions for Solving Reasoning Problems
por: Qu, Yuxiao, et al.
Publicado: (2025)
por: Qu, Yuxiao, et al.
Publicado: (2025)
MYTE: Morphology-Driven Byte Encoding for Better and Fairer Multilingual Language Modeling
por: Limisiewicz, Tomasz, et al.
Publicado: (2024)
por: Limisiewicz, Tomasz, et al.
Publicado: (2024)
Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization
por: Foroutan, Negar, et al.
Publicado: (2025)
por: Foroutan, Negar, et al.
Publicado: (2025)
Encode, Think, Decode: Scaling test-time reasoning with recursive latent thoughts
por: Koishekenov, Yeskendir, et al.
Publicado: (2025)
por: Koishekenov, Yeskendir, et al.
Publicado: (2025)
GATech at AbjadMed: Bidirectional Encoders vs. Causal Decoders: Insights from 82-Class Arabic Medical Classification
por: Khamis, Ahmed Khaled
Publicado: (2026)
por: Khamis, Ahmed Khaled
Publicado: (2026)
Ejemplares similares
-
Enhancing Rare Codes via Probability-Biased Directed Graph Attention for Long-Tail ICD Coding
por: Chen, Tianlei, et al.
Publicado: (2025) -
Large Language Models Are Bad Dice Players: LLMs Struggle to Generate Random Numbers from Statistical Distributions
por: Zhao, Minda, et al.
Publicado: (2026) -
PromptIntern: Saving Inference Costs by Internalizing Recurrent Prompt during Large Language Model Fine-tuning
por: Zou, Jiaru, et al.
Publicado: (2024) -
Graph-enhanced Large Language Models in Asynchronous Plan Reasoning
por: Lin, Fangru, et al.
Publicado: (2024) -
Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models
por: Noukhovitch, Michael, et al.
Publicado: (2024)