Anchor function: a type of benchmark functions for studying language models
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Zhongwang, Wang, Zhiwei, Yao, Junjie, Zhou, Zhangchen, Li, Xiaolong, E, Weinan, Xu, Zhi-Qin John |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
An Analysis for Reasoning Bias of Language Models with Small Initialization
por: Yao, Junjie, et al.
Publicado: (2025)
por: Yao, Junjie, et al.
Publicado: (2025)
Complexity Control Facilitates Reasoning-Based Compositional Generalization in Transformers
por: Zhang, Zhongwang, et al.
Publicado: (2025)
por: Zhang, Zhongwang, et al.
Publicado: (2025)
Understanding the Language Model to Solve the Symbolic Multi-Step Reasoning Problem from the Perspective of Buffer Mechanism
por: Wang, Zhiwei, et al.
Publicado: (2024)
por: Wang, Zhiwei, et al.
Publicado: (2024)
Reasoning Bias of Next Token Prediction Training
por: Lin, Pengxiao, et al.
Publicado: (2025)
por: Lin, Pengxiao, et al.
Publicado: (2025)
Loss Spike in Training Neural Networks
por: Li, Xiaolong, et al.
Publicado: (2023)
por: Li, Xiaolong, et al.
Publicado: (2023)
Loss Jump During Loss Switch in Solving PDEs with Neural Networks
por: Wang, Zhiwei, et al.
Publicado: (2024)
por: Wang, Zhiwei, et al.
Publicado: (2024)
An overview of condensation phenomenon in deep learning
por: Xu, Zhi-Qin John, et al.
Publicado: (2025)
por: Xu, Zhi-Qin John, et al.
Publicado: (2025)
Initialization is Critical to Whether Transformers Fit Composite Functions by Reasoning or Memorizing
por: Zhang, Zhongwang, et al.
Publicado: (2024)
por: Zhang, Zhongwang, et al.
Publicado: (2024)
Scalable Complexity Control Facilitates Reasoning Ability of LLMs
por: Hang, Liangkai, et al.
Publicado: (2025)
por: Hang, Liangkai, et al.
Publicado: (2025)
A rationale from frequency perspective for grokking in training neural network
por: Zhou, Zhangchen, et al.
Publicado: (2024)
por: Zhou, Zhangchen, et al.
Publicado: (2024)
Adaptive Preconditioners Trigger Loss Spikes in Adam
por: Bai, Zhiwei, et al.
Publicado: (2025)
por: Bai, Zhiwei, et al.
Publicado: (2025)
The SMeL Test: A simple benchmark for media literacy in language models
por: Ahdritz, Gustaf, et al.
Publicado: (2025)
por: Ahdritz, Gustaf, et al.
Publicado: (2025)
BabySLM: language-acquisition-friendly benchmark of self-supervised spoken language models
por: Lavechin, Marvin, et al.
Publicado: (2023)
por: Lavechin, Marvin, et al.
Publicado: (2023)
DevBench: A multimodal developmental benchmark for language learning
por: Tan, Alvin Wei Ming, et al.
Publicado: (2024)
por: Tan, Alvin Wei Ming, et al.
Publicado: (2024)
A dataset and benchmark for hospital course summarization with adapted large language models
por: Aali, Asad, et al.
Publicado: (2024)
por: Aali, Asad, et al.
Publicado: (2024)
A comparative study of zero-shot inference with large language models and supervised modeling in breast cancer pathology classification
por: Sushil, Madhumita, et al.
Publicado: (2024)
por: Sushil, Madhumita, et al.
Publicado: (2024)
KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding
por: Xu, Zhangchen, et al.
Publicado: (2025)
por: Xu, Zhangchen, et al.
Publicado: (2025)
Probability Signature: Bridging Data Semantics and Embedding Structure in Language Models
por: Yao, Junjie, et al.
Publicado: (2025)
por: Yao, Junjie, et al.
Publicado: (2025)
Local Linear Recovery Guarantee of Deep Neural Networks at Overparameterization
por: Zhang, Yaoyu, et al.
Publicado: (2024)
por: Zhang, Yaoyu, et al.
Publicado: (2024)
Zero-shot data citation function classification using transformer-based large language models (LLMs)
por: Byers, Neil, et al.
Publicado: (2025)
por: Byers, Neil, et al.
Publicado: (2025)
Reinforcement Learning with Rubric Anchors
por: Huang, Zenan, et al.
Publicado: (2025)
por: Huang, Zenan, et al.
Publicado: (2025)
AdvAnchor: Enhancing Diffusion Model Unlearning with Adversarial Anchors
por: Zhao, Mengnan, et al.
Publicado: (2024)
por: Zhao, Mengnan, et al.
Publicado: (2024)
Do language models plan ahead for future tokens?
por: Wu, Wilson, et al.
Publicado: (2024)
por: Wu, Wilson, et al.
Publicado: (2024)
Machine-assisted writing evaluation: Exploring pre-trained language models in analyzing argumentative moves
por: Qin, Wenjuan, et al.
Publicado: (2025)
por: Qin, Wenjuan, et al.
Publicado: (2025)
Adaptive Preference Optimization with Uncertainty-aware Utility Anchor
por: Wang, Xiaobo, et al.
Publicado: (2025)
por: Wang, Xiaobo, et al.
Publicado: (2025)
Constraining Sequential Model Editing with Editing Anchor Compression
por: Xu, Hao-Xiang, et al.
Publicado: (2025)
por: Xu, Hao-Xiang, et al.
Publicado: (2025)
Multi-modal Anchor Gated Transformer with Knowledge Distillation for Emotion Recognition in Conversation
por: Li, Jie, et al.
Publicado: (2025)
por: Li, Jie, et al.
Publicado: (2025)
Question answering system of bridge design specification based on large language model
por: Zhang, Leye, et al.
Publicado: (2024)
por: Zhang, Leye, et al.
Publicado: (2024)
Aligning language models with human preferences
por: Korbak, Tomasz
Publicado: (2024)
por: Korbak, Tomasz
Publicado: (2024)
Evaluating language models as risk scores
por: Cruz, André F., et al.
Publicado: (2024)
por: Cruz, André F., et al.
Publicado: (2024)
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning
por: Xu, Zhangchen, et al.
Publicado: (2025)
por: Xu, Zhangchen, et al.
Publicado: (2025)
Amortizing intractable inference in large language models
por: Hu, Edward J., et al.
Publicado: (2023)
por: Hu, Edward J., et al.
Publicado: (2023)
ESI: Epistemic Uncertainty Quantification via Semantic-preserving Intervention for Large Language Models
por: Li, Mingda, et al.
Publicado: (2025)
por: Li, Mingda, et al.
Publicado: (2025)
Bridging vision language model (VLM) evaluation gaps with a framework for scalable and cost-effective benchmark generation
por: Rädsch, Tim, et al.
Publicado: (2025)
por: Rädsch, Tim, et al.
Publicado: (2025)
Generative adversarial networks vs large language models: a comparative study on synthetic tabular data generation
por: Barr, Austin A., et al.
Publicado: (2025)
por: Barr, Austin A., et al.
Publicado: (2025)
Simple linear attention language models balance the recall-throughput tradeoff
por: Arora, Simran, et al.
Publicado: (2024)
por: Arora, Simran, et al.
Publicado: (2024)
The language of time: a language model perspective on time-series foundation models
por: Xie, Yi, et al.
Publicado: (2025)
por: Xie, Yi, et al.
Publicado: (2025)
Linear representations in language models can change dramatically over a conversation
por: Lampinen, Andrew Kyle, et al.
Publicado: (2026)
por: Lampinen, Andrew Kyle, et al.
Publicado: (2026)
On the generalization of language models from in-context learning and finetuning: a controlled study
por: Lampinen, Andrew K., et al.
Publicado: (2025)
por: Lampinen, Andrew K., et al.
Publicado: (2025)
Perturbed examples reveal invariances shared by language models
por: Rawal, Ruchit, et al.
Publicado: (2023)
por: Rawal, Ruchit, et al.
Publicado: (2023)
Ejemplares similares
-
An Analysis for Reasoning Bias of Language Models with Small Initialization
por: Yao, Junjie, et al.
Publicado: (2025) -
Complexity Control Facilitates Reasoning-Based Compositional Generalization in Transformers
por: Zhang, Zhongwang, et al.
Publicado: (2025) -
Understanding the Language Model to Solve the Symbolic Multi-Step Reasoning Problem from the Perspective of Buffer Mechanism
por: Wang, Zhiwei, et al.
Publicado: (2024) -
Reasoning Bias of Next Token Prediction Training
por: Lin, Pengxiao, et al.
Publicado: (2025) -
Loss Spike in Training Neural Networks
por: Li, Xiaolong, et al.
Publicado: (2023)