Converting Transformers into DGNNs Form
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Jie, Mao, Mao-Hsuan, Chiu, Bo-Wei, Sun, Min-Te |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GPG: Generalized Policy Gradient Theorem for Transformer-based Policies
von: Mao, Hangyu, et al.
Veröffentlicht: (2025)
von: Mao, Hangyu, et al.
Veröffentlicht: (2025)
A Survey of On-Policy Distillation for Large Language Models
von: Song, Mingyang, et al.
Veröffentlicht: (2026)
von: Song, Mingyang, et al.
Veröffentlicht: (2026)
Recurrent Confidence Chain: Temporal-Aware Uncertainty Quantification in Large Language Models
von: Mao, Zhenjiang, et al.
Veröffentlicht: (2026)
von: Mao, Zhenjiang, et al.
Veröffentlicht: (2026)
Thinking Deeper, Not Longer: Depth-Recurrent Transformers for Compositional Generalization
von: Chen, Hung-Hsuan
Veröffentlicht: (2026)
von: Chen, Hung-Hsuan
Veröffentlicht: (2026)
Weights to Code: Extracting Interpretable Algorithms from the Discrete Transformer
von: Zhang, Yifan, et al.
Veröffentlicht: (2026)
von: Zhang, Yifan, et al.
Veröffentlicht: (2026)
Grounding Language Plans in Demonstrations Through Counterfactual Perturbations
von: Wang, Yanwei, et al.
Veröffentlicht: (2024)
von: Wang, Yanwei, et al.
Veröffentlicht: (2024)
ToMoE: Converting Dense Large Language Models to Mixture-of-Experts through Dynamic Structural Pruning
von: Gao, Shangqian, et al.
Veröffentlicht: (2025)
von: Gao, Shangqian, et al.
Veröffentlicht: (2025)
Lossless Compression of Large Language Model-Generated Text via Next-Token Prediction
von: Mao, Yu, et al.
Veröffentlicht: (2025)
von: Mao, Yu, et al.
Veröffentlicht: (2025)
LIRE: listwise reward enhancement for preference alignment
von: Zhu, Mingye, et al.
Veröffentlicht: (2024)
von: Zhu, Mingye, et al.
Veröffentlicht: (2024)
Differential Transformer
von: Ye, Tianzhu, et al.
Veröffentlicht: (2024)
von: Ye, Tianzhu, et al.
Veröffentlicht: (2024)
Can LLMs Convert Graphs to Text-Attributed Graphs?
von: Wang, Zehong, et al.
Veröffentlicht: (2024)
von: Wang, Zehong, et al.
Veröffentlicht: (2024)
RAFT: Realistic Attacks to Fool Text Detectors
von: Wang, James, et al.
Veröffentlicht: (2024)
von: Wang, James, et al.
Veröffentlicht: (2024)
AdaDPO: Self-Adaptive Direct Preference Optimization with Balanced Gradient Updates
von: Chen, Shaolong, et al.
Veröffentlicht: (2026)
von: Chen, Shaolong, et al.
Veröffentlicht: (2026)
FAAST: Forward-Only Associative Learning via Closed-Form Fast Weights for Test-Time Supervised Adaptation
von: Bao, Guangsheng, et al.
Veröffentlicht: (2026)
von: Bao, Guangsheng, et al.
Veröffentlicht: (2026)
Automata Extraction from Transformers
von: Zhang, Yihao, et al.
Veröffentlicht: (2024)
von: Zhang, Yihao, et al.
Veröffentlicht: (2024)
Learning Novel Transformer Architecture for Time-series Forecasting
von: Zhang, Juyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Juyuan, et al.
Veröffentlicht: (2025)
Generating Fine Details of Entity Interactions
von: Gu, Xinyi, et al.
Veröffentlicht: (2025)
von: Gu, Xinyi, et al.
Veröffentlicht: (2025)
RLAC: Reinforcement Learning with Adversarial Critic for Free-Form Generation Tasks
von: Wu, Mian, et al.
Veröffentlicht: (2025)
von: Wu, Mian, et al.
Veröffentlicht: (2025)
Uncertainty-Aware Exploratory Direct Preference Optimization for Multimodal Large Language Models
von: Zhang, Huatian, et al.
Veröffentlicht: (2026)
von: Zhang, Huatian, et al.
Veröffentlicht: (2026)
Keyframe-oriented Vision Token Pruning: Enhancing Efficiency of Large Vision Language Models on Long-Form Video Processing
von: Liu, Yudong, et al.
Veröffentlicht: (2025)
von: Liu, Yudong, et al.
Veröffentlicht: (2025)
Transformers with Selective Access to Early Representations
von: Gunasekaran, Skye, et al.
Veröffentlicht: (2026)
von: Gunasekaran, Skye, et al.
Veröffentlicht: (2026)
Finite State Automata Inside Transformers with Chain-of-Thought: A Mechanistic Study on State Tracking
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
Advancing Graph Representation Learning with Large Language Models: A Comprehensive Survey of Techniques
von: Mao, Qiheng, et al.
Veröffentlicht: (2024)
von: Mao, Qiheng, et al.
Veröffentlicht: (2024)
Teaching Language Models to Critique via Reinforcement Learning
von: Xie, Zhihui, et al.
Veröffentlicht: (2025)
von: Xie, Zhihui, et al.
Veröffentlicht: (2025)
Deterministic Inference across Tensor Parallel Sizes That Eliminates Training-Inference Mismatch
von: Zhang, Ziyang, et al.
Veröffentlicht: (2025)
von: Zhang, Ziyang, et al.
Veröffentlicht: (2025)
LARGO: Latent Adversarial Reflection through Gradient Optimization for Jailbreaking LLMs
von: Li, Ran, et al.
Veröffentlicht: (2025)
von: Li, Ran, et al.
Veröffentlicht: (2025)
DNAZEN: Enhanced Gene Sequence Representations via Mixed Granularities of Coding Units
von: Mao, Lei, et al.
Veröffentlicht: (2025)
von: Mao, Lei, et al.
Veröffentlicht: (2025)
CHAMP: A Competition-level Dataset for Fine-Grained Analyses of LLMs' Mathematical Reasoning Capabilities
von: Mao, Yujun, et al.
Veröffentlicht: (2024)
von: Mao, Yujun, et al.
Veröffentlicht: (2024)
SelfIE: Self-Interpretation of Large Language Model Embeddings
von: Chen, Haozhe, et al.
Veröffentlicht: (2024)
von: Chen, Haozhe, et al.
Veröffentlicht: (2024)
Convex Dominance in Deep Learning I: A Scaling Law of Loss and Learning Rate
von: Bu, Zhiqi, et al.
Veröffentlicht: (2026)
von: Bu, Zhiqi, et al.
Veröffentlicht: (2026)
KnowHalu: Hallucination Detection via Multi-Form Knowledge Based Factual Checking
von: Zhang, Jiawei, et al.
Veröffentlicht: (2024)
von: Zhang, Jiawei, et al.
Veröffentlicht: (2024)
Token-level Accept or Reject: A Micro Alignment Approach for Large Language Models
von: Zhang, Yang, et al.
Veröffentlicht: (2025)
von: Zhang, Yang, et al.
Veröffentlicht: (2025)
Information-Theoretic Reward Decomposition for Generalizable RLHF
von: Mao, Liyuan, et al.
Veröffentlicht: (2025)
von: Mao, Liyuan, et al.
Veröffentlicht: (2025)
An Integrated Data Processing Framework for Pretraining Foundation Models
von: Sun, Yiding, et al.
Veröffentlicht: (2024)
von: Sun, Yiding, et al.
Veröffentlicht: (2024)
Fine-Grained Alignment in Vision-and-Language Navigation through Bayesian Optimization
von: Song, Yuhang, et al.
Veröffentlicht: (2024)
von: Song, Yuhang, et al.
Veröffentlicht: (2024)
On the Convergence of Moral Self-Correction in Large Language Models
von: Liu, Guangliang, et al.
Veröffentlicht: (2025)
von: Liu, Guangliang, et al.
Veröffentlicht: (2025)
BBox-Adapter: Lightweight Adapting for Black-Box Large Language Models
von: Sun, Haotian, et al.
Veröffentlicht: (2024)
von: Sun, Haotian, et al.
Veröffentlicht: (2024)
Multi-modal Anchor Gated Transformer with Knowledge Distillation for Emotion Recognition in Conversation
von: Li, Jie, et al.
Veröffentlicht: (2025)
von: Li, Jie, et al.
Veröffentlicht: (2025)
Enhancing Latent Computation in Transformers with Latent Tokens
von: Sun, Yuchang, et al.
Veröffentlicht: (2025)
von: Sun, Yuchang, et al.
Veröffentlicht: (2025)
How do Large Language Models Understand Relevance? A Mechanistic Interpretability Perspective
von: Liu, Qi, et al.
Veröffentlicht: (2025)
von: Liu, Qi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
GPG: Generalized Policy Gradient Theorem for Transformer-based Policies
von: Mao, Hangyu, et al.
Veröffentlicht: (2025) -
A Survey of On-Policy Distillation for Large Language Models
von: Song, Mingyang, et al.
Veröffentlicht: (2026) -
Recurrent Confidence Chain: Temporal-Aware Uncertainty Quantification in Large Language Models
von: Mao, Zhenjiang, et al.
Veröffentlicht: (2026) -
Thinking Deeper, Not Longer: Depth-Recurrent Transformers for Compositional Generalization
von: Chen, Hung-Hsuan
Veröffentlicht: (2026) -
Weights to Code: Extracting Interpretable Algorithms from the Discrete Transformer
von: Zhang, Yifan, et al.
Veröffentlicht: (2026)