To Softmax, or not to Softmax: that is the question when applying Active Learning for Transformer Models
Fuente:
arXiv
Saved in:
| Main Authors: | Gonsior, Julius, Falkenberg, Christian, Magino, Silvio, Reusch, Anja, Thiele, Maik, Lehner, Wolfgang |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ImitAL: Learned Active Learning Strategy on Synthetic Data
by: Gonsior, Julius, et al.
Published: (2022)
by: Gonsior, Julius, et al.
Published: (2022)
Survey of Active Learning Hyperparameters: Insights from a Large-Scale Experimental Grid
by: Gonsior, Julius, et al.
Published: (2025)
by: Gonsior, Julius, et al.
Published: (2025)
In-Context Learning with Transformers: Softmax Attention Adapts to Function Lipschitzness
by: Collins, Liam, et al.
Published: (2024)
by: Collins, Liam, et al.
Published: (2024)
Forgetting Transformer: Softmax Attention with a Forget Gate
by: Lin, Zhixuan, et al.
Published: (2025)
by: Lin, Zhixuan, et al.
Published: (2025)
Scalable-Softmax Is Superior for Attention
by: Nakanishi, Ken M.
Published: (2025)
by: Nakanishi, Ken M.
Published: (2025)
The Information Geometry of Softmax: Probing and Steering
by: Park, Kiho, et al.
Published: (2026)
by: Park, Kiho, et al.
Published: (2026)
Vertex-Softmax: Tight Transformer Verification via Exact Softmax Optimization
by: Rezazadeh, Navid, et al.
Published: (2026)
by: Rezazadeh, Navid, et al.
Published: (2026)
Partially Recentralization Softmax Loss for Vision-Language Models Robustness
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
KU-DMIS at EHRSQL 2024:Generating SQL query via question templatization in EHR
by: Kim, Hajung, et al.
Published: (2024)
by: Kim, Hajung, et al.
Published: (2024)
SQL-Exchange: Transforming SQL Queries Across Domains
by: Daviran, Mohammadreza, et al.
Published: (2025)
by: Daviran, Mohammadreza, et al.
Published: (2025)
A Survey of Large Language Models on Generative Graph Analytics: Query, Learning, and Applications
by: Shang, Wenbo, et al.
Published: (2024)
by: Shang, Wenbo, et al.
Published: (2024)
Compressible Softmax-Attended Language under Incompressible Attention
by: Lee, Wonsuk
Published: (2026)
by: Lee, Wonsuk
Published: (2026)
Softmax-free Linear Transformers
by: Lu, Jiachen, et al.
Published: (2022)
by: Lu, Jiachen, et al.
Published: (2022)
Rationalization Models for Text-to-SQL
by: Rossiello, Gaetano, et al.
Published: (2025)
by: Rossiello, Gaetano, et al.
Published: (2025)
Schema-Aware Multi-Task Learning for Complex Text-to-SQL
by: Wu, Yangjun, et al.
Published: (2024)
by: Wu, Yangjun, et al.
Published: (2024)
E3-Rewrite: Learning to Rewrite SQL for Executability, Equivalence,and Efficiency
by: Xu, Dongjie, et al.
Published: (2025)
by: Xu, Dongjie, et al.
Published: (2025)
Dimensionality Reduction in Sentence Transformer Vector Databases with Fast Fourier Transform
by: Bulgakov, Vitaly, et al.
Published: (2024)
by: Bulgakov, Vitaly, et al.
Published: (2024)
Feather-SQL: A Lightweight NL2SQL Framework with Dual-Model Collaboration Paradigm for Small Language Models
by: Pei, Wenqi, et al.
Published: (2025)
by: Pei, Wenqi, et al.
Published: (2025)
GradeSQL: Test-Time Inference with Outcome Reward Models for Text-to-SQL Generation from Large Language Models
by: Tritto, Mattia, et al.
Published: (2025)
by: Tritto, Mattia, et al.
Published: (2025)
Structure Guided Large Language Model for SQL Generation
by: Zhang, Qinggang, et al.
Published: (2024)
by: Zhang, Qinggang, et al.
Published: (2024)
Are Large Language Models the New Interface for Data Pipelines?
by: Junior, Sylvio Barbon, et al.
Published: (2024)
by: Junior, Sylvio Barbon, et al.
Published: (2024)
Generating Tables from the Parametric Knowledge of Language Models
by: Berkovitch, Yevgeni, et al.
Published: (2024)
by: Berkovitch, Yevgeni, et al.
Published: (2024)
Are Large Language Models a Good Replacement of Taxonomies?
by: Sun, Yushi, et al.
Published: (2024)
by: Sun, Yushi, et al.
Published: (2024)
Analyzing the Effectiveness of Large Language Models on Text-to-SQL Synthesis
by: Roberson, Richard, et al.
Published: (2024)
by: Roberson, Richard, et al.
Published: (2024)
TableLlama: Towards Open Large Generalist Models for Tables
by: Zhang, Tianshu, et al.
Published: (2023)
by: Zhang, Tianshu, et al.
Published: (2023)
GraLMatch: Matching Groups of Entities with Graphs and Language Models
by: Pardo, Fernando De Meer, et al.
Published: (2024)
by: Pardo, Fernando De Meer, et al.
Published: (2024)
The Interpretability Analysis of the Model Can Bring Improvements to the Text-to-SQL Task
by: Zhang, Cong
Published: (2025)
by: Zhang, Cong
Published: (2025)
DeepAnalyze: Agentic Large Language Models for Autonomous Data Science
by: Zhang, Shaolei, et al.
Published: (2025)
by: Zhang, Shaolei, et al.
Published: (2025)
Semantic Operators: A Declarative Model for Rich, AI-based Data Processing
by: Patel, Liana, et al.
Published: (2024)
by: Patel, Liana, et al.
Published: (2024)
PURPLE: Making a Large Language Model a Better SQL Writer
by: Ren, Tonghui, et al.
Published: (2024)
by: Ren, Tonghui, et al.
Published: (2024)
SQL-PaLM: Improved Large Language Model Adaptation for Text-to-SQL (extended)
by: Sun, Ruoxi, et al.
Published: (2023)
by: Sun, Ruoxi, et al.
Published: (2023)
Multi-hop Question Answering over Knowledge Graphs using Large Language Models
by: Chakraborty, Abir
Published: (2024)
by: Chakraborty, Abir
Published: (2024)
Schema Lineage Extraction at Scale: Multilingual Pipelines, Composite Evaluation, and Language-Model Benchmarks
by: Yin, Jiaqi, et al.
Published: (2025)
by: Yin, Jiaqi, et al.
Published: (2025)
AnyMatch -- Efficient Zero-Shot Entity Matching with a Small Language Model
by: Zhang, Zeyu, et al.
Published: (2024)
by: Zhang, Zeyu, et al.
Published: (2024)
Evaluating the Data Model Robustness of Text-to-SQL Systems Based on Real User Queries
by: Fürst, Jonathan, et al.
Published: (2024)
by: Fürst, Jonathan, et al.
Published: (2024)
FinSQL: Model-Agnostic LLMs-based Text-to-SQL Framework for Financial Analysis
by: Zhang, Chao, et al.
Published: (2024)
by: Zhang, Chao, et al.
Published: (2024)
Towards Improving Interpretability of Language Model Generation through a Structured Knowledge Discovery Approach
by: Liu, Shuqi, et al.
Published: (2025)
by: Liu, Shuqi, et al.
Published: (2025)
Aligning Large Language Models to a Domain-specific Graph Database for NL2GQL
by: Liang, Yuanyuan, et al.
Published: (2024)
by: Liang, Yuanyuan, et al.
Published: (2024)
Automated Data Visualization from Natural Language via Large Language Models: An Exploratory Study
by: Wu, Yang, et al.
Published: (2024)
by: Wu, Yang, et al.
Published: (2024)
ExeSQL: Self-Taught Text-to-SQL Models with Execution-Driven Bootstrapping for SQL Dialects
by: Zhang, Jipeng, et al.
Published: (2025)
by: Zhang, Jipeng, et al.
Published: (2025)
Similar Items
-
ImitAL: Learned Active Learning Strategy on Synthetic Data
by: Gonsior, Julius, et al.
Published: (2022) -
Survey of Active Learning Hyperparameters: Insights from a Large-Scale Experimental Grid
by: Gonsior, Julius, et al.
Published: (2025) -
In-Context Learning with Transformers: Softmax Attention Adapts to Function Lipschitzness
by: Collins, Liam, et al.
Published: (2024) -
Forgetting Transformer: Softmax Attention with a Forget Gate
by: Lin, Zhixuan, et al.
Published: (2025) -
Scalable-Softmax Is Superior for Attention
by: Nakanishi, Ken M.
Published: (2025)