A General Framework for Clustering and Distribution Matching with Bandit Feedback
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yavas, Recep Can, Huang, Yuqi, Tan, Vincent Y. F., Scarlett, Jonathan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Adaptive Exploration for Latent-State Bandits
von: Jin, Jikai, et al.
Veröffentlicht: (2026)
von: Jin, Jikai, et al.
Veröffentlicht: (2026)
Contrastive MIM: A Contrastive Mutual Information Framework for Unified Generative and Discriminative Representation Learning
von: Livne, Micha
Veröffentlicht: (2025)
von: Livne, Micha
Veröffentlicht: (2025)
5C Prompt Contracts: A Minimalist, Creative-Friendly, Token-Efficient Design Framework for Individual and SME LLM Usage
von: Ari, Ugur
Veröffentlicht: (2025)
von: Ari, Ugur
Veröffentlicht: (2025)
Adaptive Weighting in Knowledge Distillation: An Axiomatic Framework for Multi-Scale Teacher Ensemble Optimization
von: Flouro, Aaron R., et al.
Veröffentlicht: (2026)
von: Flouro, Aaron R., et al.
Veröffentlicht: (2026)
Multi-Teacher Ensemble Distillation: A Mathematical Framework for Probability-Domain Knowledge Aggregation
von: Flouro, Aaron R., et al.
Veröffentlicht: (2026)
von: Flouro, Aaron R., et al.
Veröffentlicht: (2026)
Federated Learning based on Self-Evolving Gaussian Clustering
von: Ožbot, Miha, et al.
Veröffentlicht: (2025)
von: Ožbot, Miha, et al.
Veröffentlicht: (2025)
Generative Pretrained Embedding and Hierarchical Irregular Time Series Representation for Daily Living Activity Recognition
von: Bouchabou, Damien, et al.
Veröffentlicht: (2024)
von: Bouchabou, Damien, et al.
Veröffentlicht: (2024)
Adversarial Learning in Games with Bandit Feedback: Logarithmic Pure-Strategy Maximin Regret
von: Ito, Shinji, et al.
Veröffentlicht: (2026)
von: Ito, Shinji, et al.
Veröffentlicht: (2026)
Repetition Makes Perfect: Recurrent Graph Neural Networks Match Message-Passing Limit
von: Rosenbluth, Eran, et al.
Veröffentlicht: (2025)
von: Rosenbluth, Eran, et al.
Veröffentlicht: (2025)
Dueling DDQN-Based Adaptive Multi-Objective Handover Optimization for LEO Satellite Networks
von: Chou, Po-Heng, et al.
Veröffentlicht: (2026)
von: Chou, Po-Heng, et al.
Veröffentlicht: (2026)
Step-E: A Differentiable Data Cleaning Framework for Robust Learning with Noisy Labels
von: Du, Wenzhang
Veröffentlicht: (2025)
von: Du, Wenzhang
Veröffentlicht: (2025)
Human-Aligned Skill Discovery: Balancing Behaviour Exploration and Alignment
von: Hussonnois, Maxence, et al.
Veröffentlicht: (2025)
von: Hussonnois, Maxence, et al.
Veröffentlicht: (2025)
Evaluating Model-Agnostic Meta-Learning on MetaWorld ML10 Benchmark: Fast Adaptation in Robotic Manipulation Tasks
von: Atamuradov, Sanjar
Veröffentlicht: (2025)
von: Atamuradov, Sanjar
Veröffentlicht: (2025)
Information flow in multilayer perceptrons: an in-depth analysis
von: Armano, Giuliano
Veröffentlicht: (2025)
von: Armano, Giuliano
Veröffentlicht: (2025)
Effects of Initialization Biases on Deep Neural Network Training Dynamics
von: Pellegrino, Nicholas, et al.
Veröffentlicht: (2025)
von: Pellegrino, Nicholas, et al.
Veröffentlicht: (2025)
Contrastive Mutual Information Learning: Toward Robust Representations without Positive-Pair Augmentations
von: Livne, Micha
Veröffentlicht: (2025)
von: Livne, Micha
Veröffentlicht: (2025)
When Redundancy Matters: Machine Teaching of Representations
von: Ferri, Cèsar, et al.
Veröffentlicht: (2024)
von: Ferri, Cèsar, et al.
Veröffentlicht: (2024)
Data structure > labels? Unsupervised heuristics for SVM hyperparameter estimation
von: Cholewa, Michał, et al.
Veröffentlicht: (2021)
von: Cholewa, Michał, et al.
Veröffentlicht: (2021)
Reinforcement Learning for Stock Transactions
von: Zhou, Ziyi, et al.
Veröffentlicht: (2025)
von: Zhou, Ziyi, et al.
Veröffentlicht: (2025)
Understanding the Limits of Deep Tabular Methods with Temporal Shift
von: Cai, Hao-Run, et al.
Veröffentlicht: (2025)
von: Cai, Hao-Run, et al.
Veröffentlicht: (2025)
RVI-SAC: Average Reward Off-Policy Deep Reinforcement Learning
von: Hisaki, Yukinari, et al.
Veröffentlicht: (2024)
von: Hisaki, Yukinari, et al.
Veröffentlicht: (2024)
Grammar-based evolutionary approach for automated workflow composition with domain-specific operators and ensemble diversity
von: Barbudo, Rafael, et al.
Veröffentlicht: (2024)
von: Barbudo, Rafael, et al.
Veröffentlicht: (2024)
Feature-aware Modulation for Learning from Temporal Tabular Data
von: Cai, Hao-Run, et al.
Veröffentlicht: (2025)
von: Cai, Hao-Run, et al.
Veröffentlicht: (2025)
Temporal Taskification in Streaming Continual Learning: A Source of Evaluation Instability
von: Filat, Nicolae, et al.
Veröffentlicht: (2026)
von: Filat, Nicolae, et al.
Veröffentlicht: (2026)
Fine-Tuning Regimes Define Distinct Continual Learning Problems
von: Iordache, Paul-Tiberiu, et al.
Veröffentlicht: (2026)
von: Iordache, Paul-Tiberiu, et al.
Veröffentlicht: (2026)
Hard Samples, Bad Labels: Robust Loss Functions That Know When to Back Off
von: Pellegrino, Nicholas, et al.
Veröffentlicht: (2025)
von: Pellegrino, Nicholas, et al.
Veröffentlicht: (2025)
Data-Driven Preference Sampling for Pareto Front Learning
von: Ye, Rongguang, et al.
Veröffentlicht: (2024)
von: Ye, Rongguang, et al.
Veröffentlicht: (2024)
Quantum Machine Learning for Predicting Anastomotic Leak: A Clinical Study
von: Novák, Vojtěch, et al.
Veröffentlicht: (2025)
von: Novák, Vojtěch, et al.
Veröffentlicht: (2025)
Nonlinear Data Integration via Kernel Methods for Data Collaboration Analysis
von: Suetake, Yamato, et al.
Veröffentlicht: (2026)
von: Suetake, Yamato, et al.
Veröffentlicht: (2026)
Adaptive Bernstein Change Detector for High-Dimensional Data Streams
von: Heyden, Marco, et al.
Veröffentlicht: (2023)
von: Heyden, Marco, et al.
Veröffentlicht: (2023)
Correction and Corruption: A Two-Rate View of Error Flow in LLM Protocols
von: Reitich, Fernando
Veröffentlicht: (2026)
von: Reitich, Fernando
Veröffentlicht: (2026)
Distinguished In Uniform: Self Attention Vs. Virtual Nodes
von: Rosenbluth, Eran, et al.
Veröffentlicht: (2024)
von: Rosenbluth, Eran, et al.
Veröffentlicht: (2024)
Adaptive Mixture Importance Sampling for Automated Ads Auction Tuning
von: Jia, Yimeng, et al.
Veröffentlicht: (2024)
von: Jia, Yimeng, et al.
Veröffentlicht: (2024)
Extracting Sentence Embeddings from Pretrained Transformer Models
von: Stankevičius, Lukas, et al.
Veröffentlicht: (2024)
von: Stankevičius, Lukas, et al.
Veröffentlicht: (2024)
Sentiment Analysis of Lithuanian Online Reviews Using Large Language Models
von: Vileikytė, Brigita, et al.
Veröffentlicht: (2024)
von: Vileikytė, Brigita, et al.
Veröffentlicht: (2024)
Bayes Conditional Distribution Estimation for Knowledge Distillation Based on Conditional Mutual Information
von: Ye, Linfeng, et al.
Veröffentlicht: (2024)
von: Ye, Linfeng, et al.
Veröffentlicht: (2024)
Harnessing non-adversarial robustness in large language models
von: Zhou, Qinghua, et al.
Veröffentlicht: (2026)
von: Zhou, Qinghua, et al.
Veröffentlicht: (2026)
A Frequency-Domain Analysis of the Multi-Armed Bandit Problem: A New Perspective on the Exploration-Exploitation Trade-off
von: Zhang, Di
Veröffentlicht: (2025)
von: Zhang, Di
Veröffentlicht: (2025)
Formulation and Therapeutic Assessment of a Zinc Oxide, Silver, and Cerium Oxide Enriched Ointment for Accelerated Wound Healing in Aged Models
von: Yousaf, Iqra, et al.
Veröffentlicht: (2025)
von: Yousaf, Iqra, et al.
Veröffentlicht: (2025)
Spatio-temporal, multi-field deep learning of shock propagation in meso-structured media
von: Fernández-Godino, M. Giselle, et al.
Veröffentlicht: (2025)
von: Fernández-Godino, M. Giselle, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Adaptive Exploration for Latent-State Bandits
von: Jin, Jikai, et al.
Veröffentlicht: (2026) -
Contrastive MIM: A Contrastive Mutual Information Framework for Unified Generative and Discriminative Representation Learning
von: Livne, Micha
Veröffentlicht: (2025) -
5C Prompt Contracts: A Minimalist, Creative-Friendly, Token-Efficient Design Framework for Individual and SME LLM Usage
von: Ari, Ugur
Veröffentlicht: (2025) -
Adaptive Weighting in Knowledge Distillation: An Axiomatic Framework for Multi-Scale Teacher Ensemble Optimization
von: Flouro, Aaron R., et al.
Veröffentlicht: (2026) -
Multi-Teacher Ensemble Distillation: A Mathematical Framework for Probability-Domain Knowledge Aggregation
von: Flouro, Aaron R., et al.
Veröffentlicht: (2026)