OpenApps: Simulating Environment Variations to Measure UI-Agent Reliability
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ullrich, Karen, Su, Jingtong, Shi, Claudia, Subramonian, Arjun, Bar, Amir, Evtimov, Ivan, Tsilivis, Nikolaos, Balestriero, Randall, Kempe, Julia, Ibrahim, Mark |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How Reinforcement Learning After Next-Token Prediction Facilitates Learning
von: Tsilivis, Nikolaos, et al.
Veröffentlicht: (2025)
von: Tsilivis, Nikolaos, et al.
Veröffentlicht: (2025)
On the Robustness of Neural Collapse and the Neural Collapse of Robustness
von: Su, Jingtong, et al.
Veröffentlicht: (2023)
von: Su, Jingtong, et al.
Veröffentlicht: (2023)
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers
von: Su, Jingtong, et al.
Veröffentlicht: (2025)
von: Su, Jingtong, et al.
Veröffentlicht: (2025)
Mission Impossible: A Statistical Perspective on Jailbreaking LLMs
von: Su, Jingtong, et al.
Veröffentlicht: (2024)
von: Su, Jingtong, et al.
Veröffentlicht: (2024)
Flavors of Margin: Implicit Bias of Steepest Descent in Homogeneous Neural Networks
von: Tsilivis, Nikolaos, et al.
Veröffentlicht: (2024)
von: Tsilivis, Nikolaos, et al.
Veröffentlicht: (2024)
The Price of Implicit Bias in Adversarially Robust Generalization
von: Tsilivis, Nikolaos, et al.
Veröffentlicht: (2024)
von: Tsilivis, Nikolaos, et al.
Veröffentlicht: (2024)
On the Geometry of Regularization in Adversarial Training: High-Dimensional Asymptotics and Generalization Bounds
von: Vilucchio, Matteo, et al.
Veröffentlicht: (2024)
von: Vilucchio, Matteo, et al.
Veröffentlicht: (2024)
Strong Model Collapse
von: Dohmatob, Elvis, et al.
Veröffentlicht: (2024)
von: Dohmatob, Elvis, et al.
Veröffentlicht: (2024)
Attacking Bayes: On the Adversarial Robustness of Bayesian Neural Networks
von: Feng, Yunzhen, et al.
Veröffentlicht: (2024)
von: Feng, Yunzhen, et al.
Veröffentlicht: (2024)
No Location Left Behind: Measuring and Improving the Fairness of Implicit Representations for Earth Data
von: Cai, Daniel, et al.
Veröffentlicht: (2025)
von: Cai, Daniel, et al.
Veröffentlicht: (2025)
auto-fpt: Automating Free Probability Theory Calculations for Machine Learning Theory
von: Subramonian, Arjun, et al.
Veröffentlicht: (2025)
von: Subramonian, Arjun, et al.
Veröffentlicht: (2025)
A Single Character can Make or Break Your LLM Evals
von: Su, Jingtong, et al.
Veröffentlicht: (2025)
von: Su, Jingtong, et al.
Veröffentlicht: (2025)
Occam's Razor for Self Supervised Learning: What is Sufficient to Learn Good Representations?
von: Ibrahim, Mark, et al.
Veröffentlicht: (2024)
von: Ibrahim, Mark, et al.
Veröffentlicht: (2024)
Theoretical and Empirical Insights into the Origins of Degree Bias in Graph Neural Networks
von: Subramonian, Arjun, et al.
Veröffentlicht: (2024)
von: Subramonian, Arjun, et al.
Veröffentlicht: (2024)
Networked Inequality: Preferential Attachment Bias in Graph Neural Network Link Prediction
von: Subramonian, Arjun, et al.
Veröffentlicht: (2023)
von: Subramonian, Arjun, et al.
Veröffentlicht: (2023)
Task Priors: Enhancing Model Evaluation by Considering the Entire Space of Downstream Tasks
von: Patel, Niket, et al.
Veröffentlicht: (2025)
von: Patel, Niket, et al.
Veröffentlicht: (2025)
Eidetic Learning: an Efficient and Provable Solution to Catastrophic Forgetting
von: Dronen, Nicholas, et al.
Veröffentlicht: (2025)
von: Dronen, Nicholas, et al.
Veröffentlicht: (2025)
SAFE: A Novel Approach to AI Weather Evaluation through Stratified Assessments of Forecasts over Earth
von: Masi, Nick, et al.
Veröffentlicht: (2025)
von: Masi, Nick, et al.
Veröffentlicht: (2025)
ALLoRA: Adaptive Learning Rate Mitigates LoRA Fatal Flaws
von: Huang, Hai, et al.
Veröffentlicht: (2024)
von: Huang, Hai, et al.
Veröffentlicht: (2024)
Weisfeiler and Leman Go Measurement Modeling: Probing the Validity of the WL Test
von: Subramonian, Arjun, et al.
Veröffentlicht: (2023)
von: Subramonian, Arjun, et al.
Veröffentlicht: (2023)
Fast and Exact Enumeration of Deep Networks Partitions Regions
von: Balestriero, Randall, et al.
Veröffentlicht: (2024)
von: Balestriero, Randall, et al.
Veröffentlicht: (2024)
LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
von: Balestriero, Randall, et al.
Veröffentlicht: (2025)
von: Balestriero, Randall, et al.
Veröffentlicht: (2025)
Learning by Reconstruction Produces Uninformative Features For Perception
von: Balestriero, Randall, et al.
Veröffentlicht: (2024)
von: Balestriero, Randall, et al.
Veröffentlicht: (2024)
Post-Hoc Guidance for Consistency Models by Joint Flow Distribution Learning
von: Hsu, Chia-Hong, et al.
Veröffentlicht: (2026)
von: Hsu, Chia-Hong, et al.
Veröffentlicht: (2026)
Uncertainty-Based Abstention in LLMs Improves Safety and Reduces Hallucinations
von: Tomani, Christian, et al.
Veröffentlicht: (2024)
von: Tomani, Christian, et al.
Veröffentlicht: (2024)
Understanding "Democratization" in NLP and ML Research
von: Subramonian, Arjun, et al.
Veröffentlicht: (2024)
von: Subramonian, Arjun, et al.
Veröffentlicht: (2024)
Stop! In the Name of Flaws: Disentangling Personal Names and Sociodemographic Attributes in NLP
von: Gautam, Vagrant, et al.
Veröffentlicht: (2024)
von: Gautam, Vagrant, et al.
Veröffentlicht: (2024)
VISReg: Variance-Invariance-Sketching Regularization for JEPA training
von: Wu, Haiyu, et al.
Veröffentlicht: (2026)
von: Wu, Haiyu, et al.
Veröffentlicht: (2026)
Curvature Tuning: Provable Training-free Model Steering From a Single Parameter
von: Hu, Leyang, et al.
Veröffentlicht: (2025)
von: Hu, Leyang, et al.
Veröffentlicht: (2025)
Characterizing Large Language Model Geometry Helps Solve Toxicity Detection and Generation
von: Balestriero, Randall, et al.
Veröffentlicht: (2023)
von: Balestriero, Randall, et al.
Veröffentlicht: (2023)
Self-Supervised Anomaly Detection in the Wild: Favor Joint Embeddings Methods
von: Otero, Daniel, et al.
Veröffentlicht: (2024)
von: Otero, Daniel, et al.
Veröffentlicht: (2024)
The Fair Language Model Paradox
von: Pinto, Andrea, et al.
Veröffentlicht: (2024)
von: Pinto, Andrea, et al.
Veröffentlicht: (2024)
Joint Embedding vs Reconstruction: Provable Benefits of Latent Space Prediction for Self Supervised Learning
von: Van Assel, Hugues, et al.
Veröffentlicht: (2025)
von: Van Assel, Hugues, et al.
Veröffentlicht: (2025)
User-Centric Design of UI for Mobile Banking Apps: Improving UI and Features for Better Customer Experience
von: Chitrakar, Luniva, et al.
Veröffentlicht: (2026)
von: Chitrakar, Luniva, et al.
Veröffentlicht: (2026)
Disentangling Geometry, Performance, and Training in Language Models
von: Kulkarni, Atharva, et al.
Veröffentlicht: (2026)
von: Kulkarni, Atharva, et al.
Veröffentlicht: (2026)
An Effective Theory of Bias Amplification
von: Subramonian, Arjun, et al.
Veröffentlicht: (2024)
von: Subramonian, Arjun, et al.
Veröffentlicht: (2024)
Learning to Think from Multiple Thinkers
von: Joshi, Nirmit, et al.
Veröffentlicht: (2026)
von: Joshi, Nirmit, et al.
Veröffentlicht: (2026)
Deep Networks Always Grok and Here is Why
von: Humayun, Ahmed Imtiaz, et al.
Veröffentlicht: (2024)
von: Humayun, Ahmed Imtiaz, et al.
Veröffentlicht: (2024)
LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures
von: Huang, Hai, et al.
Veröffentlicht: (2025)
von: Huang, Hai, et al.
Veröffentlicht: (2025)
Semantic Tube Prediction: Beating LLM Data Efficiency with JEPA
von: Huang, Hai, et al.
Veröffentlicht: (2026)
von: Huang, Hai, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
How Reinforcement Learning After Next-Token Prediction Facilitates Learning
von: Tsilivis, Nikolaos, et al.
Veröffentlicht: (2025) -
On the Robustness of Neural Collapse and the Neural Collapse of Robustness
von: Su, Jingtong, et al.
Veröffentlicht: (2023) -
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers
von: Su, Jingtong, et al.
Veröffentlicht: (2025) -
Mission Impossible: A Statistical Perspective on Jailbreaking LLMs
von: Su, Jingtong, et al.
Veröffentlicht: (2024) -
Flavors of Margin: Implicit Bias of Steepest Descent in Homogeneous Neural Networks
von: Tsilivis, Nikolaos, et al.
Veröffentlicht: (2024)