Publications
104 papers, 2001–2025. See also Google Scholar and Semantic Scholar.
2025
Hyperparameter Optimization in Machine Learning
Foundations and Trends in Machine Learning, 18(6):1054-1201, 2025.
2024
Explaining Probabilistic Models with Distributional Values
International Conference on Machine Learning (ICML), 2024.
Fortuna: A Library for Uncertainty Quantification in Deep Learning
Journal of Machine Learning Research (JMLR), Open Source Software Track, 238:1-7, 2024.
On the Choice of Learning Rate for Local SGD
Transactions on Machine Learning Research (TMLR), 2024.
Structural Pruning of Pre-trained Language Models via Neural Architecture Search
Transactions on Machine Learning Research (TMLR), 2024.
2023
A Negative Result on Gradient Matching for Selective Backprop
NeurIPS workshop on Failure Modes in the Age of Foundation Models, 2023.
Explaining Multiclass Classifiers with Categorical Values: A Case Study in Radiography
International Workshop on Trustworthy Machine Learning for Healthcare (TML4H) at ICLR, 2023.
Geographical Erasure in Language Generation
Findings of the Association for Computational Linguistics: EMNLP 2023.
Optimizing Hyperparameters with Conformal Quantile Regression
International Conference on Machine Learning (ICML), 2023.
PASHA: Efficient HPO and NAS with Progressive Resource Allocation
International Conference on Representation Learning (ICLR), 2023.
Renate: A Library for Real-world Continual Learning
Technical report, 2023.
2022
Automatic Termination for Hyperparameter Optimization
Conference on Automated Machine Learning (Main Track), 2022. (best paper award)
Continual Learning with Transformers for Image Classification
CVPR workshop on Continual Learning in Computer Vision, 2022.
Differentially private gradient boosting on linear learners for tabular data analysis
NeurIPS workshop on Trustworthy and Socially Responsible Machine Learning, 2022.
Gradient-Matching Coresets for Rehearsal-Based Continual Learning
Technical report, 2022.
Hyperparameter Optimization
In Dive Into Deep Learning, vol. 2 (Chapter 19), 2022.
Memory Efficient Continual Learning with Transformers
Annual Conference on Advances in Neural Information Processing Systems (NeurIPS), 2022.
Memory-efficient Continual Learning for Neural Text Classification
Technical report, 2022.
PASHA: Efficient HPO with Progressive Resource Allocation. [Code]
Conference on Automated Machine Learning (Late-Breaking Workshop Track), 2022.
Private Synthetic Data for Multitask Learning and Marginal Queries
Annual Conference on Advances in Neural Information Processing Systems (NeurIPS), 2022.
Syne Tune: A Library for Large Scale Hyperparameter Tuning and Reproducible Research. [GitHub]
Conference on Automated Machine Learning (Main Track), 2022.
Uncertainty Calibration in Bayesian Neural Networks via Distance-Aware Priors
Technical report, 2022.
2021
A Multi-objective Perspective on Jointly Tuning Hardware and Hyperparameters
ICLR NAS workshop, 2021.
A Resource-efficient Method for Repeated HPO and NAS Problems
ICML AutoML workshop, 2021.
Amazon SageMaker Automatic Model Tuning: Black-box Optimization at Scale
ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD), 2021. Industry track.
BORE: Bayesian Optimization by Density-Ratio Estimation
International Conference on Machine Learning (ICML), 2021. (long presentation)
Diverse Counterfactual Explanations for Anomaly Detection in Time Series
Technical report, 2021.
Dynamic Pruning of a Neural Network via Gradient Signal-to-Noise Ratio
ICML AutoML workshop, 2021.
Fair Bayesian Optimization
AAAI/ACM Conference on Artificial Intelligence, Ethics, and Society (AIES), 2021.
Gradient-matching Coresets for Continual Learning
NeurIPS workshop on Distribution Shifts: Connecting Methods and Applications, 2021.
Hyperparameter Transfer Learning with Adaptive Complexity
International Conference on Artificial Intelligence and Statistics (AISTATS), 2021.
Meta-Forecasting by Combining Global Deep Representations with Local Adaptation
Technical report, 2021.
Multi-objective Asynchronous Successive Halving
Technical report, 2021.
On the Lack of Robustness of Deep Neural Text Classifiers
Annual Meeting of the Association for Computational Linguistics (ACL), 2021. Findings.
Overfitting in Bayesian Optimization: an empirical study and early-stopping solution
ICLR NAS workshop, 2021.
Towards Robust Episodic Meta-Learning
Uncertainty in Artificial Intelligence (UAI), 2021.
2020
Bayesian Optimization by Density Ratio Estimation
NeurIPS workshop on Meta-learning, December 2020. (selected for oral presentation)
Bayesian Optimization with Fairness Constraints
ICML workshop on AutoML, 2020. (best paper award)
Cost-aware Bayesian Optimization
ICML workshop on AutoML, 2020.
LEEP: A New Measure to Evaluate Transferability of Learned Representations
International Conference on Machine Learning (ICML), 2020.
Model-based Asynchronous Hyperparameter and Neural Architecture Search
Technical report, 2020.
Multi-Objective Multi-Fidelity Hyperparameter Optimization with application to Fairness
NeurIPS workshop on Meta-learning, 2020.
Pareto-efficient Acquisition Functions for Cost-Aware Bayesian Optimization
NeurIPS workshop on Meta-learning, 2020.
2019
Constrained Bayesian Optimization with Max-Value Entropy Search
NeurIPS workshop on Meta-learning, 2019.
Learning Search Spaces for Bayesian Optimization: Another View of Hyperparameter Transfer Learning
Annual Conference on Advances in Neural Information Processing Systems (NeurIPS), 2019.
2018
A Simple Transfer Learning Extension of Hyperband
NeurIPS workshop on Meta-Learning, 2018.
Scalable Hyperparameter Transfer Learning
Annual Conference on Advances in Neural Information Processing Systems (NeurIPS), 2018.
2017
An interpretable latent variable model for attribute applicability in the Amazon catalogue
NeurIPS Symposium on Interpretable Machine Learning, 2017.
Bayesian Optimization with Tree-structured Dependencies
International Conference on Machine Learning (ICML), 2017.
Multiple Adaptive Bayesian Linear Regression for Scalable Bayesian Optimization with Warm Start
NeurIPS workshop on Meta-Learning, 2017.
2016
Adaptive Algorithms for Online Convex Optimization with Long-term Constraints
International Conference on Machine Learning (ICML), 2016.
Online Dual Decomposition for Performance and Delivery-based Distributed Ad Allocation
Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD), pp. 117-126, 2016.
Online Optimization and Regret Guarantees for Non-additive Long-term Constraints
Technical report, 2016.
2015
Incremental Variational Inference applied to Latent Dirichlet Allocation. [slides]
NeurIPS workshop on Advances in Approximate Bayesian Inference, 2015.
Incremental Variational Inference for Latent Dirichlet Allocation
Technical report, 2015.
Latent IBP compound Dirichlet Allocation
IEEE transactions in Pattern Analysis and Machine Intelligence (PAMI) 37(2):321-333, 2015.
One-Pass Ranking Models for Low-Latency Product Recommendations
Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD), pp. 1789-1798, 2015.
Online Inference for Relation Extraction with a Reduced Feature Set
Technical report, 2015.
2014
Overlapping Trace Norms in Multi-View Learning
Technical report, April 2014.
Towards Crowd-based Customer Service: A Mixed-Initiative Tool for Managing Q&A Sites
Proceedings of the ACM Conference on Human Factors in Computing Systems (CHI), pp. 2725-2734, 2014.
2013
Bringing Representativeness into Social Media Monitoring and Analysis
46th Hawaii International Conference on System Sciences (HICSS), pp. 2003-2012, 2013
Connecting Comments and Tags: Improved Modeling of Social Tagging Systems
In S. Leonardi, A. Panconesi, P. Ferragina, A. Gionis (Eds.), 6th ACM Conference on Web Search and Data Mining (WSDM), pp. 547-556, 2013.
Error Prediction with Partial Feedback
In H. Blockeel, K. Kersting, S. Nijssen, F. Zelezny (Eds.), European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML-PKDD), Lecture Notes in Computer Science (LNCS), 8189:80-94, 2013.
Log-linear Language Models based on Structured Sparsity
Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 233-243, 2013.
2012
Plackett-Luce regression: a new Bayesian model for polychotomous data
In N. de Freitas, K. P. Murphy (Eds.), Uncertainty in Artificial Intelligence (UAI) 28, pp. 84-92, 2012.
Variational Markov chain Monte Carlo for Bayesian smoothing of non-linear diffusions
Computational Statistics 27:1, 149-176, 2012.
2011
Approximate Inference for continuous-time Markov processes
In D. Barber, A. T. Cemgil, and S. Chiappa, Inference and Learning in Dynamic Models. Cambridge University Press, 2011.
Latent IBP compound Dirichlet allocation
NeurIPS 24 workshop on Bayesian nonparametrics: Hope or Hype?, 2011.
Mail2Wiki: low-cost sharing and early curation from email to wikis
In M. Foth, J. Kjeldskov, J. Paay (Eds.), Proceedings of the International Conference on Communities and Technologies (C&T) 5, pp. 98-107, 2011.
Mail2Wiki: posting and curating Wiki content from email [demo]
In P. Pu, M. J. Pazzani, E. Andre, D. Riecken (Eds.), Proceedings of the International Conference on Intelligent User Interfaces (IUI), pp 441-442, 2011.
Robust Bayesian Matrix Factorisation
Artificial Intelligence and Statistics (AISTATS) 14. JMLR Workshop and Conference Proceedings 15:425-433, 2011.
Sparse Bayesian multi-task learning
In J. Shawe-Taylor, R. S. Zemel, P. L. Bartlett, F. C. N. Pereira, K. Q. Weinberger (Eds.), Neural Information Processing Systems (NeurIPS) 24, pp. 1755-1763, 2011.
The Sequence Memoizer
Communications of the ACM, 54(2):91-98, 2011.
2010
A Comparison of Variational and Markov Chain Monte Carlo Methods for Inference in Partially Observed Stochastic Dynamic Systems
Journal of Signal Processing Systems, 61(1):51-59, 2010.
Multiple Gaussian process models [videolecture]
NeurIPS 23 workshop on New Directions in Multiple Kernel Learning, 2010. [arXiv]
2009
Prediction of hot spot residues at protein-protein interfaces by combining machine learning and energy-based methods
BMC Bioinformatics, 10: 365-382, 2009.
Sparse Probabilistic Projections
In D. Koller, D. Schuurmans, Y. Bengio, and L. Bottou (Eds.), Neural Information Processing Systems (NeurIPS) 21, pp.17-24, 2009. The MIT Press.
Stochastic Memoizer for Sequence Data
In L. Bottou and M. Littman, Proceedings of the 26th International Conference on Machine Learning (ICML), Montreal (Quebec), Canada, June 14-18, 2009, pp. 1129-1136. ACM.
Switching Regulatory Models of Cellular Stress Response
Bioinformatics, 25(10): 1280-1286, 2009. Oxford University Press.
The Variational Gaussian Approximation Revisited
Neural Computation 21(3):786-792, 2009.
2008
Improving the robustness to outliers of mixtures of probabilistic PCAs
In T. Wahio, et al. (Eds.), Pacific-Asia Conference on Knowledge Discovery and Data Mining (PAKDD) 12, Lecture notes in Artificial Intelligence (LNAI) 5012:527-535, 2008. Springer.
Mixtures of Robust Probabilistic Principal Component Analyzers
Neurocomputing, 71(7-9):1274-1282, 2008. Elsevier.
Using Subspace-Based Template Attacks to Compare and Combine Power and Electromagnetic Information Leakages
In E. Oswald and P. Rohatgi (Eds.), 10th International Workshop on Cryptographic Hardware and Embedded Systems (CHES), Washington, DC, USA, 10-13 August, 2008. Lecture Notes in Computer Science vol. 5154, pp. 411-425. Springer.
Variational Inference for Diffusion Processes
In C. Platt, D. Koller, Y. Singer and S. Roweis (Eds.), Neural Information Processing Systems (NeurIPS) 20, pp.17-24, 2008. The MIT Press.
2007
Evaluation of Variational and Markov Chain Monte Carlo Methods for Inference in Partially Observed Stochastic Dynamic Systems
Proceedings of the 17th IEEE workshop on Machine Learning for Signal Processing (MLSP), Thessaloniki, Greece, 27-28 August, 2007, pp. 306-311.
Gaussian Process Approximations of Stochastic Differential Equations
Journal of Machine Learning Research Workshop and Conference Proceedings, 1:1-16, 2007.
Mixtures of Robust Probabilistic Principal Component Analyzers
Proceedings of the 15th European Symposium on Artificial Neural Networks (ESANN), Bruges, Belgium, April 25-27, 2007, pp. 229-234. D-side.
Robust Bayesian Clustering
Neural Networks, 20:129-138, 2007. Elsevier.
2006
Automatic Adjustment of Discriminant Adaptive Nearest Neighbor
In Y.Y. Tang, P. Wang, G. Lorette and D.S. Yeung (Eds.), Proceedings of the 18th International Conference on Pattern Recognition (ICPR), Hong Kong, P.R.C., 20-24 August, 2006, vol. 2, pp. 525-555. IEEE Computer Society.
Robust Probabilistic Projections
In W. W. Cohen and A. Moore (Eds.), Proceedings of the 23rd International Conference on Machine Learning (ICML), Pittsburgh (PA), U.S.A., 25-29 June, 2006, pp. 33-40. ACM.
Template Attacks in Principal Subspaces
In L. Goubin and M. Matsui (Eds.), 8th International Workshop on Cryptographic Hardware and Embedded Systems (CHES), Yokohama, Japan, 10-13 October, 2006. Lecture Notes in Computer Science vol. 4249, pp. 1-14. Springer.
Towards Security Limits of Side-Channel Attacks
In L. Goubin and M. Matsui (Eds.), 8th International Workshop on Cryptographic Hardware and Embedded Systems (CHES), Yokohama, Japan, 10-13 October, 2006. Lecture Notes in Computer Science vol. 4249, pp. 30-45. Springer.
2005
Local Vector-based Models for Sense Discrimination
In H. Bunt, J. Geertzen and E. Thijsse (Eds.), Proceedings of the 6th International Workshop on Computational Semantics (IWCS), Tilburg, the Netherlands, January 12-14, 2005, pp. 163-174.
Manifold Constrained Finite Gaussian Mixtures
In J. Cabestany, A. Prieto and F. Sandoval Hernández (Eds.), Computational Intelligence and Bioinspired Systems - 8th International Work-Conference on Artificial Neural Networks (IWANN), Vilanova i la Geltrú (Barcelona), Spain, June 8-10, 2005. Lecture Notes in Computer Science, vol. 3512, pp.820-828. Springer.
2004
Entropy Minima and Distribution Structural Modifications in Blind Separation of Multi-model Sources
In R. Fisher, R. Preuss and U. von Toussaint, Proceedings of the 24th International Workshop on Bayesian Inference and Maximum Entropy Methods in Science and Engineering (MaxEnt), IPP Garching bei München, Germany, July 25-30, 2004, pp. 589-596. American Institute of Physics (AIP).
Flexible and Robust Bayesian Classification by Finite Mixture Models
Proceedings of the 12th European Symposium on Artificial Neural Networks (ESANN), Bruges, Belgium, April 28-30, 2004, pp. 75-80. D-side.
Prediction of Visual Perceptions with Artificial Neural Networks in a Visual Prosthesis for the Blind
Artificial Intelligence in Medicine, 32(3):183-194, 2004. Elsevier.
Supervised Nonparametric Information Theoretic Classification
In J. Kittler, M. Petrou and M. Nixon (Eds.), Proceedings of the 17th International Conference on Pattern Recognition (ICPR), Cambridge, U.K., August 23-26, 2004, vol. 3, pp. 414-417. IEEE Computer Society.
Towards a Local Separation Performances Estimator using Common ICA contrast Funtions?
Proceedings of the 12th European Symposium on Artificial Neural Networks (ESANN), Bruges, Belgium, April 28-30, 2004, pp. 211-216. D-side.
2003
Classification of Visual Sensations Generated Electrically in the Visual Field of the Blind
In D. D. Feng and E. R. Carson (Eds.), Proceedings of the 5th IFAC Symposium on Modelling and Control in Biomedical Systems, Melbourne, Australia, August 21-23, 2003, pp. 223-228. Elsevier.
Locally Linear Embedding versus Isotop
Proceedings of the 11th European Symposium on Artificial Neural Networks (ESANN), Bruges, Belgium, April 23-25, 2003, pp. 527-534.
On Convergence Problems of the EM Algorithm for Finite Gaussian Mixtures
Proceedings of the 11th European Symposium on Artificial Neural Networks (ESANN), Bruges, Belgium, April 23-25, 2003, pp. 99-106. D-side.
2002
Width Optimization of the Gaussian Kernels in Radial Basis Function Networks
Proceedings of the 10th European Symposium on Artificial Neural Networks (ESANN), Bruges, Belgium, April 24-26, 2002, pp. 425-432.
2001
Phosphene Evaluation in a Visual Prosthesis with Artificial Neural Networks
Proceedings of the 1st European Symposium on Intelligent Technologies, Hybrid Systems and their implementation on Smart Adaptive Systems (EUNITE), Puerto de la Cruz (Tenerife), Spain, December 13-14, 2001, pp. 509-515.