Bibliografia#
Le opere citate nel libro, in ordine alfabetico. Nel testo le citazioni compaiono come rimandi tra parentesi quadre: ogni rimando porta qui.
Kjersti Aas, Martin Jullum, and Anders Løland. Explaining individual predictions when features are dependent: more accurate approximations to Shapley values. Artificial Intelligence, 298:103502, 2021. URL: https://doi.org/10.1016/j.artint.2021.103502.
Martin Abadi, Andy Chu, Ian Goodfellow, H. Brendan McMahan, Ilya Mironov, Kunal Talwar, and Li Zhang. Deep learning with differential privacy. In Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security (CCS), 308–318. 2016. URL: https://arxiv.org/abs/1607.00133.
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek G. Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng. TensorFlow: a system for large-scale machine learning. In 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16), 265–283. 2016. URL: https://www.usenix.org/conference/osdi16/technical-sessions/presentation/abadi.
Pieter Abbeel and Andrew Y. Ng. Apprenticeship learning via inverse reinforcement learning. In Proceedings of the 21st International Conference on Machine Learning (ICML), 1. ACM, 2004. doi:10.1145/1015330.1015430.
Samira Abnar and Willem Zuidema. Quantifying attention flow in transformers. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 4190–4197. 2020.
Bruce Abramson. Expected-outcome: a general model of static evaluation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 12(2):182–193, 1990.
Margareta Ackerman and Shai Ben-David. Measures of clustering quality: a working set of axioms for clustering. In Advances in Neural Information Processing Systems 21, 121–128. 2008.
David H. Ackley, Geoffrey E. Hinton, and Terrence J. Sejnowski. A learning algorithm for Boltzmann machines. Cognitive Science, 9(1):147–169, 1985. URL: https://onlinelibrary.wiley.com/doi/abs/10.1207/s15516709cog0901_7, doi:10.1207/s15516709cog0901_7.
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim. Sanity checks for saliency maps. In Advances in Neural Information Processing Systems, volume 31. 2018.
Alekh Agarwal, Alina Beygelzimer, Miroslav Dudík, John Langford, and Hanna Wallach. A reductions approach to fair classification. In International Conference on Machine Learning (ICML). 2018. URL: https://arxiv.org/abs/1803.02453.
Megha Agarwal, Asfandyar Qureshi, Nikhil Sardana, Linden Li, Julian Quevedo, and Daya Khudia. LLM inference performance engineering: best practices. Databricks Blog, 2023. URL: https://www.databricks.com/blog/llm-inference-performance-engineering-best-practices.
Rishabh Agarwal, Dale Schuurmans, and Mohammad Norouzi. An optimistic perspective on offline reinforcement learning. In Proceedings of the 37th International Conference on Machine Learning (ICML), volume 119 of Proceedings of Machine Learning Research, 104–114. PMLR, 2020. URL: https://arxiv.org/abs/1907.04543.
Rishabh Agarwal, Avi Singh, Lei M. Zhang, Bernd Bohnet, Luis Rosias, Stephanie Chan, Biao Zhang, Ankesh Anand, Zaheer Abbas, Azade Nova, John D. Co-Reyes, Eric Chu, Feryal Behbahani, Aleksandra Faust, and Hugo Larochelle. Many-shot in-context learning. In Advances in Neural Information Processing Systems (NeurIPS), volume 37. 2024. Spotlight. URL: https://arxiv.org/abs/2404.11018.
Andrea Agostinelli, Timo I. Denk, Zalán Borsos, Jesse Engel, Mauro Verzetti, Antoine Caillon, Qingqing Huang, Aren Jansen, Adam Roberts, Marco Tagliasacchi, Matt Sharifi, Neil Zeghidour, and Christian Frank. Musiclm: generating music from text. arXiv preprint arXiv:2301.11325, 2023.
Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S. Gulavani, Alexey Tumanov, and Ramachandran Ramjee. Taming throughput-latency tradeoff in LLM inference with Sarathi-Serve. In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI), 117–134. 2024.
Shipra Agrawal and Navin Goyal. Analysis of Thompson sampling for the multi-armed bandit problem. In Proceedings of the 25th Annual Conference on Learning Theory (COLT), 39.1–39.26. 2012.
Orevaoghene Ahia, Sachin Kumar, Hila Gonen, Jungo Kasai, David R. Mortensen, Noah A. Smith, and Yulia Tsvetkov. Do all languages cost the same? tokenization in the era of commercial language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), 9904–9923. 2023. doi:10.18653/v1/2023.emnlp-main.614.
Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, and Sumit Sanghai. GQA: training generalized multi-query transformer models from multi-head checkpoints. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), 4895–4901. 2023. URL: https://arxiv.org/abs/2305.13245.
Mark A. Aizerman, Emmanuel M. Braverman, and Lev I. Rozonoer. Theoretical foundations of the potential function method in pattern recognition learning. Automation and Remote Control, 25:821–837, 1964.
Hirotugu Akaike. A new look at the statistical model identification. IEEE Transactions on Automatic Control, 19(6):716–723, 1974.
Guillaume Alain and Yoshua Bengio. What regularized auto-encoders learn from the data-generating distribution. Journal of Machine Learning Research, 15:3743–3773, 2014.
Guillaume Alain and Yoshua Bengio. Understanding intermediate layers using linear classifier probes. arXiv preprint arXiv:1610.01644, 2016. ICLR 2017, workshop track. URL: https://arxiv.org/abs/1610.01644.
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikołaj Bińkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karen Simonyan. Flamingo: a visual language model for few-shot learning. In Advances in Neural Information Processing Systems (NeurIPS). 2022.
Michael S. Albergo and Eric Vanden-Eijnden. Building normalizing flows with stochastic interpolants. In International Conference on Learning Representations (ICLR). 2023.
Reem Aleithan, Haoran Xue, Mohammad Mahdi Mohajer, Elijah Nnorom, Gias Uddin, and Song Wang. SWE-Bench+: enhanced coding benchmark for LLMs. arXiv preprint arXiv:2410.06992, 2024. URL: https://arxiv.org/abs/2410.06992.
Daniel Aloise, Amit Deshpande, Pierre Hansen, and Preyas Popat. NP-hardness of Euclidean sum-of-squares clustering. Machine Learning, 75(2):245–248, 2009.
Uri Alon and Eran Yahav. On the bottleneck of graph neural networks and its practical implications. In International Conference on Learning Representations. 2021.
David Alvarez-Melis and Tommi S. Jaakkola. On the robustness of interpretability methods. In ICML Workshop on Human Interpretability in Machine Learning (WHI 2018). 2018. URL: https://arxiv.org/abs/1806.08049.
Shun-Ichi Amari. Learning patterns and pattern sequences by self-organizing nets of threshold elements. IEEE Transactions on Computers, C-21(11):1197–1206, 1972. doi:10.1109/T-C.1972.223477.
Xavier Amatriain and Justin Basilico. Netflix recommendations: beyond the 5 stars (part 1). Netflix Technology Blog, April 2012. URL: https://netflixtechblog.com/netflix-recommendations-beyond-the-5-stars-part-1-55838468f429.
Gene M. Amdahl. Validity of the single processor approach to achieving large scale computing capabilities. In Proceedings of the April 18-20, 1967, Spring Joint Computer Conference (AFIPS '67 Spring), 483–485. 1967. doi:10.1145/1465482.1465560.
Saleema Amershi, Andrew Begel, Christian Bird, Robert DeLine, Harald Gall, Ece Kamar, Nachiappan Nagappan, Besmira Nushi, and Thomas Zimmermann. Software engineering for machine learning: a case study. In IEEE/ACM 41st International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP), 291–300. 2019.
Daniel J. Amit, Hanoch Gutfreund, and Haim Sompolinsky. Storing infinite numbers of patterns in a spin-glass model of neural networks. Physical Review Letters, 55(14):1530–1533, 1985. URL: https://link.aps.org/doi/10.1103/PhysRevLett.55.1530, doi:10.1103/PhysRevLett.55.1530.
Yali Amit and Donald Geman. Shape quantization and recognition with randomized trees. Neural Computation, 9(7):1545–1588, 1997.
Marco Ancona, Enea Ceolini, Cengiz Öztireli, and Markus Gross. Towards better understanding of gradient-based attribution methods for deep neural networks. In International Conference on Learning Representations (ICLR). 2018. arXiv:1711.06104.
Brian D. O. Anderson. Reverse-time diffusion equation models. Stochastic Processes and their Applications, 12(3):313–326, 1982.
James A. Anderson. A simple neural network generating an interactive memory. Mathematical Biosciences, 14(3–4):197–220, 1972. doi:10.1016/0025-5564(72)90075-2.
Marcin Andrychowicz, Anton Raichuk, Piotr Stańczyk, Manu Orsini, Sertan Girgin, Raphael Marinier, Léonard Hussenot, Matthieu Geist, Olivier Pietquin, Marcin Michalski, Sylvain Gelly, and Olivier Bachem. What matters in on-policy reinforcement learning? a large-scale empirical study. In International Conference on Learning Representations (ICLR). 2021. URL: https://arxiv.org/abs/2006.05990, arXiv:2006.05990.
Vito Walter Anelli, Daniele Malitesta, Claudio Pomo, Alejandro Bellogín, Eugenio Di Sciascio, and Tommaso Di Noia. Challenging the myth of graph collaborative filtering: a reasoned and reproducibility-driven analysis. In Proceedings of the 17th ACM Conference on Recommender Systems (RecSys '23), 350–361. ACM, 2023. doi:10.1145/3604915.3609489.
Julia Angwin, Jeff Larson, Surya Mattu, and Lauren Kirchner. Machine bias. ProPublica, 2016. URL: https://www.propublica.org/article/machine-bias-risk-assessments-in-criminal-sentencing.
Cem Anil, Esin Durmus, Nina Panickssery, Mrinank Sharma, Joe Benton, Sandipan Kundu, and others. Many-shot jailbreaking. In Advances in Neural Information Processing Systems, volume 37. 2024. URL: https://openreview.net/forum?id=cw5mgd71jW.
Abdul Fatir Ansari, Lorenzo Stella, Caner Turkmen, Xiyuan Zhang, Pedro Mercado, Huibin Shen, Oleksandr Shchur, Syama Sundar Rangapuram, Sebastian Pineda Arango, Shubham Kapoor, and others. Chronos: learning the language of time series. Transactions on Machine Learning Research (TMLR), 2024. URL: https://arxiv.org/abs/2403.07815.
Francis J. Anscombe. Graphs in statistical analysis. The American Statistician, 27(1):17–21, 1973.
Daniel W. Apley and Jingyu Zhu. Visualizing the effects of predictor variables in black box supervised learning models. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 82(4):1059–1086, 2020. URL: https://arxiv.org/abs/1612.08468.
Martin Arjovsky and Léon Bottou. Towards principled methods for training generative adversarial networks. In International Conference on Learning Representations (ICLR). 2017.
Martin Arjovsky, Soumith Chintala, and Léon Bottou. Wasserstein generative adversarial networks. In Proceedings of the 34th International Conference on Machine Learning (ICML). 2017.
Sanjeev Arora, Yingyu Liang, and Tengyu Ma. A simple but tough-to-beat baseline for sentence embeddings. In International Conference on Learning Representations (ICLR). 2017.
Simran Arora, Sabri Eyuboglu, Aman Timalsina, Isys Johnson, Michael Poli, James Zou, Atri Rudra, and Christopher Ré. Zoology: measuring and improving recall in efficient language models. arXiv preprint arXiv:2312.04927, 2023. URL: https://arxiv.org/abs/2312.04927, arXiv:2312.04927.
Simran Arora, Sabri Eyuboglu, Michael Zhang, Aman Timalsina, Silas Alberti, Dylan Zinsley, James Zou, Atri Rudra, and Christopher Ré. Simple linear attention language models balance the recall-throughput tradeoff. In Proceedings of the 41st International Conference on Machine Learning (ICML). 2024.
Marianne Arriola, Subham Sekhar Sahoo, Aaron Gokaslan, Zhihan Yang, Zhixuan Qi, Jiaqi Han, Justin T. Chiu, and Volodymyr Kuleshov. Block diffusion: interpolating between autoregressive and diffusion language models. In International Conference on Learning Representations (ICLR). 2025. URL: https://arxiv.org/abs/2503.09573.
Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. On the cross-lingual transferability of monolingual representations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 4623–4637. 2020.
David Arthur and Sergei Vassilvitskii. K-means++: the advantages of careful seeding. In Proceedings of the Eighteenth Annual ACM-SIAM Symposium on Discrete Algorithms (SODA), 1027–1035. 2007.
Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi. Self-RAG: learning to retrieve, generate, and critique through self-reflection. In International Conference on Learning Representations (ICLR). 2024.
Mahmoud Assran, Quentin Duval, Ishan Misra, Piotr Bojanowski, Pascal Vincent, Michael Rabbat, Yann LeCun, and Nicolas Ballas. Self-supervised learning from images with a joint-embedding predictive architecture. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2023. URL: https://arxiv.org/abs/2301.08243.
Mido Assran, Adrien Bardes, David Fan, Quentin Garrido, Russell Howes, Mojtaba Komeili, Matthew Muckley, Ammar Rizvi, Claire Roberts, Koustuv Sinha, Artem Zholus, Sergio Arnaud, Abha Gejji, Ada Martin, Francois Robert Hogan, Daniel Dugas, Piotr Bojanowski, Vasil Khalidov, Patrick Labatut, Francisco Massa, Marc Szafraniec, Kapil Krishnakumar, Yong Li, Xiaodong Ma, Sarath Chandar, Franziska Meier, Yann LeCun, Michael Rabbat, and Nicolas Ballas. V-jepa 2: self-supervised video models enable understanding, prediction and planning. arXiv preprint arXiv:2506.09985, 2025. URL: https://arxiv.org/abs/2506.09985.
Anish Athalye, Nicholas Carlini, and David Wagner. Obfuscated gradients give a false sense of security: circumventing defenses to adversarial examples. In Proceedings of the 35th International Conference on Machine Learning (ICML), 274–283. 2018.
Peter Auer, Nicolò Cesa-Bianchi, and Paul Fischer. Finite-time analysis of the multiarmed bandit problem. Machine Learning, 47(2):235–256, 2002.
Peter Auer, Nicolò Cesa-Bianchi, Yoav Freund, and Robert E. Schapire. The nonstochastic multiarmed bandit problem. SIAM Journal on Computing, 32(1):48–77, 2002. doi:10.1137/S0097539701398375.
J. L. Austin. How to Do Things with Words. Clarendon Press, Oxford, 1962.
Jacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow, and Rianne van den Berg. Structured denoising diffusion models in discrete state-spaces. In Advances in Neural Information Processing Systems (NeurIPS). 2021. URL: https://arxiv.org/abs/2107.03006.
Mohammad Gheshlaghi Azar, Mark Rowland, Bilal Piot, Daniel Guo, Daniele Calandriello, Michal Valko, and Rémi Munos. A general theoretical paradigm to understand learning from human preferences. arXiv preprint arXiv:2310.12036, 2023.
Jimmy Ba and Rich Caruana. Do deep nets really need to be deep? In Advances in Neural Information Processing Systems, volume 27. 2014.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016.
Kent Bach and Robert M. Harnish. Linguistic Communication and Speech Acts. MIT Press, Cambridge, MA, 1979.
Sebastian Bach, Alexander Binder, Grégoire Montavon, Frederick Klauschen, Klaus-Robert Müller, and Wojciech Samek. On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation. PLOS ONE, 10(7):e0130140, 2015. doi:10.1371/journal.pone.0130140.
Yoram Bachrach, Yehuda Finkelstein, Ran Gilad-Bachrach, Liran Katzir, Noam Koenigstein, Nir Nice, and Ulrich Paquet. Speeding up the Xbox recommender system using a Euclidean transformation for inner-product spaces. In Proceedings of the 8th ACM Conference on Recommender Systems (RecSys '14), 257–264. ACM, 2014. doi:10.1145/2645710.2645741.
Arturs Backurs and Piotr Indyk. Edit distance cannot be computed in strongly subquadratic time (unless SETH is false). In Proceedings of the Forty-Seventh Annual ACM Symposium on Theory of Computing (STOC), 51–58. 2015. doi:10.1145/2746539.2746612.
Pierre-Luc Bacon, Jean Harb, and Doina Precup. The option-critic architecture. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 31. 2017. URL: https://arxiv.org/abs/1609.05140.
Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu, and Michael Auli. Data2vec: a general framework for self-supervised learning in speech, vision and language. In International Conference on Machine Learning (ICML). 2022. URL: https://arxiv.org/abs/2202.03555.
Alexei Baevski and Abdelrahman Mohamed. Effectiveness of self-supervised pre-training for ASR. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 7694–7698. 2020. doi:10.1109/ICASSP40776.2020.9054224.
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. Wav2vec 2.0: a framework for self-supervised learning of speech representations. In Advances in Neural Information Processing Systems (NeurIPS), volume 33, 12449–12460. 2020.
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. Neural machine translation by jointly learning to align and translate. In International Conference on Learning Representations. 2015.
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-VL: a versatile vision-language model for understanding, localization, text reading, and beyond. arXiv preprint arXiv:2308.12966, 2023.
Shaojie Bai, J. Zico Kolter, and Vladlen Koltun. An empirical evaluation of generic convolutional and recurrent networks for sequence modeling. arXiv preprint arXiv:1803.01271, 2018. URL: https://arxiv.org/abs/1803.01271.
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, and others. Constitutional ai: harmlessness from ai feedback. arXiv preprint arXiv:2212.08073, 2022. URL: https://arxiv.org/abs/2212.08073.
Sergey Bakin. Adaptive Regression and Model Selection in Data Mining Problems. PhD thesis, Australian National University, Canberra, 1999.
Pierre Baldi and Kurt Hornik. Neural networks and principal component analysis: learning from examples without local minima. Neural Networks, 2(1):53–58, 1989. doi:10.1016/0893-6080(89)90014-2.
David Balduzzi, Marta Garnelo, Yoram Bachrach, Wojciech Czarnecki, Julien Perolat, Max Jaderberg, and Thore Graepel. Open-ended learning in symmetric zero-sum games. In Proceedings of the 36th International Conference on Machine Learning (ICML), volume 97 of Proceedings of Machine Learning Research, 434–443. 2019. URL: https://proceedings.mlr.press/v97/balduzzi19a.html.
David Balduzzi, Karl Tuyls, Julien Perolat, and Thore Graepel. Re-evaluating evaluation. In Advances in Neural Information Processing Systems 31 (NeurIPS). 2018. URL: https://arxiv.org/abs/1806.02643.
M. Ballerini, N. Cabibbo, R. Candelier, A. Cavagna, E. Cisbani, I. Giardina, V. Lecomte, A. Orlandi, G. Parisi, A. Procaccini, M. Viale, and V. Zdravkovic. Interaction ruling animal collective behavior depends on topological rather than metric distance: evidence from a field study. Proceedings of the National Academy of Sciences, 105(4):1232–1237, 2008.
Ron Banner, Yury Nahshan, and Daniel Soudry. Post training 4-bit quantization of convolutional networks for rapid-deployment. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019.
Fan Bao, Shen Nie, Kaiwen Xue, Yue Cao, Chongxuan Li, Hang Su, and Jun Zhu. All are worth words: a ViT backbone for diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 22669–22679. 2023. URL: https://arxiv.org/abs/2209.12152.
Adrien Bardes, Quentin Garrido, Jean Ponce, Xinlei Chen, Michael Rabbat, Yann LeCun, Mahmoud Assran, and Nicolas Ballas. Revisiting feature prediction for learning visual representations from video. Transactions on Machine Learning Research, 2024. URL: https://arxiv.org/abs/2404.08471.
Adrien Bardes, Jean Ponce, and Yann LeCun. VICReg: variance-invariance-covariance regularization for self-supervised learning. In International Conference on Learning Representations (ICLR). 2022. URL: https://arxiv.org/abs/2105.04906.
David Barina. Convergence verification of the Collatz problem. The Journal of Supercomputing, 77(3):2681–2688, 2021. doi:10.1007/s11227-020-03368-x.
Horace B. Barlow. Possible principles underlying the transformations of sensory messages. In Walter A. Rosenblith, editor, Sensory Communication. MIT Press, 1961. doi:10.7551/mitpress/9780262518420.003.0013.
Beth Barnes. Debate update: obfuscated arguments problem. AI Alignment Forum, 2020. URL: https://www.alignmentforum.org/posts/PJLABqQ962hZEqhdB/debate-update-obfuscated-arguments-problem.
Solon Barocas, Moritz Hardt, and Arvind Narayanan. Fairness and Machine Learning: Limitations and Opportunities. MIT Press, 2023. Versione online gratuita su fairmlbook.org. URL: https://fairmlbook.org/.
Andrew R. Barron. Universal approximation bounds for superpositions of a sigmoidal function. IEEE Transactions on Information Theory, 39(3):930–945, 1993.
Peter L. Bartlett. The sample complexity of pattern classification with neural networks: the size of the weights is more important than the size of the network. IEEE Transactions on Information Theory, 44(2):525–536, 1998. doi:10.1109/18.661502.
Peter L. Bartlett, Dylan J. Foster, and Matus Telgarsky. Spectrally-normalized margin bounds for neural networks. In Advances in Neural Information Processing Systems 30 (NeurIPS), 6241–6250. 2017. URL: https://arxiv.org/abs/1706.08498.
Peter L. Bartlett, Nick Harvey, Christopher Liaw, and Abbas Mehrabian. Nearly-tight VC-dimension and pseudodimension bounds for piecewise linear neural networks. Journal of Machine Learning Research, 20(63):1–17, 2019.
Peter L. Bartlett and Shahar Mendelson. Rademacher and Gaussian complexities: risk bounds and structural results. Journal of Machine Learning Research, 3:463–482, 2002.
Ali Basiri, Niosha Behnam, Ruud de Rooij, Lorin Hochstein, Luke Kosewski, Justin Reynolds, and Casey Rosenthal. Chaos engineering. IEEE Software, 33(3):35–41, 2016. doi:10.1109/MS.2016.60.
Peter W. Battaglia, Jessica B. Hamrick, Victor Bapst, and others. Relational inductive biases, deep learning, and graph networks. arXiv preprint arXiv:1806.01261, 2018.
David Bau, Bolei Zhou, Aditya Khosla, Aude Oliva, and Antonio Torralba. Network dissection: quantifying interpretability of deep visual representations. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 6541–6549. 2017.
David Bau, Jun-Yan Zhu, Hendrik Strobelt, Agata Lapedriza, Bolei Zhou, and Antonio Torralba. Understanding the role of individual units in a deep neural network. Proceedings of the National Academy of Sciences, 117(48):30071–30078, 2020. doi:10.1073/pnas.1907375117.
Mohammad Bavarian, Heewoo Jun, Nikolas Tezak, John Schulman, Christine McLeavey, Jerry Tworek, and Mark Chen. Efficient training of language models to fill in the middle. arXiv preprint arXiv:2207.14255, 2022.
Rohan Bavishi, Erich Elsen, Curtis Hawthorne, Maxwell Nye, Augustus Odena, Arushi Somani, and Sağnak Taşırlar. Fuyu-8B: a multimodal architecture for AI agents. Adept AI, blog e scheda del modello, 2023. URL: https://www.adept.ai/blog/fuyu-8b.
Dave Bayer and Persi Diaconis. Trailing the dovetail shuffle to its lair. The Annals of Applied Probability, 2(2):294–313, 1992.
Christian Beck, Sebastian Becker, Philipp Grohs, Nor Jaafari, and Arnulf Jentzen. Solving the Kolmogorov PDE by means of deep learning. Journal of Scientific Computing, 88:73, 2021. arXiv:1806.00421.
Maximilian Beck, Korbinian Pöppel, Phillip Lippe, Richard Kurle, Patrick Blies, Günter Klambauer, Sebastian Böck, and Sepp Hochreiter. xLSTM 7B: a recurrent LLM for fast and efficient inference. In Proceedings of the 42nd International Conference on Machine Learning (ICML). 2025.
Maximilian Beck, Korbinian Pöppel, Markus Spanring, Andreas Auer, Oleksandra Prudnikova, Michael Kopp, Günter Klambauer, Johannes Brandstetter, and Sepp Hochreiter. xLSTM: extended long short-term memory. In Advances in Neural Information Processing Systems (NeurIPS). 2024.
Ahmad Beirami, Alekh Agarwal, Jonathan Berant, Alexander D'Amour, Jacob Eisenstein, Chirag Nagpal, and Ananda Theertha Suresh. Theoretical guarantees on the best-of-n alignment policy. In Proceedings of the 42nd International Conference on Machine Learning (ICML). 2025. arXiv:2401.01879.
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal. Reconciling modern machine-learning practice and the classical bias–variance trade-off. Proceedings of the National Academy of Sciences, 116(32):15849–15854, 2019. doi:10.1073/pnas.1903070116.
Anthony J. Bell and Terrence J. Sejnowski. An information-maximization approach to blind separation and blind deconvolution. Neural Computation, 7(6):1129–1159, 1995.
Marc G. Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Rémi Munos. Unifying count-based exploration and intrinsic motivation. In Advances in Neural Information Processing Systems 29 (NeurIPS), 1471–1479. 2016. URL: https://arxiv.org/abs/1606.01868, arXiv:1606.01868.
Richard Bellman. Dynamic Programming. Princeton University Press, 1957.
Adel Belouchrani, Karim Abed-Meraim, Jean-François Cardoso, and Éric Moulines. A blind source separation technique using second-order statistics. IEEE Transactions on Signal Processing, 45(2):434–444, 1997.
Iz Beltagy, Matthew E. Peters, and Arman Cohan. Longformer: the long-document transformer. arXiv preprint arXiv:2004.05150, 2020.
Shai Ben-David, Ulrike von Luxburg, and Dávid Pál. A sober look at clustering stability. In Learning Theory (COLT 2006), volume 4005 of Lecture Notes in Computer Science, 5–19. Springer, 2006.
Michael Ben-Or. Another advantage of free choice (extended abstract): completely asynchronous agreement protocols. In Proceedings of the Second Annual ACM Symposium on Principles of Distributed Computing (PODC '83), 27–30. ACM, 1983. doi:10.1145/800221.806707.
Paul Benacerraf. God, the devil, and Gödel. The Monist, 51(1):9–32, 1967.
Emily M. Bender. Stochastic parrots: frequently unasked questions. Medium, May 2026. URL: https://medium.com/@emilymenonbender/stochastic-parrots-frequently-unasked-questions-49c2e7d22d11.
Emily M. Bender and Batya Friedman. Data statements for natural language processing: toward mitigating system bias and enabling better science. Transactions of the Association for Computational Linguistics, 6:587–604, 2018. doi:10.1162/tacl_a_00041.
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. On the dangers of stochastic parrots: can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (FAccT). 2021.
Emily M. Bender and Alexander Koller. Climbing towards NLU: On meaning, form, and understanding in the age of data. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), 5185–5198. 2020. doi:10.18653/v1/2020.acl-main.463.
Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer. Scheduled sampling for sequence prediction with recurrent neural networks. In Advances in Neural Information Processing Systems, volume 28. 2015.
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Jauvin. A neural probabilistic language model. Journal of Machine Learning Research, 3:1137–1155, 2003.
Yoshua Bengio, Geoffrey Hinton, Andrew Yao, Dawn Song, Pieter Abbeel, Trevor Darrell, and others. Managing extreme AI risks amid rapid progress. Science, 384(6698):842–845, 2024. doi:10.1126/science.adn0117.
Yoshua Bengio, Pascal Lamblin, Dan Popovici, and Hugo Larochelle. Greedy layer-wise training of deep networks. In Advances in Neural Information Processing Systems 19 (NIPS 2006). MIT Press, 2007.
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston. Curriculum learning. In Proceedings of the 26th Annual International Conference on Machine Learning (ICML), 41–48. 2009. doi:10.1145/1553374.1553380.
Yoshua Bengio, Nicholas Léonard, and Aaron Courville. Estimating or propagating gradients through stochastic neurons for conditional computation. arXiv preprint arXiv:1308.3432, 2013. URL: https://arxiv.org/abs/1308.3432.
Yoshua Bengio, Patrice Simard, and Paolo Frasconi. Learning long-term dependencies with gradient descent is difficult. IEEE Transactions on Neural Networks, 5(2):157–166, 1994.
Yoav Benjamini and Yosef Hochberg. Controlling the false discovery rate: a practical and powerful approach to multiple testing. Journal of the Royal Statistical Society, Series B, 57(1):289–300, 1995.
Yoav Benjamini and Daniel Yekutieli. The control of the false discovery rate in multiple testing under dependency. The Annals of Statistics, 29(4):1165–1188, 2001.
Adam L. Berger, Stephen A. Della Pietra, and Vincent J. Della Pietra. A maximum entropy approach to natural language processing. Computational Linguistics, 22(1):39–71, 1996.
Christoph Bergmeir, Rob J. Hyndman, and Bonsoo Koo. A note on the validity of cross-validation for evaluating autoregressive time series prediction. Computational Statistics & Data Analysis, 120:70–83, 2018. doi:10.1016/j.csda.2017.11.003.
James Bergstra, Rémi Bardenet, Yoshua Bengio, and Balázs Kégl. Algorithms for hyper-parameter optimization. In Advances in Neural Information Processing Systems (NeurIPS), volume 24. 2011.
James Bergstra and Yoshua Bengio. Random search for hyper-parameter optimization. Journal of Machine Learning Research, 13:281–305, 2012.
Elwyn R. Berlekamp, John H. Conway, and Richard K. Guy. Winning Ways for Your Mathematical Plays. Academic Press, London, 1982. 2 voll.
Dimitri P. Bertsekas and John N. Tsitsiklis. Neuro-Dynamic Programming. Athena Scientific, Belmont, Massachusetts, 1996.
Dimitris Bertsimas, Angela King, and Rahul Mazumder. Best subset selection via a modern optimization lens. The Annals of Statistics, 44(2):813–852, 2016.
Tamay Besiroglu, Ege Erdil, Matthew Barnett, and Josh You. Chinchilla scaling: a replication attempt. arXiv preprint arXiv:2404.10102, 2024.
Kevin Beyer, Jonathan Goldstein, Raghu Ramakrishnan, and Uri Shaft. When is “nearest neighbor” meaningful? In Database Theory, ICDT 1999, volume 1540 of Lecture Notes in Computer Science, 217–235. Springer, 1999.
Anna Maria Bianucci, Alessio Micheli, Alessandro Sperduti, and Antonina Starita. Application of cascade correlation networks for structures to chemistry. Applied Intelligence, 12(1-2):117–147, 2000. doi:10.1023/A:1008368105614.
Peter J. Bickel and David A. Freedman. Some asymptotic theory for the bootstrap. The Annals of Statistics, 9(6):1196–1217, 1981.
Dan Biderman, Jacob Portes, Jose Javier Gonzalez Ortiz, Mansheej Paul, Philip Greengard, Connor Jennings, Daniel King, Sam Havens, Vitaliy Chiley, Jonathan Frankle, Cody Blakeney, and John P. Cunningham. LoRA learns less and forgets less. Transactions on Machine Learning Research (TMLR), 2024. Featured Certification. URL: https://openreview.net/forum?id=aloEru2qCG.
Battista Biggio, Igino Corona, Davide Maiorca, Blaine Nelson, Nedim Šrndić, Pavel Laskov, Giorgio Giacinto, and Fabio Roli. Evasion attacks against machine learning at test time. In European Conference on Machine Learning and Knowledge Discovery in Databases (ECML PKDD). 2013. URL: https://arxiv.org/abs/1708.06131.
Çağdaş Bilen, Giacomo Ferroni, Francesco Tuveri, Juan Azcarreta, and Sacha Krstulović. A framework for the robust evaluation of sound event detection. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 61–65. 2020. doi:10.1109/ICASSP40776.2020.9052995.
Christopher M. Bishop. Training with noise is equivalent to Tikhonov regularization. Neural Computation, 7(1):108–116, 1995. doi:10.1162/neco.1995.7.1.108.
Mikołaj Bińkowski, Danica J. Sutherland, Michael Arbel, and Arthur Gretton. Demystifying MMD GANs. In International Conference on Learning Representations (ICLR). 2018.
Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Sergey Levine, Adrian Li-Bell, Mohith Mothukuri, Suraj Nair, Karl Pertsch, Lucy Xiaoyang Shi, James Tanner, Quan Vuong, Anna Walling, Haohuan Wang, and Ury Zhilinsky. $\pi _0$: a vision-language-action flow model for general robot control. In Proceedings of Robotics: Science and Systems (RSS). 2025. doi:10.15607/RSS.2025.XXI.010.
Kevin Black, Michael Janner, Yilun Du, Ilya Kostrikov, and Sergey Levine. Training diffusion models with reinforcement learning. In International Conference on Learning Representations (ICLR). 2024. URL: https://arxiv.org/abs/2305.13301.
Guy E. Blelloch. Prefix sums and their applications. Technical Report CMU-CS-90-190, School of Computer Science, Carnegie Mellon University, nov 1990.
Léonard Blier and Yann Ollivier. The description length of deep learning models. In Advances in Neural Information Processing Systems, volume 31. 2018.
H. D. Block and S. A. Levin. On the boundedness of an iterative procedure for solving a system of linear inequalities. Proceedings of the American Mathematical Society, 26(2):229–235, 1970. doi:10.1090/S0002-9939-1970-0265383-5.
Anselm Blumer, Andrzej Ehrenfeucht, David Haussler, and Manfred K. Warmuth. Learnability and the Vapnik-Chervonenkis dimension. Journal of the ACM, 36(4):929–965, 1989.
Daniel G. Bobrow, Ronald M. Kaplan, Martin Kay, Donald A. Norman, Henry Thompson, and Terry Winograd. GUS, a frame-driven dialog system. Artificial Intelligence, 8(2):155–173, 1977. doi:10.1016/0004-3702(77)90018-2.
Navaneeth Bodla, Bharat Singh, Rama Chellappa, and Larry S. Davis. Soft-NMS: improving object detection with one line of code. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 5562–5570. 2017. doi:10.1109/ICCV.2017.593.
Nicholas M. Boffi, Michael S. Albergo, and Eric Vanden-Eijnden. Flow map matching with stochastic interpolants: a mathematical framework for consistency models. Transactions on Machine Learning Research, 2025. URL: https://arxiv.org/abs/2406.07507.
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. Enriching word vectors with subword information. Transactions of the Association for Computational Linguistics, 5:135–146, 2017.
Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, and others. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258, 2021.
Yelysei Bondarenko, Markus Nagel, and Tijmen Blankevoort. Understanding and overcoming the challenges of efficient transformer quantization. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP), 7947–7969. 2021. doi:10.18653/v1/2021.emnlp-main.627.
Antoine Bordes, Nicolas Usunier, Alberto Garcia-Durán, Jason Weston, and Oksana Yakhnenko. Translating embeddings for modeling multi-relational data. In Advances in Neural Information Processing Systems, volume 26. 2013.
Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George van den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, Diego de Las Casas, Aurelia Guy, Jacob Menick, Roman Ring, Tom Hennigan, Saffron Huang, Loren Maggiore, Chris Jones, Albin Cassirer, Andy Brock, Michela Paganini, Geoffrey Irving, Oriol Vinyals, Simon Osindero, Karen Simonyan, Jack W. Rae, Erich Elsen, and Laurent Sifre. Improving language models by retrieving from trillions of tokens. In Proceedings of the 39th International Conference on Machine Learning (ICML), volume 162 of Proceedings of Machine Learning Research, 2206–2240. 2022. URL: https://proceedings.mlr.press/v162/borgeaud22a.html.
Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Dominik Roblek, Olivier Teboul, David Grangier, Marco Tagliasacchi, and Neil Zeghidour. AudioLM: a language modeling approach to audio generation. IEEE/ACM Transactions on Audio, Speech, and Language Processing (TASLP), 31:2523–2533, 2023.
Bernhard E. Boser, Isabelle M. Guyon, and Vladimir N. Vapnik. A training algorithm for optimal margin classifiers. In Proceedings of the Fifth Annual Workshop on Computational Learning Theory (COLT), 144–152. 1992.
H. Bourlard and Y. Kamp. Auto-association by multilayer perceptrons and singular value decomposition. Biological Cybernetics, 59(4–5):291–294, 1988.
Lucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia, Adelin Travers, Baiwu Zhang, David Lie, and Nicolas Papernot. Machine unlearning. In IEEE Symposium on Security and Privacy (S&P). 2021. URL: https://arxiv.org/abs/1912.03817.
Charles L. Bouton. Nim, a game with a complete mathematical theory. Annals of Mathematics, 3(1/4):35–39, 1901. doi:10.2307/1967631.
Samuel R. Bowman, Luke Vilnis, Oriol Vinyals, Andrew M. Dai, Rafal Jozefowicz, and Samy Bengio. Generating sentences from a continuous space. In Proceedings of the 20th SIGNLL Conference on Computational Natural Language Learning (CoNLL). 2016. URL: https://arxiv.org/abs/1511.06349.
George E. P. Box, Gwilym M. Jenkins, Gregory C. Reinsel, and Greta M. Ljung. Time Series Analysis: Forecasting and Control. Wiley, fifth edition, 2015.
Arwen Bradley and Preetum Nakkiran. Classifier-free guidance is a predictor-corrector. arXiv preprint arXiv:2408.09000, 2024.
Ralph Allan Bradley and Milton E. Terry. Rank analysis of incomplete block designs: i. the method of paired comparisons. Biometrika, 39(3/4):324–345, 1952.
Thorsten Brants. TnT: a statistical part-of-speech tagger. In Proceedings of the Sixth Conference on Applied Natural Language Processing (ANLP), 224–231. 2000. doi:10.3115/974147.974178.
Thorsten Brants and Alex Franz. Web 1T 5-gram version 1. Linguistic Data Consortium, Philadelphia, 2006. LDC2006T13. doi:10.35111/cqpa-a498.
Thorsten Brants, Ashok C. Popat, Peng Xu, Franz J. Och, and Jeffrey Dean. Large language models in machine translation. In Proceedings of the 2007 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning (EMNLP-CoNLL), 858–867. 2007. URL: https://aclanthology.org/D07-1090/.
Eric Breck, Shanqing Cai, Eric Nielsen, Michael Salib, and D. Sculley. The ML test score: a rubric for ML production readiness and technical debt reduction. In IEEE International Conference on Big Data (Big Data), 1123–1132. 2017.
Joseph L. Breeden. The simple mathematics of large language models. CRC Working Paper, Credit Research Centre, University of Edinburgh Business School, 2026. URL: https://www.crc.business-school.ed.ac.uk/sites/crc/files/2026-02/Mathematics_of_LLMs_2026.pdf.
John S. Breese, David Heckerman, and Carl Kadie. Empirical analysis of predictive algorithms for collaborative filtering. In Gregory F. Cooper and Serafín Moral, editors, Proceedings of the Fourteenth Conference on Uncertainty in Artificial Intelligence (UAI'98), 43–52. San Francisco, CA, 1998. Morgan Kaufmann. arXiv:1301.7363.
Leo Breiman. Bagging predictors. Machine Learning, 24(2):123–140, 1996.
Leo Breiman. Heuristics of instability and stabilization in model selection. The Annals of Statistics, 24(6):2350–2383, 1996.
Leo Breiman. Random forests. Machine Learning, 45(1):5–32, 2001.
Leo Breiman. Statistical modeling: the two cultures. Statistical Science, 16(3):199–231, 2001. doi:10.1214/ss/1009213726.
Leo Breiman, Jerome H. Friedman, Richard A. Olshen, and Charles J. Stone. Classification and Regression Trees. Wadsworth, 1984.
Timothy F. Bresnahan and Manuel Trajtenberg. General purpose technologies “engines of growth”? Journal of Econometrics, 65(1):83–108, 1995.
Drew Breunig. How long contexts fail. dbreunig.com, 22 giugno 2025, 2025. URL: https://www.dbreunig.com/2025/06/22/how-contexts-fail-and-how-to-fix-them.html.
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, and others. Towards monosemanticity: decomposing language models with dictionary learning. Transformer Circuits Thread, Anthropic, 2023. URL: https://transformer-circuits.pub/2023/monosemantic-features.
Glenn W. Brier. Verification of forecasts expressed in terms of probability. Monthly Weather Review, 78(1):1–3, 1950.
Andrew Brock, Jeff Donahue, and Karen Simonyan. Large scale GAN training for high fidelity natural image synthesis. In International Conference on Learning Representations (ICLR). 2019. arXiv:1809.11096.
Shaked Brody, Uri Alon, and Eran Yahav. How attentive are graph attention networks? In International Conference on Learning Representations (ICLR). 2022. URL: https://arxiv.org/abs/2105.14491.
Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Xi Chen, Krzysztof Choromanski, Tianli Ding, Danny Driess, Avinava Dubey, Chelsea Finn, and others. RT-2: vision-language-action models transfer web knowledge to robotic control. In Proceedings of the 7th Conference on Robot Learning (CoRL), volume 229 of Proceedings of Machine Learning Research, 2165–2183. 2023.
Nathan Bronson, Abutalib Aghayev, Aleksey Charapko, and Timothy Zhu. Metastable failures in distributed systems. In Proceedings of the Workshop on Hot Topics in Operating Systems (HotOS), 221–227. 2021. doi:10.1145/3458336.3465286.
Michael M. Bronstein, Joan Bruna, Taco Cohen, and Petar Veličković. Geometric deep learning: grids, groups, graphs, geodesics, and gauges. arXiv preprint arXiv:2104.13478, 2021. URL: https://arxiv.org/abs/2104.13478.
Marc Brooker. Exponential backoff and jitter. AWS Architecture Blog, March 2015. URL: https://aws.amazon.com/blogs/architecture/exponential-backoff-and-jitter/.
Tim Brooks, Aleksander Holynski, and Alexei A. Efros. InstructPix2Pix: learning to follow image editing instructions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 18392–18402. 2023. doi:10.1109/CVPR52729.2023.01764.
Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, Clarence Ng, Ricky Wang, and Aditya Ramesh. Video generation models as world simulators. OpenAI technical report, 2024. URL: https://openai.com/research/video-generation-models-as-world-simulators.
Noam Brown and Tuomas Sandholm. Superhuman AI for heads-up no-limit poker: Libratus beats top professionals. Science, 359(6374):418–424, 2018. doi:10.1126/science.aao1733.
Noam Brown and Tuomas Sandholm. Superhuman AI for multiplayer poker. Science, 365(6456):885–890, 2019.
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, and others. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33. 2020.
Jonah Brown-Cohen, Geoffrey Irving, and Georgios Piliouras. Scalable ai safety via doubly-efficient debate. arXiv preprint arXiv:2311.14125, 2023. URL: https://arxiv.org/abs/2311.14125.
Jonah Brown-Cohen, Geoffrey Irving, Georgios Piliouras, Lijie Chen, Jiawei Li, and Zhiyang Xun. Avoiding obfuscation with prover-estimator debate. arXiv preprint arXiv:2506.13609, 2025. URL: https://arxiv.org/abs/2506.13609.
Jake Bruce, Michael Dennis, Ashley Edwards, Jack Parker-Holder, Yuge Shi, Edward Hughes, Matthew Lai, Aditi Mavalankar, Richie Steigerwald, Chris Apps, Yusuf Aytar, Sarah Bechtle, Feryal Behbahani, Stephanie Chan, Nicolas Heess, Lucy Gonzalez, Simon Osindero, Sherjil Ozair, Scott Reed, Jingwei Zhang, Konrad Zolna, Jeff Clune, Nando de Freitas, Satinder Singh, and Tim Rocktäschel. Genie: generative interactive environments. In Proceedings of the 41st International Conference on Machine Learning (ICML). 2024. URL: https://arxiv.org/abs/2402.15391.
Sebastian Bruch, Siyu Gai, and Amir Ingber. An analysis of fusion functions for hybrid retrieval. ACM Transactions on Information Systems, 42(1):1–35, 2023. Pubblicato online il 18 agosto 2023; fascicolo di gennaio 2024. doi:10.1145/3596512.
Joan Bruna, Wojciech Zaremba, Arthur Szlam, and Yann LeCun. Spectral networks and locally connected networks on graphs. In International Conference on Learning Representations. 2014.
Bernd Brügmann. Monte Carlo Go. 1993. Rapporto tecnico non pubblicato.
Cristian Buciluǎ, Rich Caruana, and Alexandru Niculescu-Mizil. Model compression. In Proceedings of the 12th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD '06), 535–541. ACM, 2006. doi:10.1145/1150402.1150464.
Andreas Buja, Trevor Hastie, and Robert Tibshirani. Linear smoothers and additive models. The Annals of Statistics, 17(2):453–510, 1989. doi:10.1214/aos/1176347115.
Joy Buolamwini and Timnit Gebru. Gender shades: intersectional accuracy disparities in commercial gender classification. In Proceedings of the 1st Conference on Fairness, Accountability and Transparency (FAT*), volume 81 of Proceedings of Machine Learning Research, 77–91. 2018. URL: https://proceedings.mlr.press/v81/buolamwini18a.html.
Yuri Burda, Harrison Edwards, Amos Storkey, and Oleg Klimov. Exploration by random network distillation. In International Conference on Learning Representations (ICLR). 2019. URL: https://arxiv.org/abs/1810.12894.
Yuri Burda, Roger Grosse, and Ruslan Salakhutdinov. Importance weighted autoencoders. In International Conference on Learning Representations (ICLR). 2016.
Kenneth P. Burnham and David R. Anderson. Multimodel inference: understanding AIC and BIC in model selection. Sociological Methods & Research, 33(2):261–304, 2004. doi:10.1177/0049124104268644.
Tianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng, Jason D. Lee, Deming Chen, and Tri Dao. Medusa: simple LLM inference acceleration framework with multiple decoding heads. In Proceedings of the 41st International Conference on Machine Learning (ICML). 2024.
Aylin Caliskan, Joanna J. Bryson, and Arvind Narayanan. Semantics derived automatically from language corpora contain human-like biases. Science, 356(6334):183–186, 2017. doi:10.1126/science.aal4230.
Tadeusz Caliński and Jerzy Harabasz. A dendrite method for cluster analysis. Communications in Statistics, 3(1):1–27, 1974. doi:10.1080/03610927408827101.
Andrew Campbell, Joe Benton, Valentin De Bortoli, Tom Rainforth, George Deligiannidis, and Arnaud Doucet. A continuous time framework for discrete denoising models. In Advances in Neural Information Processing Systems (NeurIPS). 2022. URL: https://arxiv.org/abs/2205.14987.
Ricardo J. G. B. Campello, Davoud Moulavi, and Jörg Sander. Density-based clustering based on hierarchical density estimates. In Advances in Knowledge Discovery and Data Mining, PAKDD 2013, volume 7819 of Lecture Notes in Computer Science, 160–172. Springer, 2013.
John Canny. A computational approach to edge detection. IEEE Transactions on Pattern Analysis and Machine Intelligence, PAMI-8(6):679–698, 1986. doi:10.1109/TPAMI.1986.4767851.
Jaime Carbonell and Jade Goldstein. The use of MMR, diversity-based reranking for reordering documents and producing summaries. In Proceedings of the 21st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, 335–336. 1998.
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In European Conference on Computer Vision (ECCV), 213–229. 2020.
Nicholas Carlini, Steve Chien, Milad Nasr, Shuang Song, Andreas Terzis, and Florian Tramèr. Membership inference attacks from first principles. In 2022 IEEE Symposium on Security and Privacy (SP), 1897–1914. 2022. arXiv:2112.03570, doi:10.1109/SP46214.2022.9833649.
Nicholas Carlini, Jamie Hayes, Milad Nasr, Matthew Jagielski, Vikash Sehwag, Florian Tramèr, Borja Balle, Daphne Ippolito, and Eric Wallace. Extracting training data from diffusion models. In 32nd USENIX Security Symposium (USENIX Security 23), 5253–5270. USENIX Association, 2023. URL: https://arxiv.org/abs/2301.13188.
Nicholas Carlini, Matthew Jagielski, Christopher A. Choquette-Choo, Daniel Paleka, Will Pearce, Hyrum Anderson, Andreas Terzis, Kurt Thomas, and Florian Tramèr. Poisoning web-scale training datasets is practical. In IEEE Symposium on Security and Privacy (S&P). 2024. URL: https://arxiv.org/abs/2302.10149.
Nicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Úlfar Erlingsson, Alina Oprea, and Colin Raffel. Extracting training data from large language models. In 30th USENIX Security Symposium (USENIX Security 21), 2633–2650. 2021.
Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin. Unsupervised learning of visual features by contrasting cluster assignments. In Advances in Neural Information Processing Systems (NeurIPS). 2020. URL: https://arxiv.org/abs/2006.09882.
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 9650–9660. 2021. URL: https://arxiv.org/abs/2104.14294.
Rich Caruana. Multitask learning. Machine Learning, 28(1):41–75, 1997. URL: https://doi.org/10.1023/A:1007379606734.
Rich Caruana, Yin Lou, Johannes Gehrke, Paul Koch, Marc Sturm, and Noemie Elhadad. Intelligible models for healthcare: predicting pneumonia risk and hospital 30-day readmission. In ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD). 2015.
Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, Jérémy Scheurer, Javier Rando, and others. Open problems and fundamental limitations of reinforcement learning from human feedback. Transactions on Machine Learning Research, 2023. Survey Certification, Featured Certification. URL: https://openreview.net/forum?id=bx24KpJ4Eb.
Miguel Castro and Barbara Liskov. Practical Byzantine fault tolerance. In Proceedings of the Third Symposium on Operating Systems Design and Implementation (OSDI '99), 173–186. USENIX Association, 1999. URL: https://www.usenix.org/legacy/publications/library/proceedings/osdi99/castro.html.
Gavin C. Cawley and Nicola L. C. Talbot. On over-fitting in model selection and subsequent selection bias in performance evaluation. Journal of Machine Learning Research, 11(70):2079–2107, 2010. URL: https://jmlr.org/papers/v11/cawley10a.html.
Mert Cemri, Melissa Z. Pan, Shuyi Yang, Lakshya A. Agrawal, Bhavya Chopra, Rishabh Tiwari, Kurt Keutzer, Aditya Parameswaran, Dan Klein, Kannan Ramchandran, Matei Zaharia, Joseph E. Gonzalez, and Ion Stoica. Why do multi-agent llm systems fail? In Advances in Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track. 2025. URL: https://arxiv.org/abs/2503.13657.
Nicolò Cesa-Bianchi, Alex Conconi, and Claudio Gentile. On the generalization ability of on-line learning algorithms. IEEE Transactions on Information Theory, 50(9):2050–2057, 2004. doi:10.1109/TIT.2004.833339.
Gregory J. Chaitin. On the length of programs for computing finite binary sequences. Journal of the ACM, 13(4):547–569, 1966. doi:10.1145/321356.321363.
Gregory J. Chaitin. Information-theoretic limitations of formal systems. Journal of the ACM, 21(3):403–424, 1974. doi:10.1145/321832.321839.
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals. Listen, attend and spell: a neural network for large vocabulary conversational speech recognition. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 2016.
Tushar Deepak Chandra and Sam Toueg. Unreliable failure detectors for reliable distributed systems. Journal of the ACM, 43(2):225–267, 1996. doi:10.1145/226643.226647.
Anantha P. Chandrakasan. Low Power Digital CMOS Design. PhD thesis, University of California, Berkeley, 1994. Memorandum UCB/ERL M94/65, Electronics Research Laboratory.
Yapei Chang, Kyle Lo, Tanya Goyal, and Mohit Iyyer. BooookScore: a systematic exploration of book-length summarization in the era of LLMs. In International Conference on Learning Representations (ICLR). 2024. arXiv:2310.00785.
David Chanin, James Wilken-Smith, Tomáš Dulka, Hardik Bhatnagar, Satvik Golechha, and Joseph Bloom. A is for absorption: studying feature splitting and absorption in sparse autoencoders. arXiv preprint arXiv:2409.14507, 2024. URL: https://arxiv.org/abs/2409.14507.
Olivier Chapelle and Lihong Li. An empirical evaluation of Thompson sampling. In Advances in Neural Information Processing Systems, volume 24, 2249–2257. 2011.
Olivier Chapelle, Jason Weston, Léon Bottou, and Vladimir Vapnik. Vicinal risk minimization. In Advances in Neural Information Processing Systems (NIPS), volume 13. 2000.
Nitesh V. Chawla, Kevin W. Bowyer, Lawrence O. Hall, and W. Philip Kegelmeyer. SMOTE: synthetic minority over-sampling technique. Journal of Artificial Intelligence Research, 16:321–357, 2002. doi:10.1613/jair.953.
Charlie Chen, Sebastian Borgeaud, Geoffrey Irving, Jean-Baptiste Lespiau, Laurent Sifre, and John Jumper. Accelerating large language model decoding with speculative sampling. arXiv preprint arXiv:2302.01318, 2023.
Danqi Chen and Christopher D. Manning. A fast and accurate dependency parser using neural networks. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), 740–750. 2014.
Keqin Chen, Zhao Zhang, Weili Zeng, Richong Zhang, Feng Zhu, and Rui Zhao. Shikra: unleashing multimodal LLM's referential dialogue magic. arXiv preprint arXiv:2306.15195, 2023.
Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Michael Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch. Decision transformer: reinforcement learning via sequence modeling. In Advances in Neural Information Processing Systems (NeurIPS). 2021. URL: https://arxiv.org/abs/2106.01345.
Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Jiaqi Wang, Yu Qiao, Dahua Lin, and Feng Zhao. Are we on the right way for evaluating large vision-language models? arXiv preprint arXiv:2403.20330, 2024.
Lingjiao Chen, Matei Zaharia, and James Zou. How is ChatGPT's behavior changing over time? Harvard Data Science Review, 2024. doi:10.1162/99608f92.5317da47.
Lingjiao Chen, Matei Zaharia, and James Zou. FrugalGPT: how to use large language models while reducing cost and improving performance. Transactions on Machine Learning Research, 2024.
Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, and Ilya Sutskever. Generative pretraining from pixels. In International Conference on Machine Learning (ICML). 2020.
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Josh Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. Evaluating large language models trained on code. OpenAI, arXiv:2107.03374, 2021. URL: https://arxiv.org/abs/2107.03374.
Nanxin Chen, Yu Zhang, Heiga Zen, Ron J. Weiss, Mohammad Norouzi, and William Chan. WaveGrad: estimating gradients for waveform generation. In International Conference on Learning Representations (ICLR). 2021.
Ricky T. Q. Chen, Yulia Rubanova, Jesse Bettencourt, and David Duvenaud. Neural ordinary differential equations. In Advances in Neural Information Processing Systems (NeurIPS). 2018.
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, Jian Wu, Long Zhou, Shuo Ren, Yanmin Qian, Yao Qian, Jian Wu, Michael Zeng, Xiangzhan Yu, and Furu Wei. WavLM: large-scale self-supervised pre-training for full stack speech processing. IEEE Journal of Selected Topics in Signal Processing, 16(6):1505–1518, 2022. doi:10.1109/JSTSP.2022.3188113.
Sanyuan Chen, Yu Wu, Chengyi Wang, Shujie Liu, Daniel Tompkins, Zhuo Chen, Wanxiang Che, Xiangzhan Yu, and Furu Wei. BEATs: audio pre-training with acoustic tokenizers. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, 5178–5193. PMLR, 2023.
Scott Shaobing Chen, David L. Donoho, and Michael A. Saunders. Atomic decomposition by basis pursuit. SIAM Journal on Scientific Computing, 20(1):33–61, 1998. doi:10.1137/S1064827596304010.
Shouyuan Chen, Sherman Wong, Liangjian Chen, and Yuandong Tian. Extending context window of large language models via positional interpolation. arXiv preprint arXiv:2306.15595, 2023.
Stanley F. Chen and Joshua Goodman. An empirical study of smoothing techniques for language modeling. Computer Speech & Language, 13(4):359–394, 1999. doi:10.1006/csla.1999.0128.
Tianping Chen and Hong Chen. Universal approximation to nonlinear operators by neural networks with arbitrary activation functions and its application to dynamical systems. IEEE Transactions on Neural Networks, 6(4):911–917, 1995. doi:10.1109/72.392253.
Tianqi Chen and Carlos Guestrin. Xgboost: a scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD), 785–794. 2016. URL: https://arxiv.org/abs/1603.02754.
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In Proceedings of the 37th International Conference on Machine Learning (ICML), 1597–1607. 2020. URL: https://arxiv.org/abs/2002.05709.
Xinlei Chen and Kaiming He. Exploring simple siamese representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 15750–15758. 2021. URL: https://arxiv.org/abs/2011.10566.
Yanda Chen, Joe Benton, Ansh Radhakrishnan, Jonathan Uesato, Carson Denison, John Schulman, Arushi Somani, Peter Hase, Misha Wagner, Fabien Roger, Vlad Mikulik, Samuel R. Bowman, Jan Leike, Jared Kaplan, and Ethan Perez. Reasoning models don't always say what they think. arXiv preprint arXiv:2505.05410, 2025.
Yu-Hsin Chen, Joel Emer, and Vivienne Sze. Eyeriss: a spatial architecture for energy-efficient dataflow for convolutional neural networks. In Proceedings of the 43rd International Symposium on Computer Architecture (ISCA), 367–379. 2016. doi:10.1109/ISCA.2016.40.
Yu-Hsin Chen, Tushar Krishna, Joel S. Emer, and Vivienne Sze. Eyeriss: an energy-efficient reconfigurable accelerator for deep convolutional neural networks. IEEE Journal of Solid-State Circuits, 52(1):127–138, 2017. doi:10.1109/JSSC.2016.2616357.
Zhe Chen, Weiyun Wang, Hao Tian, Shenglong Ye, Zhangwei Gao, Erfei Cui, Wenwen Tong, Kongzhi Hu, Jiapeng Luo, Zheng Ma, and others. How far are we to GPT-4V? closing the gap to commercial multimodal models with open-source suites. arXiv preprint arXiv:2404.16821, 2024.
Boris Cherny. I don't prompt Claude anymore. I have loops running that prompt Claude and figuring out what to do. My job is to write loops. Frase attribuita a Cherny da Addy Osmani (7 giugno 2026), con rimando a un post di Rohan Paul del 6 giugno 2026 poi ritirato; ripresa nel repository \emph loop-engineering di Cobus Greyling, 2026. URL: cobusgreyling/loop-engineering.
E. Colin Cherry. Some experiments on the recognition of speech, with one and with two ears. Journal of the Acoustical Society of America, 25(5):975–979, 1953. doi:10.1121/1.1907229.
Kyunghyun Cho, Bart van Merriënboer, Dzmitry Bahdanau, and Yoshua Bengio. On the properties of neural machine translation: encoder–decoder approaches. In Proceedings of SSST-8, Eighth Workshop on Syntax, Semantics and Structure in Statistical Translation, 103–111. Doha, Qatar, 2014. Association for Computational Linguistics. doi:10.3115/v1/W14-4012.
Kyunghyun Cho, Bart van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. Learning phrase representations using RNN encoder–decoder for statistical machine translation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2014.
François Chollet. On the measure of intelligence. arXiv preprint arXiv:1911.01547, 2019.
François Chollet. Xception: deep learning with depthwise separable convolutions. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 1251–1258. 2017. URL: https://arxiv.org/abs/1610.02357.
François Chollet and others. Keras. 2015. Software. URL: https://keras.io.
François Chollet and Matthew Watson. Deep Learning with Python. Manning Publications, third edition, 2025. ISBN 978-1-63343-658-9. Leggibile integralmente online. URL: https://deeplearningwithpython.io/.
Noam Chomsky. Three models for the description of language. IRE Transactions on Information Theory, 2(3):113–124, 1956. doi:10.1109/TIT.1956.1056813.
Noam Chomsky. On certain formal properties of grammars. Information and Control, 2(2):137–167, 1959.
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamás Sarlós, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, David Belanger, Lucy Colwell, and Adrian Weller. Rethinking attention with Performers. In International Conference on Learning Representations (ICLR). 2021.
Alexandra Chouldechova. Fair prediction with disparate impact: a study of bias in recidivism prediction instruments. Big Data, 5(2):153–163, 2017. URL: https://arxiv.org/abs/1703.00056.
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, and others. PaLM: scaling language modeling with pathways. Journal of Machine Learning Research, 24(240):1–113, 2023.
Paul F. Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei. Deep reinforcement learning from human preferences. In Advances in Neural Information Processing Systems, volume 30. 2017. URL: https://arxiv.org/abs/1706.03741.
Casey Chu, Andrey Zhmoginov, and Mark Sandler. CycleGAN, a master of steganography. In NIPS 2017 Workshop on Machine Deception. 2017. arXiv:1712.02950.
Kurtland Chua, Roberto Calandra, Rowan McAllister, and Sergey Levine. Deep reinforcement learning in a handful of trials using probabilistic dynamics models. In Advances in Neural Information Processing Systems 31 (NeurIPS). 2018. arXiv:1805.12114.
Hyungjin Chung, Jeongsol Kim, Michael T. McCann, Marc L. Klasky, and Jong Chul Ye. Diffusion posterior sampling for general noisy inverse problems. In International Conference on Learning Representations (ICLR). 2023. URL: https://arxiv.org/abs/2209.14687.
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio. Empirical evaluation of gated recurrent neural networks on sequence modeling. 2014. URL: https://arxiv.org/abs/1412.3555, arXiv:1412.3555.
Alonzo Church. An unsolvable problem of elementary number theory. American Journal of Mathematics, 58(2):345–363, 1936.
Kenneth Church and Ramesh Patil. Coping with syntactic ambiguity or how to put the block in the box on the table. American Journal of Computational Linguistics, 8(3–4):139–149, 1982.
Dan Cireşan, Ueli Meier, Jonathan Masci, and Jürgen Schmidhuber. A committee of neural networks for traffic sign classification. In The 2011 International Joint Conference on Neural Networks (IJCNN), 1918–1921. 2011. doi:10.1109/IJCNN.2011.6033458.
Aidan Clark, Diego de las Casas, Aurelia Guy, Arthur Mensch, Michela Paganini, Jordan Hoffmann, Bogdan Damoc, Blake Hechtman, Trevor Cai, Sebastian Borgeaud, George van den Driessche, Eliza Rutherford, Tom Hennigan, Matthew Johnson, Katie Millican, Albin Cassirer, Chris Jones, Elena Buchatskaya, David Budden, Laurent Sifre, Simon Osindero, Oriol Vinyals, Jack Rae, Erich Elsen, Koray Kavukcuoglu, and Karen Simonyan. Unified scaling laws for routed language models. In Proceedings of the 39th International Conference on Machine Learning (ICML), volume 162 of Proceedings of Machine Learning Research, 4057–4086. 2022. URL: https://proceedings.mlr.press/v162/clark22a.html.
Herbert H. Clark. Using Language. Cambridge University Press, 1996. doi:10.1017/CBO9780511620539.
Jack Clark and Dario Amodei. Faulty reward functions in the wild. OpenAI Blog, December 2016. URL: https://openai.com/index/faulty-reward-functions/.
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. What does BERT look at? an analysis of BERT's attention. In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, 276–286. 2019. URL: https://aclanthology.org/W19-4828/.
Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. ELECTRA: pre-training text encoders as discriminators rather than generators. In International Conference on Learning Representations (ICLR). 2020. URL: https://arxiv.org/abs/2003.10555.
Maurice Clerc and James Kennedy. The particle swarm: explosion, stability, and convergence in a multidimensional complex space. IEEE Transactions on Evolutionary Computation, 6(1):58–73, 2002. doi:10.1109/4235.985692.
Robert B. Cleveland, William S. Cleveland, Jean E. McRae, and Irma Terpenning. STL: a seasonal-trend decomposition procedure based on loess. Journal of Official Statistics, 6(1):3–73, 1990.
William S. Cleveland. Robust locally weighted regression and smoothing scatterplots. Journal of the American Statistical Association, 74(368):829–836, 1979.
Jeremy Cohen, Elan Rosenfeld, and J. Zico Kolter. Certified adversarial robustness via randomized smoothing. In Proceedings of the 36th International Conference on Machine Learning (ICML), 1310–1320. 2019.
Ronan Collobert, Jason Weston, Léon Bottou, Michael Karlen, Koray Kavukcuoglu, and Pavel Kuksa. Natural language processing (almost) from scratch. Journal of Machine Learning Research, 12:2493–2537, 2011.
Matteo Colombo and Cory Wright. First principles in the life sciences: the free-energy principle, organicism, and mechanism. Synthese, 198:3463–3488, 2021.
Pierre Comon. Independent component analysis, a new concept? Signal Processing, 36(3):287–314, 1994.
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 8440–8451. 2020.
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. XNLI: evaluating cross-lingual sentence representations. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2475–2485. Brussels, Belgium, 2018. Association for Computational Linguistics. doi:10.18653/v1/D18-1269.
Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi, and Alexandre Défossez. Simple and controllable music generation. In Advances in Neural Information Processing Systems (NeurIPS), volume 36. 2023.
Pierre-Arnaud Coquelin and Rémi Munos. Bandit algorithms for tree search. In Proceedings of the Twenty-Third Conference on Uncertainty in Artificial Intelligence (UAI), 67–74. 2007. URL: https://arxiv.org/abs/1408.2028, arXiv:1408.2028.
Gordon V. Cormack, Charles L. A. Clarke, and Stefan Büttcher. Reciprocal rank fusion outperforms Condorcet and individual rank learning methods. In Proceedings of the 32nd International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR), 758–759. 2009. doi:10.1145/1571941.1572114.
Corinna Cortes and Vladimir Vapnik. Support-vector networks. Machine Learning, 20(3):273–297, 1995.
Rémi Coulom. Efficient selectivity and backup operators in Monte-Carlo tree search. In Computers and Games (CG 2006), volume 4630 of Lecture Notes in Computer Science, 72–83. Springer, 2006.
Martin Courtois, Malte Ostendorff, Leonhard Hennig, and Georg Rehm. Symmetric dot-product attention for efficient training of BERT language models. In Findings of the Association for Computational Linguistics: ACL 2024. 2024.
Thomas M. Cover. Geometrical and statistical properties of systems of linear inequalities with applications in pattern recognition. IEEE Transactions on Electronic Computers, EC-14(3):326–334, 1965.
Thomas M. Cover and Peter E. Hart. Nearest neighbor pattern classification. IEEE Transactions on Information Theory, 13(1):21–27, 1967.
Thomas M. Cover and Joy A. Thomas. Elements of Information Theory. Wiley, 2 edition, 2006.
Paul Covington, Jay Adams, and Emre Sargin. Deep neural networks for YouTube recommendations. In Proceedings of the 10th ACM Conference on Recommender Systems (RecSys), 191–198. 2016. doi:10.1145/2959100.2959190.
Kenneth James Williams Craik. The Nature of Explanation. Cambridge University Press, 1943.
Andrea Crisanti, Daniel J. Amit, and Hanoch Gutfreund. Saturation level of the Hopfield model for neural network. Europhysics Letters, 2(4):337–341, 1986. doi:10.1209/0295-5075/2/4/012.
Francesco Croce and Matthias Hein. Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks. In Proceedings of the 37th International Conference on Machine Learning (ICML), volume 119 of Proceedings of Machine Learning Research, 2206–2216. PMLR, 2020. URL: https://proceedings.mlr.press/v119/croce20b.html.
Ekin D. Cubuk, Barret Zoph, Dandelion Mané, Vijay Vasudevan, and Quoc V. Le. AutoAugment: learning augmentation strategies from data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 113–123. 2019.
George Cybenko. Approximation by superpositions of a sigmoidal function. Mathematics of Control, Signals and Systems, 2(4):303–314, 1989.
C. A. Dahlbom and J. S. Ryan. Common channel interoffice signaling: history and description of a new signaling system. The Bell System Technical Journal, 57(2):225–250, 1978. URL: https://archive.org/details/bstj57-2-225.
Damai Dai, Chengqi Deng, Chenggang Zhao, R. X. Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Y. Wu, Zhenda Xie, Y. K. Li, Panpan Huang, Fuli Luo, Chong Ruan, Zhifang Sui, and Wenfeng Liang. DeepSeekMoE: towards ultimate expert specialization in mixture-of-experts language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1280–1297. Bangkok, Thailand, 2024. Association for Computational Linguistics. URL: https://aclanthology.org/2024.acl-long.70/, doi:10.18653/v1/2024.acl-long.70.
Wenliang Dai, Nayeon Lee, Boxin Wang, Zhuolin Yang, Zihan Liu, Jon Barker, Tuomas Rintamaki, Mohammad Shoeybi, Bryan Catanzaro, and Wei Ping. NVLM: open frontier-class multimodal LLMs. arXiv preprint arXiv:2409.11402, 2024.
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. InstructBLIP: towards general-purpose vision-language models with instruction tuning. In Advances in Neural Information Processing Systems (NeurIPS). 2023.
Navneet Dalal and Bill Triggs. Histograms of oriented gradients for human detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2005.
Fred J. Damerau. A technique for computer detection and correction of spelling errors. Communications of the ACM, 7(3):171–176, 1964. doi:10.1145/363958.363994.
Tri Dao. FlashAttention-2: faster attention with better parallelism and work partitioning. In International Conference on Learning Representations (ICLR). 2024.
Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. FlashAttention: fast and memory-efficient exact attention with IO-awareness. In Advances in Neural Information Processing Systems (NeurIPS). 2022.
Tri Dao and Albert Gu. Transformers are SSMs: generalized models and efficient algorithms through structured state space duality. In Proceedings of the 41st International Conference on Machine Learning (ICML). 2024.
Abhimanyu Das, Weihao Kong, Rajat Sen, and Yichen Zhou. A decoder-only foundation model for time-series forecasting. In Proceedings of the 41st International Conference on Machine Learning (ICML). 2024. URL: https://arxiv.org/abs/2310.10688.
Jeffrey Dastin. Amazon scraps secret AI recruiting tool that showed bias against women. Reuters, October 2018. URL: https://www.reuters.com/article/us-amazon-com-jobs-automation-insight-idUSKCN1MK08G/.
Yann N. Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio. Identifying and attacking the saddle point problem in high-dimensional non-convex optimization. In Advances in Neural Information Processing Systems, volume 27. 2014.
David L. Davies and Donald W. Bouldin. A cluster separation measure. IEEE Transactions on Pattern Analysis and Machine Intelligence, PAMI-1(2):224–227, 1979. doi:10.1109/TPAMI.1979.4766909.
Steven B. Davis and Paul Mermelstein. Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences. IEEE Transactions on Acoustics, Speech, and Signal Processing, 28(4):357–366, 1980.
Peter Dayan. The convergence of TD(λ) for general λ. Machine Learning, 8(3):341–362, 1992.
Nicola De Cao and Thomas Kipf. MolGAN: an implicit generative model for small molecular graphs. In ICML 2018 Workshop on Theoretical Foundations and Applications of Deep Generative Models. Stockholm, Sweden, 2018. URL: https://arxiv.org/abs/1805.11973.
Marie-Catherine de Marneffe, Christopher D. Manning, Joakim Nivre, and Daniel Zeman. Universal dependencies. Computational Linguistics, 47(2):255–308, 2021.
Tim De Ryck, Florent Bonnet, Siddhartha Mishra, and Emmanuel de Bézenac. An operator preconditioning perspective on training in physics-informed machine learning. In International Conference on Learning Representations (ICLR). 2024.
Tim De Ryck, Siddhartha Mishra, and Roberto Molinaro. WPINNs: weak physics informed neural networks for approximating entropy solutions of hyperbolic conservation laws. SIAM Journal on Numerical Analysis, 62(2):811–841, 2024. URL: https://arxiv.org/abs/2207.08483.
Jeffrey Dean and Luiz André Barroso. The tail at scale. Communications of the ACM, 56(2):74–80, February 2013. doi:10.1145/2408776.2408794.
Kalyanmoy Deb, Amrit Pratap, Sameer Agarwal, and T. Meyarivan. A fast and elitist multiobjective genetic algorithm: nsga-ii. IEEE Transactions on Evolutionary Computation, 6(2):182–197, 2002.
Edoardo Debenedetti, Ilia Shumailov, Tianqi Fan, Jamie Hayes, Nicholas Carlini, Daniel Fabian, Christoph Kern, Chongyang Shi, Andreas Terzis, and Florian Tramèr. Defeating prompt injections by design. arXiv preprint arXiv:2503.18813, 2025.
Edoardo Debenedetti, Jie Zhang, Mislav Balunović, Luca Beurer-Kellner, Marc Fischer, and Florian Tramèr. AgentDojo: a dynamic environment to evaluate prompt injection attacks and defenses for LLM agents. In Advances in Neural Information Processing Systems 37 (NeurIPS 2024), Datasets and Benchmarks Track, volume 37, 82895–82920. Curran Associates, Inc., 2024. URL: https://arxiv.org/abs/2406.13352, doi:10.52202/079017-2636.
Rina Dechter and Judea Pearl. Generalized best-first search strategies and the optimality of A*. Journal of the ACM, 32(3):505–536, 1985.
Scott Deerwester, Susan T. Dumais, George W. Furnas, Thomas K. Landauer, and Richard Harshman. Indexing by latent semantic analysis. Journal of the American Society for Information Science, 41(6):391–407, 1990. doi:10.1002/(SICI)1097-4571(199009)41:6<391::AID-ASI1>3.0.CO;2-9.
Michaël Defferrard, Xavier Bresson, and Pierre Vandergheynst. Convolutional neural networks on graphs with fast localized spectral filtering. In Advances in Neural Information Processing Systems (NeurIPS), volume 29. 2016. URL: https://arxiv.org/abs/1606.09375.
Morris H. DeGroot and Stephen E. Fienberg. The comparison and evaluation of forecasters. Journal of the Royal Statistical Society: Series D (The Statistician), 32(1–2):12–22, 1983. doi:10.2307/2987588.
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Łukasz Kaiser. Universal transformers. In International Conference on Learning Representations (ICLR). 2019. URL: https://arxiv.org/abs/1807.03819.
Mostafa Dehghani, Basil Mustafa, Josip Djolonga, and others. Patch n' pack: NaViT, a vision transformer for any aspect ratio and resolution. In Advances in Neural Information Processing Systems (NeurIPS). 2023.
Grégoire Delétang, Anian Ruoss, Paul-Ambroise Duquenne, Elliot Catt, Tim Genewein, Christopher Mattern, Jordi Grau-Moya, Li Kevin Wenliang, Matthew Aitchison, Laurent Orseau, Marcus Hutter, and Joel Veness. Language modeling is compression. In International Conference on Learning Representations (ICLR). 2024.
Mete Demircigil, Judith Heusel, Matthias Löwe, Sven Upgang, and Franck Vermet. On a model of associative memory with huge storage capacity. Journal of Statistical Physics, 168(2):288–299, 2017. URL: https://arxiv.org/abs/1702.01929.
Arthur P. Dempster, Nan M. Laird, and Donald B. Rubin. Maximum likelihood from incomplete data via the EM algorithm. Journal of the Royal Statistical Society, Series B (Methodological), 39(1):1–22, 1977. URL: https://doi.org/10.1111/j.2517-6161.1977.tb01600.x.
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: a large-scale hierarchical image database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 248–255. 2009. doi:10.1109/CVPR.2009.5206848.
Xiang Deng, Jeff Da, Edwin Pan, Yannis Yiming He, Charles Ide, Kanak Garg, Niklas Lauffer, Andrew Park, Nitin Pasari, Chetan Rane, Karmini Sampath, Maya Krishnan, Srivatsa Kundurthy, Sean Hendryx, Zifan Wang, Vijay Bharadwaj, Jeff Holm, Raja Aluri, Chen Bo Calvin Zhang, Noah Jacobson, Bing Liu, and Brad Kenstler. SWE-Bench Pro: can AI agents solve long-horizon software engineering tasks? 2025. arXiv:2509.16941.
R. H. Dennard, F. H. Gaensslen, Hwa-Nien Yu, V. L. Rideout, E. Bassous, and A. R. LeBlanc. Design of ion-implanted MOSFET's with very small physical dimensions. IEEE Journal of Solid-State Circuits, 9(5):256–268, 1974. doi:10.1109/JSSC.1974.1050511.
Austin Derrow-Pinion, Jennifer She, David Wong, Oliver Lange, Todd Hester, Luis Perez, Marc Nunkesser, Seongjae Lee, Xueying Guo, Brett Wiltshire, Peter W. Battaglia, Vishal Gupta, Ang Li, Zhongwen Xu, Alvaro Sanchez-Gonzalez, Yujia Li, and Petar Velickovic. Eta prediction with graph neural networks in google maps. In Proceedings of the 30th ACM International Conference on Information and Knowledge Management. 2021.
Nicki S. Detlefsen, Martin Jørgensen, and Søren Hauberg. Reliable training and estimation of variance networks. In Advances in Neural Information Processing Systems (NeurIPS), volume 32. 2019. URL: https://arxiv.org/abs/1906.03260, arXiv:1906.03260.
Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. LLM.int8(): 8-bit matrix multiplication for transformers at scale. In Advances in Neural Information Processing Systems (NeurIPS), volume 35. 2022.
Tim Dettmers, Pasquale Minervini, Pontus Stenetorp, and Sebastian Riedel. Convolutional 2D knowledge graph embeddings. In AAAI Conference on Artificial Intelligence. 2018. URL: https://arxiv.org/abs/1707.01476.
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. QLoRA: efficient finetuning of quantized LLMs. In Advances in Neural Information Processing Systems 36 (NeurIPS). 2023. URL: https://arxiv.org/abs/2305.14314, arXiv:2305.14314.
Tim Dettmers and Luke Zettlemoyer. The case for 4-bit precision: k-bit inference scaling laws. In Proceedings of the 40th International Conference on Machine Learning (ICML). 2023. URL: https://arxiv.org/abs/2212.09720, arXiv:2212.09720.
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL-HLT, 4171–4186. 2019.
Ronald A. DeVore, Ralph Howard, and Charles Micchelli. Optimal nonlinear approximation. Manuscripta Mathematica, 63(4):469–478, 1989. doi:10.1007/BF01171759.
Terrance DeVries and Graham W. Taylor. Improved regularization of convolutional neural networks with cutout. arXiv preprint arXiv:1708.04552, 2017.
Prafulla Dhariwal, Heewoo Jun, Christine Payne, Jong Wook Kim, Alec Radford, and Ilya Sutskever. Jukebox: a generative model for music. arXiv preprint arXiv:2005.00341, 2020.
Prafulla Dhariwal and Alexander Nichol. Diffusion models beat GANs on image synthesis. In Advances in Neural Information Processing Systems, volume 34. 2021.
Amit Dhurandhar, Pin-Yu Chen, Ronny Luss, Chun-Chen Tu, Paishun Ting, Karthikeyan Shanmugam, and Payel Das. Explanations based on the missing: towards contrastive explanations with pertinent negatives. In Advances in Neural Information Processing Systems, volume 31. 2018.
Thomas J. DiCiccio and Bradley Efron. Bootstrap confidence intervals. Statistical Science, 11(3):189–228, 1996.
David A. Dickey and Wayne A. Fuller. Distribution of the estimators for autoregressive time series with a unit root. Journal of the American Statistical Association, 74(366):427–431, 1979.
William Dieterich, Christina Mendoza, and Tim Brennan. COMPAS risk scales: demonstrating accuracy equity and predictive parity. Technical Report, Northpointe Inc. Research Department, July 2016. Rapporto tecnico dell'8 luglio 2016, risposta all'inchiesta di ProPublica. URL: https://www.documentcloud.org/documents/2998391-ProPublica-Commentary-Final-070616/.
Franz Dietrich and Kai Spiekermann. Independent opinions? on the causal foundations of belief formation and jury theorems. Mind, 122(487):655–685, 2013. doi:10.1093/mind/fzt074.
Laurent Dinh, David Krueger, and Yoshua Bengio. NICE: non-linear independent components estimation. In International Conference on Learning Representations (ICLR), Workshop Track. 2015. arXiv:1410.8516.
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio. Sharp minima can generalize for deep nets. In Proceedings of the 34th International Conference on Machine Learning (ICML). 2017. URL: https://arxiv.org/abs/1703.04933, arXiv:1703.04933.
Laurent Dinh, Jascha Sohl-Dickstein, and Samy Bengio. Density estimation using Real NVP. In International Conference on Learning Representations (ICLR). 2017.
M. W. M. G. Dissanayake and N. Phan-Thien. Neural-network-based approximations for solving partial differential equations. Communications in Numerical Methods in Engineering, 10(3):195–201, 1994. doi:10.1002/cnm.1640100303.
Anuj Diwan, Layne Berry, Eunsol Choi, David Harwath, and Kyle Mahowald. Why is Winoground hard? Investigating failures in visuolinguistic compositionality. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2022.
Pedro Domingos. A unified bias-variance decomposition and its applications. In Proceedings of the 17th International Conference on Machine Learning (ICML), 231–238. 2000.
Pedro Domingos and Michael Pazzani. On the optimality of the simple bayesian classifier under zero-one loss. Machine Learning, 29(2–3):103–130, 1997. doi:10.1023/A:1007413511361.
Yihe Dong, Jean-Baptiste Cordonnier, and Andreas Loukas. Attention is not all you need: pure attention loses rank doubly exponentially with depth. In International Conference on Machine Learning (ICML). 2021. URL: https://arxiv.org/abs/2103.03404.
Yixin Dong, Charlie F. Ruan, Yaxing Cai, Ruihang Lai, Ziyi Xu, Yilong Zhao, and Tianqi Chen. XGrammar: flexible and efficient structured generation engine for large language models. In Proceedings of the 8th Conference on Machine Learning and Systems (MLSys). 2025. URL: https://arxiv.org/abs/2411.15100.
Marco Dorigo, Vittorio Maniezzo, and Alberto Colorni. Ant system: optimization by a colony of cooperating agents. IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics), 26(1):29–41, 1996.
Finale Doshi-Velez and Been Kim. Towards a rigorous science of interpretable machine learning. arXiv preprint arXiv:1702.08608, 2017. URL: https://arxiv.org/abs/1702.08608.
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: transformers for image recognition at scale. In International Conference on Learning Representations (ICLR). 2021.
D. C. Dowson and B. V. Landau. The Fréchet distance between multivariate normal distributions. Journal of Multivariate Analysis, 12(3):450–455, 1982. doi:10.1016/0047-259X(82)90077-X.
Timothy Dozat and Christopher D. Manning. Deep biaffine attention for neural dependency parsing. In International Conference on Learning Representations (ICLR). 2017.
Simon S. Du, Chi Jin, Jason D. Lee, Michael I. Jordan, Aarti Singh, and Barnabás Póczos. Gradient descent can take exponential time to escape saddle points. In Advances in Neural Information Processing Systems (NeurIPS). 2017.
Yilun Du, Conor Durkan, Robin Strudel, Joshua B. Tenenbaum, Sander Dieleman, Rob Fergus, Jascha Sohl-Dickstein, Arnaud Doucet, and Will Grathwohl. Reduce, reuse, recycle: compositional generation with energy-based diffusion models and mcmc. In International Conference on Machine Learning (ICML). 2023. URL: https://arxiv.org/abs/2302.11552.
Yilun Du and Igor Mordatch. Implicit generation and modeling with energy based models. In Advances in Neural Information Processing Systems (NeurIPS), volume 32. 2019. URL: https://proceedings.neurips.cc/paper/2019/hash/378a063b8fdb1db941e34f4bde584c7d-Abstract.html.
Rachit Dubey, Pulkit Agrawal, Deepak Pathak, Thomas L. Griffiths, and Alexei A. Efros. Investigating human priors for playing video games. In Proceedings of the 35th International Conference on Machine Learning (ICML), volume 80 of Proceedings of Machine Learning Research, 1349–1357. PMLR, 2018. URL: https://arxiv.org/abs/1802.10217.
John Duchi, Elad Hazan, and Yoram Singer. Adaptive subgradient methods for online learning and stochastic optimization. Journal of Machine Learning Research, 12:2121–2159, 2011.
Richard M. Dudley. Central limit theorems for empirical measures. The Annals of Probability, 6(6):899–929, 1978.
Philipp Dufter and Hinrich Schütze. Identifying elements essential for BERT's multilinguality. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 4423–4437. 2020.
Olive Jean Dunn. Multiple comparisons among means. Journal of the American Statistical Association, 56(293):52–64, 1961.
Conor Durkan, Artur Bekasov, Iain Murray, and George Papamakarios. Neural spline flows. In Advances in Neural Information Processing Systems (NeurIPS), volume 32. 2019.
David Duvenaud, Dougal Maclaurin, Jorge Aguilera-Iparraguirre, Rafael Gómez-Bombarelli, Timothy Hirzel, Alán Aspuru-Guzik, and Ryan P. Adams. Convolutional networks on graphs for learning molecular fingerprints. In Advances in Neural Information Processing Systems (NeurIPS), volume 28. 2015. URL: https://arxiv.org/abs/1509.09292.
Vijay Prakash Dwivedi and Xavier Bresson. A generalization of transformer networks to graphs. AAAI Workshop on Deep Learning on Graphs, 2021. URL: https://arxiv.org/abs/2012.09699.
Vijay Prakash Dwivedi, Chaitanya K. Joshi, Anh Tuan Luu, Thomas Laurent, Yoshua Bengio, and Xavier Bresson. Benchmarking graph neural networks. arXiv preprint arXiv:2003.00982, 2020. URL: https://arxiv.org/abs/2003.00982.
Vijay Prakash Dwivedi, Anh Tuan Luu, Thomas Laurent, Yoshua Bengio, and Xavier Bresson. Graph neural networks with learnable structural and positional representations. In International Conference on Learning Representations. 2022. arXiv:2110.07875.
Cynthia Dwork, Moritz Hardt, Toniann Pitassi, Omer Reingold, and Richard Zemel. Fairness through awareness. In Proceedings of the 3rd Innovations in Theoretical Computer Science Conference (ITCS), 214–226. 2012. URL: https://arxiv.org/abs/1104.3913, doi:10.1145/2090236.2090255.
Cynthia Dwork, Nancy Lynch, and Larry Stockmeyer. Consensus in the presence of partial synchrony. Journal of the ACM, 35(2):288–323, 1988. doi:10.1145/42282.42283.
Cynthia Dwork, Frank McSherry, Kobbi Nissim, and Adam Smith. Calibrating noise to sensitivity in private data analysis. In Theory of Cryptography Conference (TCC), 265–284. 2006.
Cynthia Dwork and Aaron Roth. The algorithmic foundations of differential privacy. Foundations and Trends in Theoretical Computer Science, 9(3–4):211–407, 2014. URL: https://doi.org/10.1561/0400000042, doi:10.1561/0400000042.
Gintare Karolina Dziugaite and Daniel M. Roy. Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data. In Proceedings of the 33rd Conference on Uncertainty in Artificial Intelligence (UAI). 2017. URL: https://arxiv.org/abs/1703.11008.
Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi. High fidelity neural audio compression. Transactions on Machine Learning Research (TMLR), 2023.
Weinan E and Bing Yu. The deep Ritz method: a deep learning-based numerical algorithm for solving variational problems. Communications in Mathematics and Statistics, 6(1):1–12, 2018. URL: https://arxiv.org/abs/1710.00211.
Brown Ebouky, Andrea Bartezzaghi, and Mattia Rigotti. Eliciting reasoning in language models with cognitive tools. arXiv preprint arXiv:2506.12115, 2025.
Carl Eckart and Gale Young. The approximation of one matrix by another of lower rank. Psychometrika, 1(3):211–218, 1936. doi:10.1007/BF02288367.
Adrien Ecoffet, Joost Huizinga, Joel Lehman, Kenneth O. Stanley, and Jeff Clune. First return, then explore. Nature, 590:580–586, 2021. URL: https://arxiv.org/abs/2004.12919, doi:10.1038/s41586-020-03157-9.
Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody, Steven Truitt, Dasha Metropolitansky, Robert Osazuwa Ness, and Jonathan Larson. From local to global: a graph RAG approach to query-focused summarization. 2024. arXiv:2404.16130.
Bradley Efron. Bootstrap methods: another look at the jackknife. The Annals of Statistics, 7(1):1–26, 1979. doi:10.1214/aos/1176344552.
Bradley Efron. The Jackknife, the Bootstrap and Other Resampling Plans. Number 38 in CBMS-NSF Regional Conference Series in Applied Mathematics. SIAM, Philadelphia, 1982.
Bradley Efron. Better bootstrap confidence intervals. Journal of the American Statistical Association, 82(397):171–185, 1987.
David Eigen, Marc'Aurelio Ranzato, and Ilya Sutskever. Learning factored representations in a deep mixture of experts. arXiv preprint arXiv:1312.4314, 2013.
Albert Einstein. Über die von der molekularkinetischen Theorie der Wärme geforderte Bewegung von in ruhenden Flüssigkeiten suspendierten Teilchen. Annalen der Physik, 17:549–560, 1905.
Yanai Elazar, Shauli Ravfogel, Alon Jacovi, and Yoav Goldberg. Amnesic probing: behavioral explanation with amnesic counterfactuals. Transactions of the Association for Computational Linguistics, 9:160–175, 2021. doi:10.1162/tacl_a_00359.
Ronen Eldan and Ohad Shamir. The power of depth for feedforward neural networks. In 29th Annual Conference on Learning Theory (COLT). 2016. URL: https://arxiv.org/abs/1512.03965, arXiv:1512.03965.
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah. Toy models of superposition. Transformer Circuits Thread, Anthropic, 2022. URL: https://transformer-circuits.pub/2022/toy_model/index.html.
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. A mathematical framework for transformer circuits. Transformer Circuits Thread, 2021. URL: https://transformer-circuits.pub/2021/framework/index.html.
Charles Elkan. The foundations of cost-sensitive learning. In Proceedings of the 17th International Joint Conference on Artificial Intelligence (IJCAI), 973–978. 2001.
Jeffrey L. Elman. Finding structure in time. Cognitive Science, 14(2):179–211, 1990.
Cooper Elsworth, Keguo Huang, David Patterson, Ian Schneider, Robert Sedivy, Savannah Goodman, Ben Townsend, Parthasarathy Ranganathan, Jeff Dean, Amin Vahdat, Ben Gomes, and James Manyika. Measuring the environmental impact of delivering AI at Google scale. 2025. arXiv:2508.15734, doi:10.48550/arXiv.2508.15734.
Scott Emmons, Benjamin Eysenbach, Ilya Kostrikov, and Sergey Levine. RvS: what is essential for offline RL via supervised learning? In International Conference on Learning Representations (ICLR). 2022. URL: https://arxiv.org/abs/2112.10751.
Bernd Engelmann, Evelyn Hayden, and Dirk Tasche. Measuring the discriminative power of rating systems. Discussion Paper, Series 2: Banking and Financial Supervision 01/2003, Deutsche Bundesbank, 2003.
Logan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Firdaus Janoos, Larry Rudolph, and Aleksander Madry. Implementation matters in deep policy gradients: a case study on PPO and TRPO. In International Conference on Learning Representations (ICLR). 2020. URL: https://arxiv.org/abs/2005.12729, arXiv:2005.12729.
Danielle Ensign, Sorelle A. Friedler, Scott Neville, Carlos Scheidegger, and Suresh Venkatasubramanian. Runaway feedback loops in predictive policing. In Sorelle A. Friedler and Christo Wilson, editors, Proceedings of the 1st Conference on Fairness, Accountability and Transparency, volume 81 of Proceedings of Machine Learning Research, 160–171. PMLR, 2018. URL: https://proceedings.mlr.press/v81/ensign18a.html.
Dumitru Erhan, Yoshua Bengio, Aaron Courville, Pierre-Antoine Manzagol, Pascal Vincent, and Samy Bengio. Why does unsupervised pre-training help deep learning? Journal of Machine Learning Research, 11:625–660, 2010.
Lee D. Erman, Frederick Hayes-Roth, Victor R. Lesser, and D. Raj Reddy. The Hearsay-II speech-understanding system: integrating knowledge to resolve uncertainty. ACM Computing Surveys, 12(2):213–253, 1980.
Shahul Es, Jithin James, Luis Espinosa-Anke, and Steven Schockaert. RAGAs: automated evaluation of retrieval augmented generation. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (EACL): System Demonstrations, 150–158. 2024.
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, Kyle Lacey, Alex Goodwin, Yannik Marek, and Robin Rombach. Scaling rectified flow transformers for high-resolution image synthesis. In Proceedings of the 41st International Conference on Machine Learning (ICML). 2024. URL: https://arxiv.org/abs/2403.03206.
Patrick Esser, Robin Rombach, and Björn Ommer. Taming transformers for high-resolution image synthesis. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2021.
Martin Ester, Hans-Peter Kriegel, Jörg Sander, and Xiaowei Xu. A density-based algorithm for discovering clusters in large spatial databases with noise. In Proceedings of the Second International Conference on Knowledge Discovery and Data Mining (KDD), 226–231. 1996.
Richard Evans and Jim Gao. Deepmind ai reduces google data centre cooling bill by 40%. DeepMind blog, 7 2016. URL: https://deepmind.google/discover/blog/deepmind-ai-reduces-google-data-centre-cooling-bill-by-40/.
Zach Evans, CJ Carr, Josiah Taylor, Scott H. Hawley, and Jordi Pons. Fast timing-conditioned latent audio diffusion. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, 12652–12665. PMLR, 2024.
Utku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro, and Erich Elsen. Rigging the lottery: making all tickets winners. In Proceedings of the 37th International Conference on Machine Learning (ICML), volume 119 of Proceedings of Machine Learning Research, 2943–2952. 2020.
Mark Everingham, Luc Van Gool, Christopher K. I. Williams, John Winn, and Andrew Zisserman. The PASCAL visual object classes (VOC) challenge. International Journal of Computer Vision, 88(2):303–338, 2010.
Benjamin Eysenbach, Abhishek Gupta, Julian Ibarz, and Sergey Levine. Diversity is all you need: learning skills without a reward function. In International Conference on Learning Representations (ICLR). 2019. URL: https://arxiv.org/abs/1802.06070.
Scott E. Fahlman, Geoffrey E. Hinton, and Terrence J. Sejnowski. Massively parallel architectures for AI: NETL, Thistle, and Boltzmann machines. In Proceedings of the National Conference on Artificial Intelligence (AAAI-83), 109–113. Washington, DC, 1983. AAAI Press.
Jianqing Fan. Design-adaptive nonparametric regression. Journal of the American Statistical Association, 87(420):998–1004, 1992.
Manuel Faysse, Hugues Sibille, Tony Wu, Bilel Omrani, Gautier Viaud, Céline Hudelot, and Pierre Colombo. ColPali: efficient document retrieval with vision language models. In International Conference on Learning Representations (ICLR). 2025.
William Fedus, Barret Zoph, and Noam Shazeer. Switch Transformers: scaling to trillion parameter models with simple and efficient sparsity. Journal of Machine Learning Research, 23(120):1–39, 2022.
Alvan R. Feinstein and Domenic V. Cicchetti. High agreement but low kappa: i. the problems of two paradoxes. Journal of Clinical Epidemiology, 43(6):543–549, 1990. doi:10.1016/0895-4356(90)90158-L.
William Feller. On the theory of stochastic processes, with particular reference to applications. In Proceedings of the First Berkeley Symposium on Mathematical Statistics and Probability, 403–432. University of California Press, 1949.
Fangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan, and Wei Wang. Language-agnostic BERT sentence embedding. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, 878–891. 2022.
Guhao Feng, Bohang Zhang, Yuntian Gu, Haotian Ye, Di He, and Liwei Wang. Towards revealing the mystery behind chain of thought: a theoretical perspective. In Advances in Neural Information Processing Systems (NeurIPS), volume 36. 2023. URL: https://arxiv.org/abs/2305.15408.
Maurizio Ferrari Dacrema, Paolo Cremonesi, and Dietmar Jannach. Are we really making much progress? a worrying analysis of recent neural recommendation approaches. In Proceedings of the 13th ACM Conference on Recommender Systems (RecSys '19), 101–109. 2019. Best Long Paper; arXiv:1907.06902. doi:10.1145/3298689.3347058.
Richard P. Feynman. Space-time approach to non-relativistic quantum mechanics. Reviews of Modern Physics, 20(2):367–387, 1948.
Richard E. Fikes and Nils J. Nilsson. STRIPS: a new approach to the application of theorem proving to problem solving. Artificial Intelligence, 2(3–4):189–208, 1971. doi:10.1016/0004-3702(71)90010-5.
Tim Finin, Richard Fritzson, Don McKay, and Robin McEntire. KQML as an agent communication language. In Proceedings of the Third International Conference on Information and Knowledge Management (CIKM '94), 456–463. ACM, 1994. doi:10.1145/191246.191322.
Arlington M. Fink. Equilibrium in a stochastic n-person game. Journal of Science of the Hiroshima University, Series A-I (Mathematics), 28(1):89–93, 1964. doi:10.32917/hmj/1206139508.
Chelsea Finn, Pieter Abbeel, and Sergey Levine. Model-agnostic meta-learning for fast adaptation of deep networks. In Proceedings of the 34th International Conference on Machine Learning (ICML). 2017. URL: https://arxiv.org/abs/1703.03400.
John R. Firth. A synopsis of linguistic theory, 1930–1955. In Studies in Linguistic Analysis, pages 1–32. Blackwell, Oxford, 1957.
Michael J. Fischer, Nancy A. Lynch, and Michael S. Paterson. Impossibility of distributed consensus with one faulty process. Journal of the ACM, 32(2):374–382, April 1985. doi:10.1145/3149.214121.
Martin A. Fischler and Robert C. Bolles. Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography. Communications of the ACM, 24(6):381–395, 1981.
Aaron Fisher, Cynthia Rudin, and Francesca Dominici. All models are wrong, but many are useful: learning a variable's importance by studying an entire class of prediction models simultaneously. Journal of Machine Learning Research, 20(177):1–81, 2019. URL: https://arxiv.org/abs/1801.01489.
Ronald A. Fisher. The use of multiple measurements in taxonomic problems. Annals of Eugenics, 7(2):179–188, 1936. doi:10.1111/j.1469-1809.1936.tb02137.x.
Seth Flaxman, Sharad Goel, and Justin M. Rao. Filter bubbles, echo chambers, and online news consumption. Public Opinion Quarterly, 80(S1):298–320, 2016. doi:10.1093/poq/nfw006.
Jakob N. Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson. Counterfactual multi-agent policy gradients. In Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, 2974–2982. 2018.
Riccardo Fogliato, Max G'Sell, and Alexandra Chouldechova. Fairness evaluation in presence of biased noisy labels. In Proceedings of the Twenty Third International Conference on Artificial Intelligence and Statistics (AISTATS), volume 108 of Proceedings of Machine Learning Research, 2325–2336. PMLR, 2020. URL: https://proceedings.mlr.press/v108/fogliato20a.html, arXiv:2003.13808.
Daniel Ford. Introducing contextual retrieval. Anthropic, September 2024. Pubblicato il 19 settembre 2024 (prima all'indirizzo https://www.anthropic.com/news/contextual-retrieval). Consultato il 2 ottobre 2026. URL: https://www.anthropic.com/engineering/contextual-retrieval.
Edward B. Fowlkes and Colin L. Mallows. A method for comparing two hierarchical clusterings. Journal of the American Statistical Association, 78(383):553–569, 1983. doi:10.1080/01621459.1983.10478008.
Jonathan Frankle and Michael Carbin. The lottery ticket hypothesis: finding sparse, trainable neural networks. In International Conference on Learning Representations (ICLR). 2019. arXiv:1803.03635.
Jonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, and Michael Carbin. Linear mode connectivity and the lottery ticket hypothesis. In International Conference on Machine Learning (ICML). 2020. URL: https://arxiv.org/abs/1912.05671.
Jonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, and Michael Carbin. Pruning neural networks at initialization: why are we missing the mark? In International Conference on Learning Representations (ICLR). 2021.
Elias Frantar and Dan Alistarh. SparseGPT: massive language models can be accurately pruned in one-shot. In Proceedings of the 40th International Conference on Machine Learning (ICML), 10323–10337. 2023.
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. GPTQ: accurate post-training quantization for generative pre-trained transformers. In International Conference on Learning Representations (ICLR). 2023.
Torkel Franzén. Gödel's Theorem: An Incomplete Guide to Its Use and Abuse. A K Peters, 2005.
Yoav Freund and Robert E. Schapire. A decision-theoretic generalization of on-line learning and an application to boosting. Journal of Computer and System Sciences, 55(1):119–139, 1997.
Sara Fridovich-Keil, Alex Yu, Matthew Tancik, Qinhong Chen, Benjamin Recht, and Angjoo Kanazawa. Plenoxels: radiance fields without neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 5491–5500. 2022. doi:10.1109/CVPR52688.2022.00542.
Batya Friedman and Helen Nissenbaum. Bias in computer systems. ACM Transactions on Information Systems, 14(3):330–347, July 1996. doi:10.1145/230538.230561.
Jerome Friedman, Trevor Hastie, and Robert Tibshirani. Additive logistic regression: a statistical view of boosting. The Annals of Statistics, 28(2):337–407, 2000.
Jerome H. Friedman. Regularized discriminant analysis. Journal of the American Statistical Association, 84(405):165–175, 1989. doi:10.1080/01621459.1989.10478752.
Jerome H. Friedman. Greedy function approximation: a gradient boosting machine. The Annals of Statistics, 29(5):1189–1232, 2001.
Jerome H. Friedman and Bogdan E. Popescu. Predictive learning via rule ensembles. The Annals of Applied Statistics, 2(3):916–954, 2008. doi:10.1214/07-AOAS148.
Karl Friston, Christopher Thornton, and Andy Clark. Free-energy minimization and the dark-room problem. Frontiers in Psychology, 3:130, 2012. doi:10.3389/fpsyg.2012.00130.
Daniel Y. Fu, Tri Dao, Khaled K. Saab, Armin W. Thomas, Atri Rudra, and Christopher Ré. Hungry hungry hippos: towards language modeling with state space models. In International Conference on Learning Representations (ICLR). 2023.
Yao Fu, Leyang Xue, Yeqi Huang, Andrei-Octavian Brabete, Dmitrii Ustiugov, Yuvraj Patel, and Luo Mai. ServerlessLLM: low-latency serverless inference for large language models. In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI), 135–153. 2024.
Scott Fujimoto, David Meger, and Doina Precup. Off-policy deep reinforcement learning without exploration. In International Conference on Machine Learning (ICML). 2019. URL: https://arxiv.org/abs/1812.02900.
Scott Fujimoto, Herke van Hoof, and David Meger. Addressing function approximation error in actor-critic methods. In International Conference on Machine Learning (ICML). 2018. URL: https://arxiv.org/abs/1802.09477.
Kunihiko Fukushima. Visual feature extraction by a multilayered network of analog threshold elements. IEEE Transactions on Systems Science and Cybernetics, 5(4):322–333, 1969. doi:10.1109/TSSC.1969.300225.
Kunihiko Fukushima. Neocognitron: a self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position. Biological Cybernetics, 36:193–202, 1980.
Simon Funk. Netflix update: try this at home. Blog post, December 2006. URL: https://sifter.org/~simon/journal/20061211.html.
Philip Gage. A new algorithm for data compression. The C Users Journal, 12(2):23–38, 1994.
Trevor Gale, Deepak Narayanan, Cliff Young, and Matei Zaharia. MegaBlocks: efficient sparse training with mixture-of-experts. In Proceedings of Machine Learning and Systems (MLSys), volume 5, 288–304. 2023. URL: https://proceedings.mlsys.org/paper_files/paper/2023/hash/5a54f79333768effe7e8927bcccffe40-Abstract-mlsys2023.html.
Francis Galton. Regression towards mediocrity in hereditary stature. The Journal of the Anthropological Institute of Great Britain and Ireland, 15:246–263, 1886. doi:10.2307/2841583.
João Gama, Indrė Žliobaitė, Albert Bifet, Mykola Pechenizkiy, and Abdelhamid Bouchachia. A survey on concept drift adaptation. ACM Computing Surveys, 46(4):44:1–44:37, 2014.
Leo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, and Jeffrey Wu. Scaling and evaluating sparse autoencoders. 2024. arXiv:2406.04093.
Leo Gao, John Schulman, and Jacob Hilton. Scaling laws for reward model overoptimization. In Proceedings of the 40th International Conference on Machine Learning (ICML), volume 202 of Proceedings of Machine Learning Research, 10835–10866. PMLR, 2023. URL: https://proceedings.mlr.press/v202/gao23h.html, arXiv:2210.10760.
Luyu Gao, Xueguang Ma, Jimmy Lin, and Jamie Callan. Precise zero-shot dense retrieval without relevance labels. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL), 1762–1777. 2023.
Everette S. Gardner, Jr. and Ed McKenzie. Forecasting trends in time series. Management Science, 31(10):1237–1246, 1985. doi:10.1287/mnsc.31.10.1237.
Aurélien Garivier and Olivier Cappé. The KL-UCB algorithm for bounded stochastic bandits and beyond. In Proceedings of the 24th Annual Conference on Learning Theory, volume 19 of Proceedings of Machine Learning Research, 359–376. PMLR, 2011.
Aurélien Garivier and Eric Moulines. On upper-confidence bound policies for switching bandit problems. In Algorithmic Learning Theory (ALT 2011), volume 6925 of Lecture Notes in Computer Science, 174–188. Springer, 2011. doi:10.1007/978-3-642-24412-4_16.
Quentin Garrido, Randall Balestriero, Laurent Najman, and Yann LeCun. RankMe: assessing the downstream performance of pretrained self-supervised representations by their rank. In International Conference on Machine Learning (ICML). 2023. URL: https://arxiv.org/abs/2210.02885.
Leon A. Gatys, Alexander S. Ecker, and Matthias Bethge. Texture synthesis using convolutional neural networks. In Advances in Neural Information Processing Systems, volume 28. Curran Associates, Inc., 2015. URL: https://proceedings.neurips.cc/paper_files/paper/2015/file/a5e00132373a7031000fd987a3c9f87b-Paper.pdf.
Leon A. Gatys, Alexander S. Ecker, and Matthias Bethge. Image style transfer using convolutional neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2414–2423. 2016.
Carl Friedrich Gauss. Summarische Übersicht der zur bestimmung der bahnen der beiden neuen hauptplaneten angewandten methoden. In Werke, Band VI, pages 148–165. Königliche Gesellschaft der Wissenschaften zu Göttingen, 1874.
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford. Datasheets for datasets. Communications of the ACM, 64(12):86–92, 2021. doi:10.1145/3458723.
Jonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer, Siddharth Singh, Brian R. Bartoldson, Bhavya Kailkhura, Abhinav Bhatele, and Tom Goldstein. Scaling up test-time compute with latent reasoning: a recurrent depth approach. In Advances in Neural Information Processing Systems, volume 38, 41340–41391. 2025. doi:10.52202/085713-1380.
Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A. Wichmann. Shortcut learning in deep neural networks. Nature Machine Intelligence, 2(11):665–673, 2020. doi:10.1038/s42256-020-00257-z.
Zorik Gekhman, Gal Yona, Roee Aharoni, Matan Eyal, Amir Feder, Roi Reichart, and Jonathan Herzig. Does fine-tuning LLMs on new knowledge encourage hallucinations? In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP), 7765–7784. Miami, Florida, USA, 2024. Association for Computational Linguistics. URL: https://aclanthology.org/2024.emnlp-main.444/, doi:10.18653/v1/2024.emnlp-main.444.
Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter. Audio set: an ontology and human-labeled dataset for audio events. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 776–780. 2017.
Zhengyang Geng, Mingyang Deng, Xingjian Bai, J. Zico Kolter, and Kaiming He. Mean flows for one-step generative modeling. In Advances in Neural Information Processing Systems, volume 38. 2025. URL: https://arxiv.org/abs/2505.13447.
Mathieu Germain, Karol Gregor, Iain Murray, and Hugo Larochelle. MADE: masked autoencoder for distribution estimation. In Proceedings of the 32nd International Conference on Machine Learning (ICML), volume 37 of Proceedings of Machine Learning Research, 881–889. 2015.
Felix A. Gers, Jürgen Schmidhuber, and Fred Cummins. Learning to forget: continual prediction with LSTM. Neural Computation, 12(10):2451–2471, 2000. doi:10.1162/089976600300015015.
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. Transformer feed-forward layers are key-value memories. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP), 5484–5495. 2021.
Marjan Ghazvininejad, Omer Levy, Yinhan Liu, and Luke Zettlemoyer. Mask-predict: parallel decoding of conditional masked language models. In Conference on Empirical Methods in Natural Language Processing (EMNLP-IJCNLP). 2019. URL: https://arxiv.org/abs/1904.09324.
Amir Gholami, Zhewei Yao, Sehoon Kim, Coleman Hooper, Michael W. Mahoney, and Kurt Keutzer. AI and memory wall. IEEE Micro, 44(3):33–39, 2024. doi:10.1109/MM.2024.3373763.
Martin Giles. The GANfather: the man who's given machines the gift of imagination. MIT Technology Review, February 2018. URL: https://www.technologyreview.com/2018/02/21/145289/the-ganfather-the-man-whos-given-machines-the-gift-of-imagination/.
Nicolas Gillis and François Glineur. Low-rank matrix approximation with weights or missing data is NP-hard. SIAM Journal on Matrix Analysis and Applications, 32(4):1149–1165, 2011. doi:10.1137/110820361.
Justin Gilmer, Samuel S. Schoenholz, Patrick F. Riley, Oriol Vinyals, and George E. Dahl. Neural message passing for quantum chemistry. In Proceedings of the 34th International Conference on Machine Learning (ICML), 1263–1272. 2017. URL: https://arxiv.org/abs/1704.01212.
Corrado Gini. Variabilità e mutabilità: contributo allo studio delle distribuzioni e delle relazioni statistiche. Studi economico-giuridici della R. Università di Cagliari. Tipografia di Paolo Cuppini, Bologna, 1912.
Corrado Gini. Sulla misura della concentrazione e della variabilità dei caratteri. Atti del Reale Istituto Veneto di Scienze, Lettere ed Arti, 73:1203–1248, 1914.
Ross Girshick. Fast R-CNN. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 1440–1448. 2015. doi:10.1109/ICCV.2015.169.
Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2014.
Xavier Glorot and Yoshua Bengio. Understanding the difficulty of training deep feedforward neural networks. In Proceedings of the 13th International Conference on Artificial Intelligence and Statistics (AISTATS), volume 9 of Proceedings of Machine Learning Research, 249–256. 2010.
Xavier Glorot, Antoine Bordes, and Yoshua Bengio. Deep sparse rectifier neural networks. In Proceedings of the 14th International Conference on Artificial Intelligence and Statistics (AISTATS). 2011.
Tilmann Gneiting, Fadoua Balabdaoui, and Adrian E. Raftery. Probabilistic forecasts, calibration and sharpness. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 69(2):243–268, 2007. doi:10.1111/j.1467-9868.2007.00587.x.
Joseph A. Goguen and José Meseguer. Security policies and security models. In Proceedings of the 1982 IEEE Symposium on Security and Privacy, 11–20. Oakland, CA, USA, 1982. IEEE. doi:10.1109/SP.1982.10014.
David Goldberg, David Nichols, Brian M. Oki, and Douglas Terry. Using collaborative filtering to weave an information tapestry. Communications of the ACM, 35(12):61–70, 1992. doi:10.1145/138859.138867.
Alex Goldstein, Adam Kapelner, Justin Bleich, and Emil Pitkin. Peeking inside the black box: visualizing statistical learning with plots of individual conditional expectation. Journal of Computational and Graphical Statistics, 24(1):44–65, 2015. URL: https://doi.org/10.1080/10618600.2014.907095.
Eric Goles and Jorge Olivos. Periodic behaviour of generalized threshold functions. Discrete Mathematics, 30(2):187–189, 1980. doi:10.1016/0012-365X(80)90121-1.
Philippe Golle. Revisiting the uniqueness of simple demographics in the US population. In Proceedings of the 5th ACM Workshop on Privacy in Electronic Society (WPES '06), 77–80. ACM, 2006. doi:10.1145/1179601.1179615.
Gene H. Golub and Charles F. Van Loan. Matrix Computations. Johns Hopkins University Press, 4 edition, 2013.
Yuan Gong, Yu-An Chung, and James Glass. AST: audio spectrogram transformer. In Interspeech, 571–575. 2021.
Ian Goodfellow. NIPS 2016 tutorial: generative adversarial networks. arXiv:1701.00160, 2017. tutorial NIPS 2016.
Ian Goodfellow, Yoshua Bengio, and Aaron Courville. Deep learning. MIT Press, 2016.
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets. In Advances in Neural Information Processing Systems, volume 27. 2014.
Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy. Explaining and harnessing adversarial examples. In International Conference on Learning Representations (ICLR). 2015. URL: https://arxiv.org/abs/1412.6572.
Joshua T. Goodman. A bit of progress in language modeling: extended version. Technical Report MSR-TR-2001-72, Microsoft Research, 2001. URL: https://arxiv.org/abs/cs/0108005.
Neil J. Gordon, David J. Salmond, and Adrian F. M. Smith. Novel approach to nonlinear/non-Gaussian Bayesian state estimation. IEE Proceedings F (Radar and Signal Processing), 140(2):107–113, 1993. doi:10.1049/ip-f-2.1993.0015.
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He. Accurate, large minibatch SGD: training ImageNet in 1 hour. arXiv preprint arXiv:1706.02677, 2017.
Peter D. Grünwald and Paul M. B. Vitányi. Shannon information and kolmogorov complexity. arXiv preprint cs/0410002, 2004.
Clive W. J. Granger. Investigating causal relations by econometric models and cross-spectral methods. Econometrica, 37(3):424–438, 1969.
Will Grathwohl, Ricky T. Q. Chen, Jesse Bettencourt, Ilya Sutskever, and David Duvenaud. FFJORD: free-form continuous dynamics for scalable reversible generative models. In International Conference on Learning Representations (ICLR). 2019.
Will Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud, Mohammad Norouzi, and Kevin Swersky. Your classifier is secretly an energy based model and you should treat it like one. In International Conference on Learning Representations (ICLR). 2020. URL: https://arxiv.org/abs/1912.03263.
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, and others. The Llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024. URL: https://arxiv.org/abs/2407.21783.
Alex Graves. Sequence transduction with recurrent neural networks. In ICML 2012 Workshop on Representation Learning. 2012. arXiv:1211.3711.
Alex Graves. Adaptive computation time for recurrent neural networks. arXiv preprint arXiv:1603.08983, 2016. URL: https://arxiv.org/abs/1603.08983.
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber. Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks. In Proceedings of the 23rd International Conference on Machine Learning (ICML), 369–376. 2006.
Riccardo Grazzi, Julien Siems, Arber Zela, Jörg K. H. Franke, Frank Hutter, and Massimiliano Pontil. Unlocking state-tracking in linear RNNs through negative eigenvalues. In International Conference on Learning Representations (ICLR). 2025. URL: https://arxiv.org/abs/2411.12537, arXiv:2411.12537.
Klaus Greff, Rupesh K. Srivastava, Jan Koutník, Bas R. Steunebrink, and Jürgen Schmidhuber. LSTM: a search space odyssey. IEEE Transactions on Neural Networks and Learning Systems, 28(10):2222–2232, 2017.
Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz. Not what you've signed up for: compromising real-world LLM-integrated applications with indirect prompt injection. In Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security (AISec '23), 79–90. Association for Computing Machinery, 2023. URL: https://arxiv.org/abs/2302.12173, doi:10.1145/3605764.3623985.
Arthur Gretton, Karsten M. Borgwardt, Malte J. Rasch, Bernhard Schölkopf, and Alexander Smola. A kernel two-sample test. Journal of Machine Learning Research, 13:723–773, 2012.
Cobus Greyling. Loop-engineering. Repository GitHub, 2026. URL: cobusgreyling/loop-engineering.
H. Paul Grice. Logic and conversation. In Peter Cole and Jerry L. Morgan, editors, Syntax and Semantics, Vol. 3: Speech Acts, pages 41–58. Academic Press, New York, 1975.
D. Griffin and J. Lim. Signal estimation from modified short-time Fourier transform. IEEE Transactions on Acoustics, Speech, and Signal Processing, 32(2):236–243, April 1984. doi:10.1109/TASSP.1984.1164317.
Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre H. Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, Bilal Piot, Koray Kavukcuoglu, Rémi Munos, and Michal Valko. Bootstrap your own latent: a new approach to self-supervised learning. In Advances in Neural Information Processing Systems (NeurIPS), volume 33, 21271–21284. 2020. URL: https://arxiv.org/abs/2006.07733.
Léo Grinsztajn, Edouard Oyallon, and Gaël Varoquaux. Why do tree-based models still outperform deep learning on typical tabular data? In Advances in Neural Information Processing Systems 35 (NeurIPS 2022), Datasets and Benchmarks Track. 2022. arXiv:2207.08815.
Tamara G. Grossmann, Urszula Julia Komorowska, Jonas Latz, and Carola-Bibiane Schönlieb. Can physics-informed neural networks beat the finite element method? IMA Journal of Applied Mathematics, 89(1):143–174, 2024. doi:10.1093/imamat/hxae011.
Aditya Grover and Jure Leskovec. Node2vec: scalable feature learning for networks. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD), 855–864. 2016. URL: https://arxiv.org/abs/1607.00653.
Patrick M. Grundy. Mathematics and games. Eureka, 2:6–8, 1939. Ristampato in Eureka 27 (1964), pp. 9–11.
Peter Grünwald. A tutorial introduction to the minimum description length principle. 2004. Ripreso come capitoli 1-2 di Grünwald, Myung e Pitt (a cura di), Advances in Minimum Description Length, MIT Press, 2005; le pagine di quel volume non le ho riscontrate. URL: https://arxiv.org/abs/math/0406077, arXiv:math/0406077.
Albert Gu and Tri Dao. Mamba: linear-time sequence modeling with selective state spaces. In First Conference on Language Modeling (COLM). 2024. Outstanding Paper Award; preprint arXiv:2312.00752 (2023).
Albert Gu, Tri Dao, Stefano Ermon, Atri Rudra, and Christopher Ré. HiPPO: recurrent memory with optimal polynomial projections. In Advances in Neural Information Processing Systems (NeurIPS). 2020.
Albert Gu, Karan Goel, and Christopher Ré. Efficiently modeling long sequences with structured state spaces. In International Conference on Learning Representations (ICLR). 2022.
Albert Gu, Ankit Gupta, Karan Goel, and Christopher Ré. On the parameterization and initialization of diagonal state space models. In Advances in Neural Information Processing Systems (NeurIPS). 2022.
Albert Gu, Isys Johnson, Karan Goel, Khaled Saab, Tri Dao, Atri Rudra, and Christopher Ré. Combining recurrent, convolutional, and continuous-time models with linear state space layers. In Advances in Neural Information Processing Systems, volume 34, 572–585. 2021.
Chenchen Gu, Xiang Lisa Li, Rohith Kuditipudi, Percy Liang, and Tatsunori Hashimoto. Auditing prompt caching in language model APIs. In Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, 20477–20496. PMLR, 2025. URL: https://proceedings.mlr.press/v267/gu25b.html.
Jiatao Gu, Tianrong Chen, David Berthelot, Huangjie Zheng, Yuyang Wang, Ruixiang Zhang, Laurent Dinh, Miguel Angel Bautista, Josh Susskind, and Shuangfei Zhai. STARFlow: scaling latent normalizing flows for high-resolution image synthesis. arXiv preprint arXiv:2506.06276, 2025. URL: https://arxiv.org/abs/2506.06276.
Shixiang Gu, Timothy Lillicrap, Ilya Sutskever, and Sergey Levine. Continuous deep Q-learning with model-based acceleration. In Proceedings of the 33rd International Conference on Machine Learning (ICML), volume 48 of Proceedings of Machine Learning Research, 2829–2838. PMLR, 2016. URL: https://arxiv.org/abs/1603.00748.
Tianyu Gu, Brendan Dolan-Gavitt, and Siddharth Garg. BadNets: identifying vulnerabilities in the machine learning model supply chain. arXiv preprint arXiv:1708.06733, 2017.
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, and Ruoming Pang. Conformer: convolution-augmented transformer for speech recognition. In Proceedings of Interspeech, 5036–5040. 2020. doi:10.21437/Interspeech.2020-3015.
Ishaan Gulrajani, Faruk Ahmed, Martin Arjovsky, Vincent Dumoulin, and Aaron Courville. Improved training of Wasserstein GANs. In Advances in Neural Information Processing Systems, volume 30. 2017.
Chuan Guo, Tom Goldstein, Awni Hannun, and Laurens van der Maaten. Certified data removal from machine learning models. In Proceedings of the 37th International Conference on Machine Learning (ICML). 2020. arXiv:1911.03030.
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. On calibration of modern neural networks. In Proceedings of the 34th International Conference on Machine Learning (ICML). 2017. URL: https://arxiv.org/abs/1706.04599.
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, Qihao Zhu, Runxin Xu, Ruoyu Zhang, Shirong Ma, Xiao Bi, and others. DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature, 645(8081):633–638, 2025. Versione estesa: arXiv:2501.12948. doi:10.1038/s41586-025-09422-z.
Mandy Guo, Qinlan Shen, Yinfei Yang, Heming Ge, Daniel Cer, Gustavo Hernandez Abrego, Keith Stevens, Noah Constant, Yun-Hsuan Sung, Brian Strope, and Ray Kurzweil. Effective parallel corpus mining using bilingual sentence embeddings. In Proceedings of the Third Conference on Machine Translation: Research Papers, 165–176. Brussels, Belgium, 2018. Association for Computational Linguistics. doi:10.18653/v1/W18-6317.
Ankit Gupta, Albert Gu, and Jonathan Berant. Diagonal state spaces are as effective as structured state spaces. In Advances in Neural Information Processing Systems (NeurIPS). 2022.
Udit Gupta, Young Geun Kim, Sylvia Lee, Jordan Tse, Hsien-Hsin S. Lee, Gu-Yeon Wei, David Brooks, and Carole-Jean Wu. Chasing carbon: the elusive environmental footprint of computing. In 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA), 854–867. IEEE, 2021. doi:10.1109/HPCA51647.2021.00076.
John L. Gustafson. Reevaluating Amdahl's law. Communications of the ACM, 31(5):532–533, 1988. doi:10.1145/42411.42415.
Michael Gutmann and Aapo Hyvärinen. Noise-contrastive estimation: a new estimation principle for unnormalized statistical models. In Proceedings of the 13th International Conference on Artificial Intelligence and Statistics (AISTATS), volume 9. 2010. URL: https://proceedings.mlr.press/v9/gutmann10a.html.
Michael U. Gutmann and Aapo Hyvärinen. Noise-contrastive estimation of unnormalized statistical models, with applications to natural image statistics. Journal of Machine Learning Research, 13:307–361, 2012. URL: https://jmlr.org/papers/v13/gutmann12a.html.
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. REALM: retrieval-augmented language model pre-training. In Proceedings of the 37th International Conference on Machine Learning (ICML), volume 119 of Proceedings of Machine Learning Research, 3929–3938. 2020. URL: https://proceedings.mlr.press/v119/guu20a.html.
Kurt Gödel. Über formal unentscheidbare Sätze der Principia Mathematica und verwandter Systeme I. Monatshefte für Mathematik und Physik, 38:173–198, 1931.
Aurélien Géron. Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems. O'Reilly Media, third edition, 2022. ISBN 978-1-098-12597-4.
Aurélien Géron. Hands-On Machine Learning with Scikit-Learn and PyTorch. O'Reilly Media, first edition, 2025. ISBN 979-8-341-60797-2. URL: https://homl.info/.
Juan Gómez-Luna, Izzat El Hajj, Ivan Fernandez, Christina Giannoula, Geraldo F. Oliveira, and Onur Mutlu. Benchmarking a new paradigm: an experimental analysis of a real processing-in-memory architecture. arXiv preprint arXiv:2105.03814, 2021.
David Ha and Jürgen Schmidhuber. Recurrent world models facilitate policy evolution. In Advances in Neural Information Processing Systems 31 (NeurIPS). 2018. Circolato anche come \emph World Models, arXiv:1803.10122. URL: https://arxiv.org/abs/1803.10122.
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. Soft actor-critic: off-policy maximum entropy deep reinforcement learning with a stochastic actor. In International Conference on Machine Learning (ICML). 2018. URL: https://arxiv.org/abs/1801.01290.
Tuomas Haarnoja, Aurick Zhou, Kristian Hartikainen, George Tucker, Sehoon Ha, Jie Tan, Vikash Kumar, Henry Zhu, Abhishek Gupta, Pieter Abbeel, and Sergey Levine. Soft actor-critic algorithms and applications. 2018. URL: https://arxiv.org/abs/1812.05905, arXiv:1812.05905.
Jeremy Hadfield, Barry Zhang, Kenneth Lien, Florian Scholz, Jeremy Fox, and Daniel Ford. How we built our multi-agent research system. Anthropic Engineering, June 2025. Pubblicato il 13 giugno 2025. Consultato il 2 ottobre 2026. URL: https://www.anthropic.com/engineering/multi-agent-research-system.
Raia Hadsell, Sumit Chopra, and Yann LeCun. Dimensionality reduction by learning an invariant mapping. In IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR), volume 2, 1735–1742. 2006. doi:10.1109/CVPR.2006.100.
Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson. Learning latent dynamics for planning from pixels. In Proceedings of the 36th International Conference on Machine Learning (ICML), volume 97 of PMLR. 2019. Introduce PlaNet e il \emph recurrent state-space model (RSSM), arXiv:1811.04551. URL: https://arxiv.org/abs/1811.04551.
Danijar Hafner, Timothy Lillicrap, Mohammad Norouzi, and Jimmy Ba. Mastering atari with discrete world models. In International Conference on Learning Representations (ICLR). 2021.
Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap. Mastering diverse control tasks through world models. Nature, 640:647–653, 2025. Preprint 2023: \emph Mastering Diverse Domains through World Models, arXiv:2301.04104. doi:10.1038/s41586-025-08744-2.
Danijar Hafner, Wilson Yan, and Timothy Lillicrap. Training agents inside of scalable world models. arXiv preprint arXiv:2509.24527, 2025. URL: https://arxiv.org/abs/2509.24527.
William L. Hamilton. Graph Representation Learning. Synthesis Lectures on Artificial Intelligence and Machine Learning. Morgan & Claypool, 2020.
William L. Hamilton, Kevin Clark, Jure Leskovec, and Dan Jurafsky. Inducing domain-specific sentiment lexicons from unlabeled corpora. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing (EMNLP), 595–605. 2016. doi:10.18653/v1/D16-1057.
William L. Hamilton, Rex Ying, and Jure Leskovec. Inductive representation learning on large graphs. In Advances in Neural Information Processing Systems (NeurIPS), volume 30. 2017. URL: https://arxiv.org/abs/1706.02216.
David K. Hammond, Pierre Vandergheynst, and Rémi Gribonval. Wavelets on graphs via spectral graph theory. Applied and Computational Harmonic Analysis, 30(2):129–150, 2011.
Jiequn Han, Arnulf Jentzen, and Weinan E. Solving high-dimensional partial differential equations using deep learning. Proceedings of the National Academy of Sciences, 115(34):8505–8510, 2018. doi:10.1073/pnas.1718942115.
Song Han, Xingyu Liu, Huizi Mao, Jing Pu, Ardavan Pedram, Mark A. Horowitz, and William J. Dally. EIE: efficient inference engine on compressed deep neural network. In 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA), 243–254. 2016. doi:10.1109/ISCA.2016.30.
Song Han, Jeff Pool, John Tran, and William J. Dally. Learning both weights and connections for efficient neural networks. In Advances in Neural Information Processing Systems (NIPS), volume 28. 2015. URL: https://arxiv.org/abs/1506.02626.
James Hannan. Approximation to Bayes risk in repeated play. In Melvin Dresher, Albert W. Tucker, and Philip Wolfe, editors, Contributions to the Theory of Games, Volume III, number 39 in Annals of Mathematics Studies, pages 97–139. Princeton University Press, 1957.
Awni Y. Hannun, Pranav Rajpurkar, Masoumeh Haghpanahi, Geoffrey H. Tison, Codie Bourn, Mintu P. Turakhia, and Andrew Y. Ng. Cardiologist-level arrhythmia detection and classification in ambulatory electrocardiograms using a deep neural network. Nature Medicine, 25(1):65–69, 2019. doi:10.1038/s41591-018-0268-3.
Shibo Hao, Sainbayar Sukhbaatar, DiJia Su, Xian Li, Zhiting Hu, Jason Weston, and Yuandong Tian. Training large language models to reason in a continuous latent space. In Conference on Language Modeling (COLM). 2025. URL: https://openreview.net/forum?id=Itxz7S4Ip3.
Jean Harb, Pierre-Luc Bacon, Martin Klissarov, and Doina Precup. When waiting is not an option: learning options with a deliberation cost. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 32. 2018. URL: https://arxiv.org/abs/1709.04571, doi:10.1609/aaai.v32i1.11831.
Moritz Hardt, Eric Price, and Nathan Srebro. Equality of opportunity in supervised learning. In Advances in Neural Information Processing Systems (NeurIPS). 2016. URL: https://arxiv.org/abs/1610.02413.
F. Maxwell Harper and Joseph A. Konstan. The MovieLens datasets: history and context. ACM Transactions on Interactive Intelligent Systems, 5(4):1–19, 2015.
Chris Harris and Mike Stephens. A combined corner and edge detector. In Proceedings of the 4th Alvey Vision Conference, 147–151. 1988.
Peter E. Hart, Nils J. Nilsson, and Bertram Raphael. A formal basis for the heuristic determination of minimum cost paths. IEEE Transactions on Systems Science and Cybernetics, 4(2):100–107, 1968. doi:10.1109/TSSC.1968.300136.
Richard Hartley and Andrew Zisserman. Multiple View Geometry in Computer Vision. Cambridge University Press, second edition, 2004.
Trevor Hastie, Andreas Buja, and Robert Tibshirani. Penalized discriminant analysis. The Annals of Statistics, 23(1):73–102, 1995. doi:10.1214/aos/1176324456.
Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J. Tibshirani. Surprises in high-dimensional ridgeless least squares interpolation. The Annals of Statistics, 50(2):949–986, 2022.
Trevor Hastie and Robert Tibshirani. Generalized additive models. Statistical Science, 1(3):297–318, 1986. doi:10.1214/ss/1177013604.
Trevor Hastie, Robert Tibshirani, and Andreas Buja. Flexible discriminant analysis by optimal scoring. Journal of the American Statistical Association, 89(428):1255–1270, 1994. doi:10.1080/01621459.1994.10476866.
Trevor Hastie, Robert Tibshirani, and Jerome Friedman. The Elements of Statistical Learning: Data Mining, Inference, and Prediction. Springer Series in Statistics. Springer, second edition, 2009. doi:10.1007/978-0-387-84858-7.
Trevor Hastie, Robert Tibshirani, and Ryan Tibshirani. Best subset, forward stepwise or lasso? analysis and recommendations based on extensive comparisons. Statistical Science, 35(4):579–592, 2020.
Vasileios Hatzivassiloglou and Kathleen R. McKeown. Predicting the semantic orientation of adjectives. In 35th Annual Meeting of the Association for Computational Linguistics and 8th Conference of the European Chapter of the Association for Computational Linguistics, 174–181. 1997. doi:10.3115/976909.979640.
Taher H. Haveliwala and Sepandar D. Kamvar. The second eigenvalue of the Google matrix. Technical Report, Stanford University, 2003.
Adi Haviv, Ori Ram, Ofir Press, Peter Izsak, and Omer Levy. Transformer language models without positional encodings still learn positional information. In Findings of the Association for Computational Linguistics: EMNLP 2022. 2022. URL: https://arxiv.org/abs/2203.16634, arXiv:2203.16634.
Elad Hazan. Introduction to online convex optimization. Foundations and Trends in Optimization, 2(3-4):157–325, 2016. doi:10.1561/2400000013.
Elad Hazan, Amit Agarwal, and Satyen Kale. Logarithmic regret algorithms for online convex optimization. Machine Learning, 69(2-3):169–192, 2007.
Horace He. Defeating nondeterminism in LLM inference. Thinking Machines Lab, 10 settembre 2025, 2025. URL: https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/.
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 16000–16009. 2022. URL: https://arxiv.org/abs/2111.06377.
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 9729–9738. 2020. URL: https://arxiv.org/abs/1911.05722.
Kaiming He, Ross Girshick, and Piotr Dollár. Rethinking ImageNet pre-training. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 4917–4926. 2019. doi:10.1109/ICCV.2019.00502.
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. Mask R-CNN. In Proceedings of the IEEE International Conference on Computer Vision (ICCV). 2017.
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Delving deep into rectifiers: surpassing human-level performance on ImageNet classification. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 1026–1034. 2015. doi:10.1109/ICCV.2015.123.
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 770–778. 2016.
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Identity mappings in deep residual networks. In Computer Vision – ECCV 2016, volume 9908 of Lecture Notes in Computer Science, 630–645. Springer, 2016. arXiv:1603.05027.
Tong He, Zhi Zhang, Hang Zhang, Zhongyue Zhang, Junyuan Xie, and Mu Li. Bag of tricks for image classification with convolutional neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 558–567. 2019. doi:10.1109/CVPR.2019.00065.
Xiangnan He, Kuan Deng, Xiang Wang, Yan Li, Yongdong Zhang, and Meng Wang. Lightgcn: simplifying and powering graph convolution network for recommendation. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval, 639–648. 2020.
Xiangnan He, Lizi Liao, Hanwang Zhang, Liqiang Nie, Xia Hu, and Tat-Seng Chua. Neural collaborative filtering. In Proceedings of the 26th International Conference on World Wide Web, 173–182. 2017.
Yanzhang He, Tara N. Sainath, Rohit Prabhavalkar, Ian McGraw, Raziel Alvarez, Ding Zhao, David Rybach, Anjuli Kannan, Yonghui Wu, Ruoming Pang, Qiao Liang, Deepti Bhatia, Yuan Shangguan, Bo Li, Golan Pundak, Khe Chai Sim, Tom Bagby, Shuo-yiin Chang, Kanishka Rao, and Alexander Gruenstein. Streaming end-to-end speech recognition for mobile devices. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 2019. arXiv:1811.06621.
H. S. Heaps. Information Retrieval: Computational and Theoretical Aspects. Library and Information Science. Academic Press, New York, 1978.
Donald O. Hebb. The Organization of Behavior: A Neuropsychological Theory. Wiley, New York, 1949.
Matthew Henderson, Blaise Thomson, and Jason D. Williams. The second dialog state tracking challenge. In Proceedings of the 15th Annual Meeting of the Special Interest Group on Discourse and Dialogue (SIGDIAL), 263–272. Philadelphia, PA, USA, 2014. Association for Computational Linguistics. doi:10.3115/v1/W14-4337.
Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger. Deep reinforcement learning that matters. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 32, 3207–3214. 2018. URL: https://arxiv.org/abs/1709.06560, doi:10.1609/aaai.v32i1.11694.
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. In International Conference on Learning Representations (ICLR). 2021. URL: https://arxiv.org/abs/2009.03300, arXiv:2009.03300.
Dan Hendrycks and Kevin Gimpel. Gaussian error linear units (GELUs). arXiv preprint arXiv:1606.08415, 2016.
Dan Hendrycks, Norman Mu, Ekin D. Cubuk, Barret Zoph, Justin Gilmer, and Balaji Lakshminarayanan. AugMix: a simple data processing method to improve robustness and uncertainty. In International Conference on Learning Representations (ICLR). 2020. URL: https://openreview.net/forum?id=S1gmrxHFvB.
Mark Herbster and Manfred K. Warmuth. Tracking the best expert. Machine Learning, 32(2):151–178, 1998.
Gustav Herdan. Type-Token Mathematics: A Textbook of Mathematical Linguistics. Number 4 in Janua Linguarum, Series Maior. Mouton, 's-Gravenhage, 1960.
Jonathan L. Herlocker, Joseph A. Konstan, Al Borchers, and John Riedl. An algorithmic framework for performing collaborative filtering. In Proceedings of the 22nd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR '99), 230–237. ACM, 1999. doi:10.1145/312624.312682.
Matteo Hessel, Joseph Modayil, Hado van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Gheshlaghi Azar, and David Silver. Rainbow: combining improvements in deep reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence. 2018.
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local Nash equilibrium. In Advances in Neural Information Processing Systems, volume 30. 2017.
John Hewitt and Percy Liang. Designing and interpreting probes with control tasks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 2733–2743. 2019. URL: https://arxiv.org/abs/1909.03368.
John Hewitt and Christopher D. Manning. A structural probe for finding syntax in word representations. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL-HLT), 4129–4138. 2019.
Balázs Hidasi, Alexandros Karatzoglou, Linas Baltrunas, and Domonkos Tikk. Session-based recommendations with recurrent neural networks. In 4th International Conference on Learning Representations (ICLR 2016). 2016. arXiv:1511.06939.
Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner. Beta-VAE: learning basic visual concepts with a constrained variational framework. In International Conference on Learning Representations (ICLR). 2017.
Nicholas J. Higham. Accuracy and Stability of Numerical Algorithms. SIAM, 2 edition, 2002.
W. Daniel Hillis and Guy L. Steele, Jr. Data parallel algorithms. Communications of the ACM, 29(12):1170–1183, 1986. doi:10.1145/7902.7903.
Andrew Hines, Jan Skoglund, Anil C. Kokaram, and Naomi Harte. ViSQOL: an objective speech quality model. EURASIP Journal on Audio, Speech, and Music Processing, 2015(1):13, 2015. doi:10.1186/s13636-015-0054-9.
Keegan Hines, Gary Lopez, Matthew Hall, Federico Zarfati, Yonatan Zunger, and Emre Kiciman. Defending against indirect prompt injection attacks with spotlighting. 2024. URL: https://arxiv.org/abs/2403.14720, arXiv:2403.14720.
David V. Hinkley. Inference about the change-point from cumulative sum tests. Biometrika, 58(3):509–523, 1971.
Geoffrey Hinton, Li Deng, Dong Yu, George E. Dahl, Abdel-rahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Tara N. Sainath, and Brian Kingsbury. Deep neural networks for acoustic modeling in speech recognition: the shared views of four research groups. IEEE Signal Processing Magazine, 29(6):82–97, 2012. doi:10.1109/MSP.2012.2205597.
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2015. NIPS 2014 Deep Learning Workshop. URL: https://arxiv.org/abs/1503.02531.
Geoffrey E. Hinton. Training products of experts by minimizing contrastive divergence. Neural Computation, 14(8):1771–1800, 2002. URL: https://direct.mit.edu/neco/article/14/8/1771/6687, doi:10.1162/089976602760128018.
Geoffrey E. Hinton, Simon Osindero, and Yee-Whye Teh. A fast learning algorithm for deep belief nets. Neural Computation, 18(7):1527–1554, 2006. doi:10.1162/neco.2006.18.7.1527.
Geoffrey E. Hinton and Drew van Camp. Keeping the neural networks simple by minimizing the description length of the weights. In Proceedings of the Sixth Annual Conference on Computational Learning Theory (COLT), 5–13. 1993.
Jonathan Ho and Stefano Ermon. Generative adversarial imitation learning. In Advances in Neural Information Processing Systems 29 (NeurIPS). 2016. arXiv:1606.03476.
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems, volume 33, 6840–6851. 2020. URL: https://arxiv.org/abs/2006.11239.
Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598, 2022. Versione breve presentata al NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications. URL: https://arxiv.org/abs/2207.12598.
Tin Kam Ho. The random subspace method for constructing decision forests. IEEE Transactions on Pattern Analysis and Machine Intelligence, 20(8):832–844, 1998.
Sepp Hochreiter. Untersuchungen zu dynamischen neuronalen netzen. Master's thesis, Institut für Informatik, Technische Universität München, 1991. Relatore: Jürgen Schmidhuber, Lehrstuhl Prof. Brauer.
Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory. Neural Computation, 9(8):1735–1780, 1997.
Wassily Hoeffding. Probability inequalities for sums of bounded random variables. Journal of the American Statistical Association, 58(301):13–30, 1963.
Josef Hofbauer and Karl Sigmund. Evolutionary Games and Population Dynamics. Cambridge University Press, Cambridge, 1998. doi:10.1017/CBO9781139173179.
Matthew D. Hoffman and Matthew J. Johnson. ELBO surgery: yet another way to carve up the variational evidence lower bound. In NIPS Workshop on Advances in Approximate Bayesian Inference. 2016.
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, and others. Training compute-optimal large language models. In Advances in Neural Information Processing Systems 35 (NeurIPS). 2022.
Douglas R. Hofstadter. Gödel, Escher, Bach: An Eternal Golden Braid. Basic Books, 1979. Ed. italiana: \it Gödel, Escher, Bach: un'eterna ghirlanda brillante, Adelphi, 1984.
John H. Holland. Adaptation in Natural and Artificial Systems. University of Michigan Press, 1975.
Noah Hollmann, Samuel Müller, Lennart Purucker, Arjun Krishnakumar, Max Körfer, Shi Bin Hoo, Robin Tibor Schirrmeister, and Frank Hutter. Accurate predictions on small data with a tabular foundation model. Nature, 637:319–326, 2025.
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. The curious case of neural text degeneration. In International Conference on Learning Representations. 2020. URL: https://openreview.net/forum?id=rygGQyrFvH.
Jia-Wei Hong and H. T. Kung. I/o complexity: the red-blue pebble game. In Proceedings of the Thirteenth Annual ACM Symposium on Theory of Computing (STOC), 326–333. ACM, 1981. doi:10.1145/800076.802486.
Sanghyun Hong, Pietro Frigo, Yiğitcan Kaya, Cristiano Giuffrida, and Tudor Dumitraş. Terminal brain damage: exposing the graceless degradation in deep neural networks under hardware fault attacks. In 28th USENIX Security Symposium (USENIX Security 19), 497–514. 2019.
Giles Hooker, Lucas Mentch, and Siyu Zhou. Unrestricted permutation forces extrapolation: variable importance requires at least one more model, or there is no free variable importance. Statistics and Computing, 2021. URL: https://arxiv.org/abs/1905.03151.
Sara Hooker, Dumitru Erhan, Pieter-Jan Kindermans, and Been Kim. A benchmark for interpretability methods in deep neural networks. In Advances in Neural Information Processing Systems 32 (NeurIPS 2019). 2019. URL: https://arxiv.org/abs/1806.10758.
John J. Hopfield. Neural networks and physical systems with emergent collective computational abilities. Proceedings of the National Academy of Sciences, 79(8):2554–2558, 1982. URL: https://www.pnas.org/doi/10.1073/pnas.79.8.2554, doi:10.1073/pnas.79.8.2554.
Berthold K. P. Horn and Brian G. Schunck. Determining optical flow. Artificial Intelligence, 17(1–3):185–203, 1981.
Kurt Hornik. Approximation capabilities of multilayer feedforward networks. Neural Networks, 4(2):251–257, 1991.
Mark Horowitz. Computing's energy problem (and what we can do about it). In IEEE International Solid-State Circuits Conference (ISSCC), 10–14. 2014.
Harold Hotelling. Analysis of a complex of statistical variables into principal components. Journal of Educational Psychology, 24(6):417–441, 1933.
Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. MobileNets: efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861, 2017. URL: https://arxiv.org/abs/1704.04861.
Jeremy Howard and Sebastian Ruder. Universal language model fine-tuning for text classification. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (ACL), 328–339. 2018.
Ronald A. Howard. Dynamic Programming and Markov Processes. MIT Press, 1960.
Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, Yang Zhang, and Boris Ginsburg. RULER: what's the real context size of your long-context language models? In Conference on Language Modeling (COLM). 2024. URL: https://arxiv.org/abs/2404.06654.
Cheng-Yu Hsieh, Jieyu Zhang, Zixian Ma, Aniruddha Kembhavi, and Ranjay Krishna. SugarCrepe: fixing hackable benchmarks for vision-language compositionality. In Advances in Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track. 2023.
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. HuBERT: self-supervised speech representation learning by masked prediction of hidden units. IEEE/ACM Transactions on Audio, Speech, and Language Processing (TASLP), 29:3451–3460, 2021.
Anthony Hu, Lloyd Russell, Hudson Yeo, Zak Murez, George Fedoseev, Alex Kendall, Jamie Shotton, and Gianluca Corrado. GAIA-1: a generative world model for autonomous driving. arXiv preprint arXiv:2309.17080, 2023. URL: https://arxiv.org/abs/2309.17080.
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: low-rank adaptation of large language models. In International Conference on Learning Representations (ICLR). 2022. arXiv:2106.09685.
Yifan Hu, Yehuda Koren, and Chris Volinsky. Collaborative filtering for implicit feedback datasets. In Proceedings of the 2008 Eighth IEEE International Conference on Data Mining (ICDM '08), 263–272. 2008. doi:10.1109/ICDM.2008.22.
Zheyuan Hu, Khemraj Shukla, George Em Karniadakis, and Kenji Kawaguchi. Tackling the curse of dimensionality with physics-informed neural networks. Neural Networks, 176:106369, 2024. doi:10.1016/j.neunet.2024.106369.
Zhigang Hu, Alper Buyuktosunoglu, Viji Srinivasan, Victor Zyuban, Hans Jacobson, and Pradip Bose. Microarchitectural techniques for power gating of execution units. In Proceedings of the 2004 International Symposium on Low Power Electronics and Design (ISLPED), 32–37. 2004. doi:10.1109/LPE.2004.1349303.
Tianyu Hua, Wenxiao Wang, Zihui Xue, Sucheng Ren, Yue Wang, and Hang Zhao. On feature decorrelation in self-supervised learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 9578–9588. IEEE, 2021. arXiv:2105.00470. Pagine della versione IEEE; la versione CVF open access impagina 9598–9608. doi:10.1109/ICCV48922.2021.00946.
Weizhe Hua, Zihang Dai, Hanxiao Liu, and Quoc V. Le. Transformer quality in linear time. In Proceedings of the 39th International Conference on Machine Learning (ICML), 9099–9117. 2022.
Cheng-Zhi Anna Huang, Ashish Vaswani, Jakob Uszkoreit, Noam Shazeer, Ian Simon, Curtis Hawthorne, Andrew M. Dai, Matthew D. Hoffman, Monica Dinculescu, and Douglas Eck. Music transformer: generating music with long-term structure. In International Conference on Learning Representations (ICLR). 2019.
Di Huang, Xishan Zhang, Rui Zhang, Tian Zhi, Deyuan He, Jiaming Guo, Chang Liu, Qi Guo, Zidong Du, Shaoli Liu, Tianshi Chen, and Yunji Chen. DWM: a decomposable Winograd method for convolution acceleration. In Proceedings of the AAAI Conference on Artificial Intelligence. 2020. arXiv:2002.00552.
Gao Huang, Zhuang Liu, Laurens van der Maaten, and Kilian Q. Weinberger. Densely connected convolutional networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 4700–4708. 2017.
Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou. Large language models cannot self-correct reasoning yet. In International Conference on Learning Representations (ICLR). 2024.
Xun Huang and Serge Belongie. Arbitrary style transfer in real-time with adaptive instance normalization. In IEEE International Conference on Computer Vision (ICCV). 2017.
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Mia Xu Chen, Dehao Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le, Yonghui Wu, and Zhifeng Chen. GPipe: efficient training of giant neural networks using pipeline parallelism. In Advances in Neural Information Processing Systems (NeurIPS). 2019.
Yiwen Huang, Aaron Gokaslan, Volodymyr Kuleshov, and James Tompkin. The GAN is dead; long live the GAN! a modern GAN baseline. In Advances in Neural Information Processing Systems, volume 37, 44177–44215. 2024. URL: https://arxiv.org/abs/2501.05441, doi:10.52202/079017-1402.
Yu-Siang Huang and Yi-Hsuan Yang. Pop Music Transformer: beat-based modeling and generation of expressive pop piano compositions. In Proceedings of the 28th ACM International Conference on Multimedia (MM), 1180–1188. 2020.
Yuzhen Huang, Jinghan Zhang, Zifei Shan, and Junxian He. Compression represents intelligence linearly. In Conference on Language Modeling (COLM). 2024.
Zhiheng Huang, Wei Xu, and Kai Yu. Bidirectional LSTM-CRF models for sequence tagging. 2015. URL: https://arxiv.org/abs/1508.01991, arXiv:1508.01991.
David H. Hubel and Torsten N. Wiesel. Receptive fields of single neurones in the cat's striate cortex. The Journal of Physiology, 148(3):574–591, 1959.
David H. Hubel and Torsten N. Wiesel. Receptive fields, binocular interaction and functional architecture in the cat's visual cortex. The Journal of Physiology, 160(1):106–154, 1962. doi:10.1113/jphysiol.1962.sp006837.
Lawrence Hubert and Phipps Arabie. Comparing partitions. Journal of Classification, 2(1):193–218, 1985. doi:10.1007/BF01908075.
Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert, Meg Tong, Monte MacDiarmid, Tamera Lanham, Daniel M. Ziegler, Tim Maxwell, Newton Cheng, and others. Sleeper agents: training deceptive LLMs that persist through safety training. arXiv preprint arXiv:2401.05566, 2024. URL: https://arxiv.org/abs/2401.05566.
Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant. Risks from learned optimization in advanced machine learning systems. 2019. URL: https://arxiv.org/abs/1906.01820, arXiv:1906.01820.
David Hume. A Treatise of Human Nature. John Noon, London, 1739.
Ferenc Huszár. How (not) to train your generative model: scheduled sampling, likelihood, adversary? 2015. arXiv:1511.05101.
W. John Hutchins. The Georgetown-IBM experiment demonstrated in January 1954. In Robert E. Frederking and Kathryn B. Taylor, editors, Machine Translation: From Real Users to Research (AMTA 2004), volume 3265 of Lecture Notes in Computer Science, 102–114. Springer, 2004. doi:10.1007/978-3-540-30194-3_12.
M. F. Hutchinson. A stochastic estimator of the trace of the influence matrix for Laplacian smoothing splines. Communications in Statistics – Simulation and Computation, 18(3):1059–1076, 1989.
Chip Huyen. Designing Machine Learning Systems: An Iterative Process for Production-Ready Applications. O'Reilly Media, 2022.
Laurent Hyafil and Ronald L. Rivest. Constructing optimal binary decision trees is NP-complete. Information Processing Letters, 5(1):15–17, 1976.
Rob J. Hyndman and George Athanasopoulos. Forecasting: Principles and Practice. OTexts, third edition, 2021. URL: https://otexts.com/fpp3/.
Rob J. Hyndman and Anne B. Koehler. Another look at measures of forecast accuracy. International Journal of Forecasting, 22(4):679–688, 2006. doi:10.1016/j.ijforecast.2006.03.001.
Aapo Hyvärinen. Estimation of non-normalized statistical models by score matching. Journal of Machine Learning Research, 6:695–709, 2005. URL: https://jmlr.org/papers/v6/hyvarinen05a.html.
Aapo Hyvärinen and Erkki Oja. Independent component analysis: algorithms and applications. Neural Networks, 13(4-5):411–430, 2000.
Forrest N. Iandola, Song Han, Matthew W. Moskewicz, Khalid Ashraf, William J. Dally, and Kurt Keutzer. SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and <0.5MB model size. arXiv preprint arXiv:1602.07360, 2016. URL: https://arxiv.org/abs/1602.07360.
Sergey Ioffe and Christian Szegedy. Batch normalization: accelerating deep network training by reducing internal covariate shift. In Proceedings of the 32nd International Conference on Machine Learning (ICML). 2015.
Geoffrey Irving, Paul Christiano, and Dario Amodei. Ai safety via debate. arXiv preprint arXiv:1805.00899, 2018. URL: https://arxiv.org/abs/1805.00899.
Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A. Efros. Image-to-image translation with conditional adversarial networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2017.
Andrei Ivanov, Nikoli Dryden, Tal Ben-Nun, Shigang Li, and Torsten Hoefler. Data movement is all you need: a case study on optimizing transformers. In Proceedings of Machine Learning and Systems, volume 3, 711–732. 2021. URL: https://proceedings.mlsys.org/paper_files/paper/2021/file/bc86e95606a6392f51f95a8de106728d-Paper.pdf.
Gautier Izacard and Edouard Grave. Leveraging passage retrieval with generative models for open domain question answering. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume (EACL), 874–880. Association for Computational Linguistics, 2021. URL: https://aclanthology.org/2021.eacl-main.74/, doi:10.18653/v1/2021.eacl-main.74.
Tommi Jaakkola, Michael I. Jordan, and Satinder P. Singh. On the convergence of stochastic iterative dynamic programming algorithms. Neural Computation, 6:1185–1201, 1994.
Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, Hartwig Adam, and Dmitry Kalenichenko. Quantization and training of neural networks for efficient integer-arithmetic-only inference. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2704–2713. 2018.
Robert A. Jacobs, Michael I. Jordan, Steven J. Nowlan, and Geoffrey E. Hinton. Adaptive mixtures of local experts. Neural Computation, 3(1):79–87, 1991.
Sarthak Jain and Byron C. Wallace. Attention is not explanation. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), 3543–3556. 2019. URL: https://arxiv.org/abs/1902.10186.
Gareth James, Daniela Witten, Trevor Hastie, Robert Tibshirani, and Jonathan Taylor. An Introduction to Statistical Learning with Applications in Python. Springer Texts in Statistics. Springer, 2023. doi:10.1007/978-3-031-38747-0.
Kevin Jamieson and Ameet Talwalkar. Non-stochastic best arm identification and hyperparameter optimization. In Proceedings of the 19th International Conference on Artificial Intelligence and Statistics, 240–248. 2016.
Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine. When to trust your model: model-based policy optimization. In Advances in Neural Information Processing Systems (NeurIPS). 2019. URL: https://arxiv.org/abs/1906.08253.
Michael Janner, Qiyang Li, and Sergey Levine. Offline reinforcement learning as one big sequence modeling problem. In Advances in Neural Information Processing Systems, volume 34, 1273–1286. 2021. URL: https://arxiv.org/abs/2106.02039.
Dominik Janzing, Lenon Minorics, and Patrick Blöbaum. Feature relevance quantification in explainable AI: a causal problem. In Proceedings of the Twenty Third International Conference on Artificial Intelligence and Statistics (AISTATS), volume 108 of Proceedings of Machine Learning Research, 2907–2916. 2020.
Sadeep Jayasumana, Srikumar Ramalingam, Andreas Veit, Daniel Glasner, Ayan Chakrabarti, and Sanjiv Kumar. Rethinking FID: towards a better evaluation metric for image generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 9307–9315. 2024. URL: https://arxiv.org/abs/2401.09603, doi:10.1109/CVPR52733.2024.00889.
Samy Jelassi, David Brandfonbrener, Sham M. Kakade, and Eran Malach. Repeat after me: Transformers are better than state space models at copying. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, 21502–21521. 2024.
Frederick Jelinek and Robert L. Mercer. Interpolated estimation of Markov source parameters from sparse data. In Edzard S. Gelsema and Laveen N. Kanal, editors, Pattern Recognition in Practice: Proceedings of an International Workshop Held in Amsterdam, May 21–23, 1980, 381–397. North-Holland, 1980.
Russell Jeter, Christopher Josef, Supreeth Shashikumar, and Shamim Nemati. Does the “artificial intelligence clinician” learn optimal treatment strategies for sepsis in intensive care? 2019. URL: https://arxiv.org/abs/1902.03271, arXiv:1902.03271.
Yitong Ji, Aixin Sun, Jie Zhang, and Chenliang Li. A critical study on data leakage in recommender system offline evaluation. ACM Transactions on Information Systems, 41(3):75:1–75:27, 2023. arXiv:2010.11060. doi:10.1145/3569930.
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In Proceedings of the 38th International Conference on Machine Learning (ICML), volume 139, 4904–4916. 2021.
Ye Jia, Yu Zhang, Ron J. Weiss, Quan Wang, Jonathan Shen, Fei Ren, Zhifeng Chen, Patrick Nguyen, Ruoming Pang, Ignacio Lopez Moreno, and Yonghui Wu. Transfer learning from speaker verification to multispeaker text-to-speech synthesis. In Advances in Neural Information Processing Systems (NeurIPS). 2018.
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, and others. Mixtral of experts. arXiv preprint arXiv:2401.04088, 2024.
Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. SWE-bench: can language models resolve real-world GitHub issues? In International Conference on Learning Representations (ICLR). 2024.
Chi Jin, Rong Ge, Praneeth Netrapalli, Sham M. Kakade, and Michael I. Jordan. How to escape saddle points efficiently. In Proceedings of the 34th International Conference on Machine Learning (ICML). 2017.
Li Jing, Pascal Vincent, Yann LeCun, and Yuandong Tian. Understanding dimensional collapse in contrastive self-supervised learning. In International Conference on Learning Representations (ICLR). 2022. URL: https://arxiv.org/abs/2110.09348.
Glenn Jocher and Jing Qiu. Ultralytics YOLO. 2026. Software, dalla versione YOLOv5 (2020) a YOLO26 (2026). URL: https://docs.ultralytics.com.
Glenn Jocher, Jing Qiu, Mengyu Liu, Shuai Lyu, Fatih Cagatay Akyon, and Muhammet Esat Kalfaoglu. Ultralytics YOLO26: unified real-time end-to-end vision models. arXiv preprint arXiv:2606.03748, 2026.
Jeff Johnson, Matthijs Douze, and Hervé Jégou. Billion-scale similarity search with GPUs. IEEE Transactions on Big Data, 7(3):535–547, 2021. doi:10.1109/TBDATA.2019.2921572.
Justin Johnson, Alexandre Alahi, and Li Fei-Fei. Perceptual losses for real-time style transfer and super-resolution. In European Conference on Computer Vision, 694–711. 2016.
Melvin Johnson, Mike Schuster, Quoc V. Le, Maxim Krikun, Yonghui Wu, Zhifeng Chen, Nikhil Thorat, Fernanda Viégas, Martin Wattenberg, Greg Corrado, Macduff Hughes, and Jeffrey Dean. Google's multilingual neural machine translation system: enabling zero-shot translation. Transactions of the Association for Computational Linguistics, 5:339–351, 2017.
Nicholas K. Jong, Todd Hester, and Peter Stone. The utility of temporal abstraction in reinforcement learning. In Proceedings of the 7th International Conference on Autonomous Agents and Multiagent Systems (AAMAS), 299–306. 2008.
Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov. Bag of tricks for efficient text classification. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics (EACL), 427–431. 2017.
Norman P. Jouppi, Cliff Young, Nishant Patil, David Patterson, and others. In-datacenter performance analysis of a tensor processing unit. In Proceedings of the 44th Annual International Symposium on Computer Architecture, 1–12. 2017.
Rafal Jozefowicz, Wojciech Zaremba, and Ilya Sutskever. An empirical exploration of recurrent network architectures. In Proceedings of the 32nd International Conference on Machine Learning (ICML), 2342–2350. 2015.
Biing-Hwang Juang and A. Gray. Multiple stage vector quantization for speech coding. In IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), volume 7, 597–600. 1982. doi:10.1109/ICASSP.1982.1171604.
John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, and others. Highly accurate protein structure prediction with alphafold. Nature, 596(7873):583–589, 2021. doi:10.1038/s41586-021-03819-2.
Daniel Jurafsky and James H. Martin. Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition with Language Models. Edizione online degli autori, third edition, 2026. Bozza della terza edizione, rilasciata il 19 agosto 2026. URL: https://web.stanford.edu/~jurafsky/slp3/.
Christian Jutten and Jeanny Hérault. Blind separation of sources, part I: an adaptive algorithm based on neuromimetic architecture. Signal Processing, 24(1):1–10, 1991.
Karthikeyan K, Zihan Wang, Stephen Mayhew, and Dan Roth. Cross-lingual ability of multilingual BERT: an empirical study. In International Conference on Learning Representations. 2020.
Mark Kac. On distributions of certain Wiener functionals. Transactions of the American Mathematical Society, 65(1):1–13, 1949. doi:10.1090/S0002-9947-1949-0027960-X.
Adam Tauman Kalai, Ofir Nachum, Santosh S. Vempala, and Edwin Zhang. Why language models hallucinate. arXiv preprint arXiv:2509.04664, 2025.
Adam Tauman Kalai and Santosh S. Vempala. Calibrated language models must hallucinate. In Proceedings of the 56th Annual ACM Symposium on Theory of Computing (STOC '24), 160–171. ACM, 2024. doi:10.1145/3618260.3649777.
Dmitry Kalashnikov, Alex Irpan, Peter Pastor, Julian Ibarz, Alexander Herzog, Eric Jang, Deirdre Quillen, Ethan Holly, Mrinal Kalakrishnan, Vincent Vanhoucke, and Sergey Levine. Scalable deep reinforcement learning for vision-based robotic manipulation. In Proceedings of the 2nd Conference on Robot Learning (CoRL), volume 87 of Proceedings of Machine Learning Research, 651–673. PMLR, 2018. URL: https://arxiv.org/abs/1806.10293.
Nal Kalchbrenner and Phil Blunsom. Recurrent continuous translation models. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing (EMNLP), 1700–1709. 2013.
Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aäron van den Oord, Sander Dieleman, and Koray Kavukcuoglu. Efficient neural audio synthesis. In Proceedings of the 35th International Conference on Machine Learning. 2018.
Rudolf E. Kalman. A new approach to linear filtering and prediction problems. Journal of Basic Engineering, 82(1):35–45, 1960. doi:10.1115/1.3662552.
Faisal Kamiran and Toon Calders. Data preprocessing techniques for classification without discrimination. Knowledge and Information Systems, 33(1):1–33, 2012.
Bingyi Kang, Yang Yue, Rui Lu, Zhijie Lin, Yang Zhao, Kaixin Wang, Gao Huang, and Jiashi Feng. How far is video generation from world model: a physical law perspective. In Proceedings of the 42nd International Conference on Machine Learning (ICML). 2025. URL: https://arxiv.org/abs/2411.02385.
Minguk Kang, Jun-Yan Zhu, Richard Zhang, Jaesik Park, Eli Shechtman, Sylvain Paris, and Taesung Park. Scaling up GANs for text-to-image synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 10124–10134. 2023. URL: https://arxiv.org/abs/2303.05511, doi:10.1109/CVPR52729.2023.00976.
Wang-Cheng Kang and Julian McAuley. Self-attentive sequential recommendation. In 2018 IEEE International Conference on Data Mining (ICDM), 197–206. 2018. doi:10.1109/ICDM.2018.00035.
Anjuli Kannan, Yonghui Wu, Patrick Nguyen, Tara N. Sainath, Zhifeng Chen, and Rohit Prabhavalkar. An analysis of incorporating an external language model into a sequence-to-sequence model. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 2018. arXiv:1712.01996.
I. Kanter and H. Sompolinsky. Associative recall of memory without errors. Physical Review A, 35(1):380–392, 1987. doi:10.1103/PhysRevA.35.380.
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020.
Siddharth Karamcheti, Suraj Nair, Ashwin Balakrishna, Percy Liang, Thomas Kollar, and Dorsa Sadigh. Prismatic VLMs: investigating the design space of visually-conditioned language models. arXiv preprint arXiv:2402.07865, 2024.
Petr Karnakov, Sergey Litvinov, and Petros Koumoutsakos. Solving inverse problems in physics by optimizing a discrete loss: fast and accurate learning without neural networks. PNAS Nexus, 3(1):pgae005, 2024. doi:10.1093/pnasnexus/pgae005.
George Em Karniadakis, Ioannis G. Kevrekidis, Lu Lu, Paris Perdikaris, Sifan Wang, and Liu Yang. Physics-informed machine learning. Nature Reviews Physics, 3(6):422–440, 2021. doi:10.1038/s42254-021-00314-5.
Zohar Karnin, Tomer Koren, and Oren Somekh. Almost optimal exploration in multi-armed bandits. In Proceedings of the 30th International Conference on Machine Learning (ICML), volume 28 of PMLR, 1238–1246. 2013.
Andrej Karpathy. +1 for “context engineering” over “prompt engineering”. Post su X (Twitter), 25 giugno 2025, 2025. URL: https://x.com/karpathy/status/1937902205765607626.
Andrej Karpathy. AGI is still a decade away. Intervista di Dwarkesh Patel, 10 2025. URL: https://www.dwarkesh.com/p/andrej-karpathy.
Vladimir Karpukhin, Barlas Oğuz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 6769–6781. 2020. doi:10.18653/v1/2020.emnlp-main.550.
Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen. Progressive growing of GANs for improved quality, stability, and variation. In International Conference on Learning Representations (ICLR). 2018.
Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine. Elucidating the design space of diffusion-based generative models. In Advances in Neural Information Processing Systems (NeurIPS). 2022. URL: https://arxiv.org/abs/2206.00364.
Tero Karras, Miika Aittala, Samuli Laine, Erik Härkönen, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Alias-free generative adversarial networks. In Advances in Neural Information Processing Systems, volume 34, 852–863. 2021. URL: https://arxiv.org/abs/2106.12423.
Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2019.
Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Analyzing and improving the image quality of StyleGAN. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2020.
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret. Transformers are RNNs: fast autoregressive transformers with linear attention. In Proceedings of the 37th International Conference on Machine Learning (ICML). 2020.
Guy Katz, Clark Barrett, David L. Dill, Kyle Julian, and Mykel J. Kochenderfer. Reluplex: an efficient SMT solver for verifying deep neural networks. In Computer Aided Verification (CAV), 97–117. 2017.
Slava M. Katz. Estimation of probabilities from sparse data for the language model component of a speech recognizer. IEEE Transactions on Acoustics, Speech, and Signal Processing, 35(3):400–401, 1987. doi:10.1109/TASSP.1987.1165125.
Emilie Kaufmann, Nathaniel Korda, and Rémi Munos. Thompson sampling: an asymptotically optimal finite-time analysis. In Algorithmic Learning Theory (ALT), 199–213. 2012.
Koray Kavukcuoglu, Marc'Aurelio Ranzato, and Yann LeCun. Fast inference in sparse coding algorithms with applications to object recognition. Technical Report CBLL-TR-2008-12-01, Computational and Biological Learning Laboratory, Courant Institute, New York University, 2008. arXiv:1010.3467.
Guolin Ke, Qi Meng, Thomas Finley, Taifeng Wang, Wei Chen, Weidong Ma, Qiwei Ye, and Tie-Yan Liu. Lightgbm: a highly efficient gradient boosting decision tree. In Advances in Neural Information Processing Systems (NeurIPS), volume 30. 2017.
S. Sathiya Keerthi and Chih-Jen Lin. Asymptotic behaviors of support vector machines with Gaussian kernel. Neural Computation, 15(7):1667–1689, 2003.
Alex Kendall, Yarin Gal, and Roberto Cipolla. Multi-task learning using uncertainty to weigh losses for scene geometry and semantics. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 7482–7491. 2018. URL: https://arxiv.org/abs/1705.07115.
James Kennedy and Russell Eberhart. Particle swarm optimization. In Proceedings of ICNN'95: International Conference on Neural Networks, volume 4, 1942–1948. Perth, WA, Australia, 1995. IEEE.
James Kennedy and Rui Mendes. Population structure and particle swarm performance. In Proceedings of the 2002 Congress on Evolutionary Computation (CEC'02), volume 2, 1671–1676. IEEE, 2002. doi:10.1109/CEC.2002.1004493.
Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis. 3d Gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics, 2023. URL: https://arxiv.org/abs/2308.04079.
Mark D. Kernighan, Kenneth W. Church, and William A. Gale. A spelling correction program based on a noisy channel model. In COLING 1990 Volume 2: Papers Presented to the 13th International Conference on Computational Linguistics, 205–210. 1990. doi:10.3115/997939.997975.
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang. On large-batch training for deep learning: generalization gap and sharp minima. In 5th International Conference on Learning Representations (ICLR). 2017. URL: https://arxiv.org/abs/1609.04836, arXiv:1609.04836.
Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. Generalization through memorization: nearest neighbor language models. In International Conference on Learning Representations (ICLR). 2020. URL: https://openreview.net/forum?id=HklBjCEKvH.
Ehsan Kharazmi, Zhongqiang Zhang, and George Em Karniadakis. Variational physics-informed neural networks for solving partial differential equations. arXiv preprint arXiv:1912.00873, 2019. URL: https://arxiv.org/abs/1912.00873.
Ehsan Kharazmi, Zhongqiang Zhang, and George Em Karniadakis. Hp-VPINNs: variational physics-informed neural networks with domain decomposition. Computer Methods in Applied Mechanics and Engineering, 374:113547, 2021. doi:10.1016/j.cma.2020.113547.
Omar Khattab and Matei Zaharia. ColBERT: efficient and effective passage search via contextualized late interaction over BERT. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR), 39–48. 2020.
Kevin Kilgour, Mauricio Zuluaga, Dominik Roblek, and Matthew Sharifi. Fréchet audio distance: a reference-free metric for evaluating music enhancement algorithms. In Interspeech, 2350–2354. 2019. doi:10.21437/Interspeech.2019-2219.
Been Kim, Rajiv Khanna, and Oluwasanmi O. Koyejo. Examples are not enough, learn to criticize! criticism for interpretability. In Advances in Neural Information Processing Systems 29 (NeurIPS), 2280–2288. 2016.
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viégas, and Rory Sayres. Interpretability beyond feature attribution: quantitative testing with concept activation vectors (TCAV). In Proceedings of the 35th International Conference on Machine Learning (ICML), volume 80 of Proceedings of Machine Learning Research, 2668–2677. 2018.
David Kim. Context-engineering. Repository GitHub, 2025. URL: davidkimai/Context-Engineering.
Dongjun Kim, Chieh-Hsin Lai, Wei-Hsiang Liao, Naoki Murata, Yuhta Takida, Toshimitsu Uesaka, Yutong He, Yuki Mitsufuji, and Stefano Ermon. Consistency trajectory models: learning probability flow ODE trajectory of diffusion. In International Conference on Learning Representations (ICLR). 2024. URL: https://arxiv.org/abs/2310.02279.
Duckhwan Kim, Jaeha Kung, Sek Chai, Sudhakar Yalamanchili, and Saibal Mukhopadhyay. Neurocube: a programmable digital neuromorphic architecture with high-density 3D memory. In 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA), 380–392. 2016. doi:10.1109/ISCA.2016.41.
Elliot Myunghoon Kim, Avi Garg, Kenny Peng, and Nikhil Garg. Correlated errors in large language models. In Proceedings of the 42nd International Conference on Machine Learning (ICML), volume 267 of Proceedings of Machine Learning Research, 30038–30066. 2025. URL: https://arxiv.org/abs/2506.07962.
Jaehyeon Kim, Jungil Kong, and Juhee Son. Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech. In Proceedings of the 38th International Conference on Machine Learning (ICML), volume 139, 5530–5540. 2021.
Moo Jin Kim, Chelsea Finn, and Percy Liang. Fine-tuning vision-language-action models: optimizing speed and success. In Proceedings of Robotics: Science and Systems (RSS). 2025. doi:10.15607/RSS.2025.XXI.017.
Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, Quan Vuong, Thomas Kollar, Benjamin Burchfiel, Russ Tedrake, Dorsa Sadigh, Sergey Levine, Percy Liang, and Chelsea Finn. OpenVLA: an open-source vision-language-action model. In Proceedings of the 8th Conference on Robot Learning (CoRL), volume 270 of Proceedings of Machine Learning Research, 2679–2713. 2024. arXiv:2406.09246.
Yoon Kim and Alexander M. Rush. Sequence-level knowledge distillation. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing (EMNLP), 1317–1327. 2016. doi:10.18653/v1/D16-1139.
Diederik P. Kingma and Jimmy Ba. Adam: a method for stochastic optimization. In International Conference on Learning Representations (ICLR). 2015.
Diederik P. Kingma and Prafulla Dhariwal. Glow: generative flow with invertible 1x1 convolutions. In Advances in Neural Information Processing Systems (NeurIPS). 2018.
Diederik P. Kingma and Ruiqi Gao. Understanding diffusion objectives as the ELBO with simple data augmentation. In Advances in Neural Information Processing Systems (NeurIPS). 2023.
Diederik P. Kingma, Tim Salimans, Rafal Jozefowicz, Xi Chen, Ilya Sutskever, and Max Welling. Improved variational inference with inverse autoregressive flow. In Advances in Neural Information Processing Systems (NeurIPS), volume 29. 2016. URL: https://arxiv.org/abs/1606.04934.
Diederik P. Kingma, Tim Salimans, Ben Poole, and Jonathan Ho. Variational diffusion models. In Advances in Neural Information Processing Systems (NeurIPS). 2021.
Diederik P. Kingma and Max Welling. Auto-encoding variational bayes. In Proceedings of the 2nd International Conference on Learning Representations (ICLR). 2014. URL: https://arxiv.org/abs/1312.6114.
Diederik P. Kingma and Max Welling. An introduction to variational autoencoders. Foundations and Trends in Machine Learning, 12(4):307–392, 2019. URL: https://arxiv.org/abs/1906.02691.
Thomas N. Kipf and Max Welling. Semi-supervised classification with graph convolutional networks. In International Conference on Learning Representations (ICLR). 2017. URL: https://arxiv.org/abs/1609.02907.
John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein. A watermark for large language models. In Proceedings of the 40th International Conference on Machine Learning (ICML), 17061–17084. 2023.
Polina Kirichenko, Pavel Izmailov, and Andrew Gordon Wilson. Why normalizing flows fail to detect out-of-distribution data. In Advances in Neural Information Processing Systems (NeurIPS), volume 33, 20578–20589. 2020.
Alexander Kirillov, Kaiming He, Ross Girshick, Carsten Rother, and Piotr Dollár. Panoptic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2019.
Scott Kirkpatrick, C. Daniel Gelatt, and Mario P. Vecchi. Optimization by simulated annealing. Science, 220(4598):671–680, 1983.
Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya. Reformer: the efficient transformer. In International Conference on Learning Representations (ICLR). 2020. URL: https://arxiv.org/abs/2001.04451.
Stephen C. Kleene. Representation of events in nerve nets and finite automata. In Claude E. Shannon and John McCarthy, editors, Automata Studies, number 34 in Annals of Mathematics Studies, pages 3–41. Princeton University Press, 1956.
Jon Kleinberg, Sendhil Mullainathan, and Manish Raghavan. Inherent trade-offs in the fair determination of risk scores. In Proceedings of the 8th Innovations in Theoretical Computer Science Conference (ITCS), 43:1–43:23. 2017. doi:10.4230/LIPIcs.ITCS.2017.43.
Jon M. Kleinberg. An impossibility theorem for clustering. In Advances in Neural Information Processing Systems 15 (NIPS 2002). 2002.
Anton Klenitskiy and Alexey Vasilev. Turning dross into gold loss: is BERT4Rec really better than SASRec? In Proceedings of the 17th ACM Conference on Recommender Systems (RecSys '23), 1120–1125. ACM, 2023. doi:10.1145/3604915.3610644.
Peter E. Kloeden and Eckhard Platen. Numerical Solution of Stochastic Differential Equations. Springer, Berlin, Heidelberg, 1992. doi:10.1007/978-3-662-12616-5.
Reinhard Kneser and Hermann Ney. Improved backing-off for m-gram language modeling. In Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP-95), volume 1, 181–184. Detroit, MI, 1995. doi:10.1109/ICASSP.1995.479394.
Donald E. Knuth. The Art of Computer Programming, Volume 2: Seminumerical Algorithms. Addison-Wesley, third edition, 1997.
Donald E. Knuth and Ronald W. Moore. An analysis of alpha-beta pruning. Artificial Intelligence, 6(4):293–326, 1975. doi:10.1016/0004-3702(75)90019-3.
Dmitry Kobak and George C. Linderman. Initialization is critical for preserving global data structure in both t-sne and umap. Nature Biotechnology, 39(2):156–157, 2021. doi:10.1038/s41587-020-00809-z.
Dmitrii Kochkov, Janni Yuval, Ian Langmore, Peter Norgaard, Jamie Smith, Griffin Mooers, Milan Klöwer, James Lottes, Stephan Rasp, Peter Düben, Sam Hatfield, Peter Battaglia, Alvaro Sanchez-Gonzalez, Matthew Willson, Michael P. Brenner, and Stephan Hoyer. Neural general circulation models for weather and climate. Nature, 632(8027):1060–1066, 2024. doi:10.1038/s41586-024-07744-y.
Levente Kocsis and Csaba Szepesvári. Bandit based Monte-Carlo planning. In Machine Learning: ECML 2006, volume 4212 of Lecture Notes in Computer Science, 282–293. Springer, 2006. URL: https://doi.org/10.1007/11871842_29.
Philipp Koehn and Rebecca Knowles. Six challenges for neural machine translation. In Proceedings of the First Workshop on Neural Machine Translation, 28–39. 2017.
W. Koenig, H. K. Dunn, and L. Y. Lacy. The sound spectrograph. The Journal of the Acoustical Society of America, 18(1):19–49, July 1946. doi:10.1121/1.1916342.
Teuvo Kohonen. Correlation matrix memories. IEEE Transactions on Computers, C-21(4):353–359, 1972. doi:10.1109/TC.1972.5008975.
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. Large language models are zero-shot reasoners. In Advances in Neural Information Processing Systems (NeurIPS), volume 35. 2022. URL: https://arxiv.org/abs/2205.11916.
Andrei N. Kolmogorov. Three approaches to the quantitative definition of information. Problems of Information Transmission, 1(1):1–7, 1965.
Vladimir Koltchinskii. Rademacher penalties and structural risk minimization. IEEE Transactions on Information Theory, 47(5):1902–1914, 2001.
Matthieu Komorowski, Leo A. Celi, Omar Badawi, Anthony C. Gordon, and A. Aldo Faisal. The artificial intelligence clinician learns optimal treatment strategies for sepsis in intensive care. Nature Medicine, 24(11):1716–1720, 2018. doi:10.1038/s41591-018-0213-5.
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae. Hifi-gan: generative adversarial networks for efficient and high fidelity speech synthesis. In Advances in Neural Information Processing Systems 33 (NeurIPS), 17022–17033. 2020. URL: https://arxiv.org/abs/2010.05646.
Qiuqiang Kong, Yin Cao, Turab Iqbal, Yuxuan Wang, Wenwu Wang, and Mark D. Plumbley. PANNs: large-scale pretrained audio neural networks for audio pattern recognition. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 28:2880–2894, 2020. doi:10.1109/TASLP.2020.3030497.
Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro. DiffWave: a versatile diffusion model for audio synthesis. In International Conference on Learning Representations (ICLR). 2021.
Tomasz Korbak, Ethan Perez, and Christopher L. Buckley. RL with KL penalties is better viewed as Bayesian inference. In Findings of the Association for Computational Linguistics: EMNLP 2022. 2022. URL: https://arxiv.org/abs/2205.11275.
Yehuda Koren. Factorization meets the neighborhood: a multifaceted collaborative filtering model. In Proceedings of the 14th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD '08), 426–434. ACM, 2008. doi:10.1145/1401890.1401944.
Yehuda Koren, Robert Bell, and Chris Volinsky. Matrix factorization techniques for recommender systems. Computer, 42(8):30–37, 2009.
Richard E. Korf. Depth-first iterative-deepening: an optimal admissible tree search. Artificial Intelligence, 27(1):97–109, 1985.
Vijay Anand Korthikanti, Jared Casper, Sangkug Lym, Lawrence McAfee, Michael Andersch, Mohammad Shoeybi, and Bryan Catanzaro. Reducing activation recomputation in large transformer models. In Proceedings of Machine Learning and Systems 5 (MLSys 2023), 341–353. Curran Associates, 2023. URL: https://proceedings.mlsys.org/paper_files/paper/2023/file/80083951326cf5b35e5100260d64ed81-Paper-mlsys2023.pdf.
Ilya Kostrikov, Ashvin Nair, and Sergey Levine. Offline reinforcement learning with implicit q-learning. In International Conference on Learning Representations (ICLR). 2022. URL: https://arxiv.org/abs/2110.06169.
Olga Kovaleva, Saurabh Kulshreshtha, Anna Rogers, and Anna Rumshisky. BERT busters: outlier dimensions that disrupt transformers. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, 3392–3405. 2021. doi:10.18653/v1/2021.findings-acl.300.
Hendrik A. Kramers. Brownian motion in a field of force and the diffusion model of chemical reactions. Physica, 7(4):284–304, 1940. doi:10.1016/S0031-8914(40)90098-2.
Dominik Kreuzberger, Niklas Kühl, and Sebastian Hirschl. Machine learning operations (mlops): overview, definition, and architecture. IEEE Access, 11:31866–31879, 2023.
Walid Krichene and Steffen Rendle. On sampled metrics for item recommendation. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD '20), 1748–1757. 2020. doi:10.1145/3394486.3403226.
Aditi S. Krishnapriyan, Amir Gholami, Shandian Zhe, Robert M. Kirby, and Michael W. Mahoney. Characterizing possible failure modes in physics-informed neural networks. In Advances in Neural Information Processing Systems (NeurIPS), volume 34. 2021. URL: https://arxiv.org/abs/2109.01050.
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. Imagenet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems, volume 25. 2012.
Anders Krogh and Jesper Vedelsby. Neural network ensembles, cross validation, and active learning. In Advances in Neural Information Processing Systems 7 (NIPS 1994), 231–238. MIT Press, 1995.
Dmitry Krotov and John J. Hopfield. Dense associative memory for pattern recognition. In Advances in Neural Information Processing Systems 29 (NIPS). 2016. URL: https://arxiv.org/abs/1606.01164.
Taku Kudo. Subword regularization: improving neural network translation models with multiple subword candidates. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (ACL), 66–75. 2018.
Taku Kudo and John Richardson. SentencePiece: a simple and language independent subword tokenizer and detokenizer for neural text processing. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations (EMNLP), 66–71. 2018.
Tejas D. Kulkarni, Karthik R. Narasimhan, Ardavan Saeedi, and Joshua B. Tenenbaum. Hierarchical deep reinforcement learning: integrating temporal abstraction and intrinsic motivation. In Advances in Neural Information Processing Systems, volume 29. 2016. URL: https://arxiv.org/abs/1604.06057.
Solomon Kullback and Richard A. Leibler. On information and sufficiency. The Annals of Mathematical Statistics, 22(1):79–86, 1951.
Ananya Kumar, Aditi Raghunathan, Robbie Jones, Tengyu Ma, and Percy Liang. Fine-tuning can distort pretrained features and underperform out-of-distribution. In International Conference on Learning Representations (ICLR). 2022.
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine. Conservative q-learning for offline reinforcement learning. In Advances in Neural Information Processing Systems (NeurIPS). 2020. URL: https://arxiv.org/abs/2006.04779.
I. Elizabeth Kumar, Suresh Venkatasubramanian, Carlos Scheidegger, and Sorelle Friedler. Problems with Shapley-value-based explanations as feature importance measures. In Proceedings of the 37th International Conference on Machine Learning (ICML), volume 119 of Proceedings of Machine Learning Research, 5491–5500. 2020.
Rithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar, and Kundan Kumar. High-fidelity audio compression with improved RVQGAN. In Advances in Neural Information Processing Systems, volume 36, 27980–27993. 2023.
H. T. Kung. Why systolic architectures? Computer, 15(1):37–46, 1982.
Vitaly Kurin, Alessandro De Palma, Ilya Kostrikov, Shimon Whiteson, and M. Pawan Kumar. In defense of the unitary scalarization for deep multi-task learning. In Advances in Neural Information Processing Systems (NeurIPS), volume 35. 2022.
Denis Kwiatkowski, Peter C. B. Phillips, Peter Schmidt, and Yongcheol Shin. Testing the null hypothesis of stationarity against the alternative of a unit root. Journal of Econometrics, 54(1-3):159–178, 1992.
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with PagedAttention. In ACM SIGOPS 29th Symposium on Operating Systems Principles (SOSP). 2023.
Tuomas Kynkäänniemi, Miika Aittala, Tero Karras, Samuli Laine, Timo Aila, and Jaakko Lehtinen. Applying guidance in a limited interval improves sample and distribution quality in diffusion models. In Advances in Neural Information Processing Systems (NeurIPS). 2024. URL: https://arxiv.org/abs/2404.07724.
Tuomas Kynkäänniemi, Tero Karras, Samuli Laine, Jaakko Lehtinen, and Timo Aila. Improved precision and recall metric for assessing generative models. In Advances in Neural Information Processing Systems (NeurIPS). 2019.
Krishna K. Ladha. The Condorcet jury theorem, free speech, and correlated votes. American Journal of Political Science, 36(3):617–634, 1992. doi:10.2307/2111584.
John D. Lafferty, Andrew McCallum, and Fernando C. N. Pereira. Conditional random fields: probabilistic models for segmenting and labeling sequence data. In Proceedings of the Eighteenth International Conference on Machine Learning (ICML 2001), 282–289. Morgan Kaufmann, 2001. URL: https://dl.acm.org/doi/10.5555/645530.655813.
Isaac E. Lagaris, Aristidis Likas, and Dimitrios I. Fotiadis. Artificial neural networks for solving ordinary and partial differential equations. IEEE Transactions on Neural Networks, 9(5):987–1000, 1998. URL: https://doi.org/10.1109/72.712178, doi:10.1109/72.712178.
Aakash Lahoti, Kevin Y. Li, Berlin Chen, Caitlin Wang, Aviv Bick, J. Zico Kolter, Tri Dao, and Albert Gu. Mamba-3: improved sequence modeling using state space principles. In International Conference on Learning Representations (ICLR). 2026.
Chieh-Hsin Lai, Yang Song, Dongjun Kim, Yuki Mitsufuji, and Stefano Ermon. The Principles of Diffusion Models. MIT Press, 2026. arXiv:2510.21890. URL: https://arxiv.org/abs/2510.21890.
Tze Leung Lai and Herbert Robbins. Asymptotically efficient adaptive allocation rules. Advances in Applied Mathematics, 6(1):4–22, 1985.
Remi Lam, Alvaro Sanchez-Gonzalez, Matthew Willson, Peter Wirnsberger, Meire Fortunato, Ferran Alet, Suman Ravuri, Timo Ewalds, Zach Eaton-Rosen, Weihua Hu, Alexander Merose, Stephan Hoyer, George Holland, Oriol Vinyals, Jacklynn Stott, Alexander Pritzel, Shakir Mohamed, and Peter Battaglia. Learning skillful medium-range global weather forecasting. Science, 382(6677):1416–1421, 2023.
Guillaume Lample, Miguel Ballesteros, Sandeep Subramanian, Kazuya Kawakami, and Chris Dyer. Neural architectures for named entity recognition. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 260–270. San Diego, California, 2016. Association for Computational Linguistics. doi:10.18653/v1/N16-1030.
Guillaume Lample and Alexis Conneau. Cross-lingual language model pretraining. In Advances in Neural Information Processing Systems, volume 32. 2019.
Leslie Lamport. The part-time parliament. ACM Transactions on Computer Systems, 16(2):133–169, May 1998. doi:10.1145/279227.279229.
Leslie Lamport, Robert Shostak, and Marshall Pease. The byzantine generals problem. ACM Transactions on Programming Languages and Systems, 4(3):382–401, 1982. doi:10.1145/357172.357176.
Paul Langevin. Sur la théorie du mouvement brownien. Comptes rendus hebdomadaires des séances de l'Académie des sciences, 146:530–533, 1908.
Lauro Langosco, Jack Koch, Lee Sharkey, Jacob Pfau, and David Krueger. Goal misgeneralization in deep reinforcement learning. In Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, 12004–12019. PMLR, 2022. URL: https://proceedings.mlr.press/v162/langosco22a.html.
Tamera Lanham, Anna Chen, Ansh Radhakrishnan, Benoit Steiner, Carson Denison, Danny Hernandez, Dustin Li, Esin Durmus, Evan Hubinger, Jackson Kernion, and others. Measuring faithfulness in chain-of-thought reasoning. arXiv preprint arXiv:2307.13702, 2023. URL: https://arxiv.org/abs/2307.13702.
Hugo Larochelle and Iain Murray. The neural autoregressive distribution estimator. In Proceedings of the 14th International Conference on Artificial Intelligence and Statistics (AISTATS), volume 15 of Proceedings of Machine Learning Research, 29–37. 2011.
Andrew Lavin and Scott Gray. Fast algorithms for convolutional neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 4013–4021. 2016. arXiv:1509.09308.
David Lazer, Ryan Kennedy, Gary King, and Alessandro Vespignani. The parable of Google Flu: traps in big data analysis. Science, 343(6176):1203–1205, 2014.
Charline Le Lan and Laurent Dinh. Perfect density models cannot guarantee anomaly detection. Entropy, 23(12):1690, 2021. doi:10.3390/e23121690.
Yann LeCun. Predictive learning. Keynote, Conference on Neural Information Processing Systems (NIPS), Barcellona, 2016. La formula della torta e' riportata verbatim, e attribuita a questo keynote, in arXiv:2502.03038. Le slide primarie non sono state reperite (agosto 2026): due URL canonici danno 404 e la scheda NIPS non riporta il talk. All'ISSCC 2019 LeCun sostituisce «unsupervised» con «self-supervised».
Yann LeCun. A path towards autonomous machine intelligence (version 0.9.2, 2022-06-27). Position paper, OpenReview, 2022. URL: https://openreview.net/pdf?id=BZ5a1r-kVsf.
Yann LeCun, Bernhard Boser, John S. Denker, Donnie Henderson, Richard E. Howard, Wayne Hubbard, and Lawrence D. Jackel. Backpropagation applied to handwritten zip code recognition. Neural Computation, 1(4):541–551, 1989.
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324, 1998.
Yann LeCun, Léon Bottou, Genevieve B. Orr, and Klaus-Robert Müller. Efficient backprop. In Genevieve B. Orr and Klaus-Robert Müller, editors, Neural Networks: Tricks of the Trade, volume 1524 of Lecture Notes in Computer Science, pages 9–50. Springer, Berlin, Heidelberg, 1998. doi:10.1007/3-540-49430-8_2.
Yann LeCun, Sumit Chopra, Raia Hadsell, Marc'Aurelio Ranzato, and Fu Jie Huang. A tutorial on energy-based learning. In Gökhan Bakır, Thomas Hofmann, Bernhard Schölkopf, Alexander J. Smola, and Ben Taskar, editors, Predicting Structured Data. MIT Press, 2006. URL: http://yann.lecun.com/exdb/publis/pdf/lecun-06.pdf.
Yann LeCun, John S. Denker, and Sara A. Solla. Optimal brain damage. In Advances in Neural Information Processing Systems 2 (NIPS 1989), 598–605. Morgan Kaufmann, 1990.
Yann LeCun and Ishan Misra. Self-supervised learning: the dark matter of intelligence. Meta AI Blog, 3 2021. URL: https://ai.meta.com/blog/self-supervised-learning-the-dark-matter-of-intelligence/.
Christian Ledig, Lucas Theis, Ferenc Huszár, Jose Caballero, Andrew Cunningham, Alejandro Acosta, Andrew Aitken, Alykhan Tejani, Johannes Totz, Zehan Wang, and Wenzhe Shi. Photo-realistic single image super-resolution using a generative adversarial network. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2017.
Hyuk Lee and In Seok Kang. Neural algorithm for solving differential equations. Journal of Computational Physics, 91(1):110–131, 1990. doi:10.1016/0021-9991(90)90007-N.
Jason D. Lee, Max Simchowitz, Michael I. Jordan, and Benjamin Recht. Gradient descent only converges to minimizers. In Conference on Learning Theory (COLT), 1246–1257. 2016.
Kenton Lee, Ming-Wei Chang, and Kristina Toutanova. Latent retrieval for weakly supervised open domain question answering. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL), 6086–6096. Florence, Italy, 2019. Association for Computational Linguistics. URL: https://aclanthology.org/P19-1612/, doi:10.18653/v1/P19-1612.
Te-Won Lee, Mark Girolami, and Terrence J. Sejnowski. Independent component analysis using an extended infomax algorithm for mixed subgaussian and supergaussian sources. Neural Computation, 11(2):417–441, 1999.
Shane Legg and Marcus Hutter. Universal intelligence: a definition of machine intelligence. Minds and Machines, 17(4):391–444, 2007. doi:10.1007/s11023-007-9079-x.
Sicong Leng, Hang Zhang, Guanzheng Chen, Xin Li, Shijian Lu, Chunyan Miao, and Lidong Bing. Mitigating object hallucinations in large vision-language models through visual contrastive decoding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2024.
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. GShard: scaling giant models with conditional computation and automatic sharding. In International Conference on Learning Representations (ICLR). 2021.
Moshe Leshno, Vladimir Ya. Lin, Allan Pinkus, and Shimon Schocken. Multilayer feedforward networks with a nonpolynomial activation function can approximate any function. Neural Networks, 6(6):861–867, 1993. doi:10.1016/S0893-6080(05)80131-5.
Vladimir Iosifovich Levenshtein. Binary codes capable of correcting deletions, insertions, and reversals. Soviet Physics Doklady, 10(8):707–710, 1966. Traduzione inglese dell'originale russo apparso in Doklady Akademii Nauk SSSR, 163(4):845–848, 1965.
Yaniv Leviathan, Matan Kalman, and Yossi Matias. Fast inference from transformers via speculative decoding. In Proceedings of the 40th International Conference on Machine Learning (ICML), 19274–19286. 2023.
David A. Levin, Yuval Peres, and Elizabeth L. Wilmer. Markov Chains and Mixing Times. American Mathematical Society, 2 edition, 2017.
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu. Offline reinforcement learning: tutorial, review, and perspectives on open problems. 2020. URL: https://arxiv.org/abs/2005.01643, arXiv:2005.01643.
Omer Levy and Yoav Goldberg. Neural word embedding as implicit matrix factorization. In Advances in Neural Information Processing Systems (NeurIPS), volume 27. 2014.
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. Retrieval-augmented generation for knowledge-intensive NLP tasks. In Advances in Neural Information Processing Systems, volume 33, 9459–9474. 2020. URL: https://arxiv.org/abs/2005.11401.
Guohao Li, Matthias Müller, Ali Thabet, and Bernard Ghanem. Deepgcns: can gcns go as deep as cnns? In Proceedings of the IEEE/CVF International Conference on Computer Vision. 2019.
Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan. A diversity-promoting objective function for neural conversation models. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL-HLT), 110–119. 2016.
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models. In Proceedings of the 40th International Conference on Machine Learning (ICML), 19730–19742. 2023.
Kenneth Li, Aspen K. Hopkins, David Bau, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg. Emergent world representations: exploring a sequence model trained on a synthetic task. In International Conference on Learning Representations (ICLR). 2023. URL: https://arxiv.org/abs/2210.13382.
Liam Li, Kevin Jamieson, Afshin Rostamizadeh, Ekaterina Gonina, Jonathan Ben-tzur, Moritz Hardt, Benjamin Recht, and Ameet Talwalkar. A system for massively parallel hyperparameter tuning. In Proceedings of Machine Learning and Systems, volume 2, 230–246. 2020.
Lihong Li, Wei Chu, John Langford, and Robert E. Schapire. A contextual-bandit approach to personalized news article recommendation. In Proceedings of the 19th International Conference on World Wide Web, 661–670. ACM, 2010.
Lisha Li, Kevin Jamieson, Giulia DeSalvo, Afshin Rostamizadeh, and Ameet Talwalkar. Hyperband: a novel bandit-based approach to hyperparameter optimization. Journal of Machine Learning Research, 18(185):1–52, 2018.
Ming Li and Paul M. B. Vitányi. An Introduction to Kolmogorov Complexity and Its Applications. Springer, fourth edition, 2019. doi:10.1007/978-3-030-11298-1.
Mu Li, David G. Andersen, Jun Woo Park, Alexander J. Smola, Amr Ahmed, Vanja Josifovski, James Long, Eugene J. Shekita, and Bor-Yiing Su. Scaling distributed machine learning with the parameter server. In 11th USENIX Symposium on Operating Systems Design and Implementation (OSDI 14), 583–598. USENIX Association, 2014. URL: https://www.usenix.org/conference/osdi14/technical-sessions/presentation/li_mu.
Ninghui Li, Tiancheng Li, and Suresh Venkatasubramanian. $t$-closeness: privacy beyond $k$-anonymity and ℓ-diversity. In 2007 IEEE 23rd International Conference on Data Engineering (ICDE), 106–115. 2007. doi:10.1109/ICDE.2007.367856.
Qimai Li, Zhichao Han, and Xiao-Ming Wu. Deeper insights into graph convolutional networks for semi-supervised learning. In Proceedings of the AAAI Conference on Artificial Intelligence. 2018.
Xiang Li, Shuo Chen, Xiaolin Hu, and Jian Yang. Understanding the disharmony between dropout and batch normalization by variance shift. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2682–2690. 2019.
Yanghao Li, Naiyan Wang, Jiaying Liu, and Xiaodi Hou. Demystifying neural style transfer. In Proceedings of the 26th International Joint Conference on Artificial Intelligence (IJCAI), 2230–2236. 2017.
Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji-Rong Wen. Evaluating object hallucination in large vision-language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), 292–305. 2023.
Yuhui Li, Fangyun Wei, Chao Zhang, and Hongyang Zhang. EAGLE: speculative sampling requires rethinking feature uncertainty. In Proceedings of the 41st International Conference on Machine Learning (ICML). 2024.
Zhiyuan Li, Hong Liu, Denny Zhou, and Tengyu Ma. Chain of thought empowers transformers to solve inherently serial problems. In International Conference on Learning Representations (ICLR), 11911–11943. 2024. URL: https://openreview.net/forum?id=3EWTEy9MTM.
Zhuowan Li, Cheng Li, Mingyang Zhang, Qiaozhu Mei, and Michael Bendersky. Retrieval augmented generation or long-context LLMs? a comprehensive study and hybrid approach. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track, 881–893. Miami, Florida, US, 2024. Association for Computational Linguistics. doi:10.18653/v1/2024.emnlp-industry.66.
Zongyi Li, Nikola Kovachki, Kamyar Azizzadenesheli, Burigede Liu, Kaushik Bhattacharya, Andrew Stuart, and Anima Anandkumar. Fourier neural operator for parametric partial differential equations. In International Conference on Learning Representations (ICLR). 2021. URL: https://arxiv.org/abs/2010.08895.
Zongyi Li, Hongkai Zheng, Nikola Kovachki, David Jin, Haoxuan Chen, Burigede Liu, Kamyar Azizzadenesheli, and Anima Anandkumar. Physics-informed neural operator for learning partial differential equations. ACM/IMS Journal of Data Science, 1(3):1–27, 2024. doi:10.1145/3648506.
Weixin Liang, Yuhui Zhang, Yongchan Kwon, Serena Yeung, and James Zou. Mind the gap: understanding the modality gap in multi-modal contrastive representation learning. In Advances in Neural Information Processing Systems, volume 35. 2022.
Opher Lieber, Barak Lenz, Hofit Bata, and others. Jamba: a hybrid transformer-mamba language model. arXiv preprint arXiv:2403.19887, 2024.
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. Continuous control with deep reinforcement learning. In International Conference on Learning Representations (ICLR). 2016. URL: https://arxiv.org/abs/1509.02971.
Bryan Lim, Sercan Ö. Arık, Nicolas Loeff, and Tomas Pfister. Temporal fusion transformers for interpretable multi-horizon time series forecasting. International Journal of Forecasting, 37(4):1748–1764, 2021. URL: https://arxiv.org/abs/1912.09363.
Chin-Yew Lin. Rouge: a package for automatic evaluation of summaries. In Text Summarization Branches Out: Proceedings of the ACL-04 Workshop, 74–81. 2004. URL: https://aclanthology.org/W04-1013/.
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, and Song Han. AWQ: activation-aware weight quantization for LLM compression and acceleration. In Proceedings of Machine Learning and Systems (MLSys). 2024.
Long-Ji Lin. Self-improving reactive agents based on reinforcement learning, planning and teaching. Machine Learning, 8(3–4):293–321, 1992. doi:10.1007/BF00992699.
Min Lin, Qiang Chen, and Shuicheng Yan. Network in network. arXiv preprint arXiv:1312.4400, 2013.
Shanchuan Lin, Bingchen Liu, Jiashi Li, and Xiao Yang. Common diffusion noise schedules and sample steps are flawed. In IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). 2024. URL: https://arxiv.org/abs/2305.08891.
Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2017.
Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár. Focal loss for dense object detection. In Proceedings of the IEEE International Conference on Computer Vision (ICCV). 2017.
Zhiqiu Lin, Deepak Pathak, Baiqi Li, Jiayao Li, Xide Xia, Graham Neubig, Pengchuan Zhang, and Deva Ramanan. Evaluating text-to-visual generation with image-to-text generation. In European Conference on Computer Vision (ECCV). 2024.
Greg Linden, Brent Smith, and Jeremy York. Amazon.com recommendations: item-to-item collaborative filtering. IEEE Internet Computing, 7(1):76–80, 2003. doi:10.1109/MIC.2003.1167344.
Jack Lindsey, Wes Gurnee, Emmanuel Ameisen, Brian Chen, Adam Pearce, Nicholas L. Turner, Craig Citro, David Abrahams, Shan Carter, Basil Hosmer, Jonathan Marcus, Michael Sklar, Adly Templeton, Trenton Bricken, Callum McDougall, Hoagy Cunningham, Thomas Henighan, Adam Jermyn, Andy Jones, Andrew Persic, Zhenyi Qi, T. Ben Thompson, Sam Zimmerman, Kelley Rivoire, Thomas Conerly, Chris Olah, and Joshua Batson. On the biology of a large language model. Transformer Circuits Thread, Anthropic, 2025. URL: https://transformer-circuits.pub/2025/attribution-graphs/biology.html.
Seppo Linnainmaa. Algoritmin kumulatiivinen pyöristysvirhe yksittäisten pyöristysvirheiden Taylor-kehitelmänä. Master's thesis, University of Helsinki, 1970. In finlandese; versione inglese in BIT 16(2):146–160, 1976.
Jacques-Louis Lions. Ariane 5: flight 501 failure. report by the inquiry board. Technical Report, European Space Agency and Centre National d'Études Spatiales, Paris, 7 1996. 19 luglio 1996.
Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. Flow matching for generative modeling. In International Conference on Learning Representations (ICLR). 2023. URL: https://arxiv.org/abs/2210.02747.
Zachary C. Lipton. The mythos of model interpretability. Queue, 16(3):31–57, 2018. doi:10.1145/3236386.3241340.
Zachary C. Lipton, Yu-Xiang Wang, and Alexander J. Smola. Detecting and correcting for label shift with black box predictors. In Proceedings of the 35th International Conference on Machine Learning (ICML). 2018. URL: https://arxiv.org/abs/1802.03916, arXiv:1802.03916.
Christian List and Robert E. Goodin. Epistemic democracy: generalizing the Condorcet jury theorem. Journal of Political Philosophy, 9(3):277–306, 2001. doi:10.1111/1467-9760.00128.
John D. C. Little. A proof for the queuing formula: $L = \lambda W$. Operations Research, 9(3):383–387, 1961. doi:10.1287/opre.9.3.383.
W. A. Little. The existence of persistent states in the brain. Mathematical Biosciences, 19(1–2):101–120, 1974. doi:10.1016/0025-5564(74)90031-5.
Bo Liu, Rui Wang, Lemeng Wu, Yihao Feng, Peter Stone, and Qiang Liu. Longhorn: state space models are amortized online learners. In International Conference on Learning Representations (ICLR). 2025. URL: https://openreview.net/forum?id=8jOqCcLzeO, arXiv:2407.14207.
Chia-Wei Liu, Ryan Lowe, Iulian V. Serban, Michael Noseworthy, Laurent Charlin, and Joelle Pineau. How NOT to evaluate your dialogue system: an empirical study of unsupervised evaluation metrics for dialogue response generation. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2122–2132. Austin, Texas, 2016. Association for Computational Linguistics. doi:10.18653/v1/D16-1230.
Haohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei, Xubo Liu, Danilo Mandic, Wenwu Wang, and Mark D. Plumbley. AudioLDM: text-to-audio generation with latent diffusion models. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, 21450–21474. PMLR, 2023.
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2024.
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. In Advances in Neural Information Processing Systems, volume 36. 2023.
Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang. Is your code generated by ChatGPT really correct? rigorous evaluation of large language models for code generation. In Advances in Neural Information Processing Systems (NeurIPS), volume 36. 2023. URL: https://arxiv.org/abs/2305.01210.
Liyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Jiawei Han. On the variance of the adaptive learning rate and beyond. In International Conference on Learning Representations (ICLR). 2020.
Mingjie Liu, Shizhe Diao, Ximing Lu, Jian Hu, Xin Dong, Yejin Choi, Jan Kautz, and Yi Dong. ProRL: prolonged reinforcement learning expands reasoning boundaries in large language models. In Advances in Neural Information Processing Systems, volume 38, 17998–18031. 2025. doi:10.52202/085713-0608.
Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. Lost in the middle: how language models use long contexts. Transactions of the Association for Computational Linguistics (TACL), 12:157–173, 2024.
Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, and Alexander C. Berg. SSD: single shot multibox detector. In European Conference on Computer Vision, 21–37. 2016.
Weitang Liu, Xiaoyun Wang, John Owens, and Yixuan Li. Energy-based out-of-distribution detection. In Advances in Neural Information Processing Systems 33 (NeurIPS). 2020.
Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, Shudan Zhang, Xiang Deng, Aohan Zeng, Zhengxiao Du, Chenhui Zhang, Sheng Shen, Tianjun Zhang, Yu Su, Huan Sun, Minlie Huang, Yuxiao Dong, and Jie Tang. AgentBench: evaluating LLMs as agents. In International Conference on Learning Representations (ICLR). 2024.
Xingchao Liu, Chengyue Gong, and Qiang Liu. Flow straight and fast: learning to generate and transfer data with rectified flow. In International Conference on Learning Representations (ICLR). 2023. arXiv:2209.03003.
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. RoBERTa: a robustly optimized BERT pretraining approach. arXiv preprint arXiv:1907.11692, 2019.
Zhuang Liu, Mingjie Sun, Tinghui Zhou, Gao Huang, and Trevor Darrell. Rethinking the value of network pruning. In International Conference on Learning Representations (ICLR). 2019.
Zihan Liu, Genta Indra Winata, Andrea Madotto, and Pascale Fung. Exploring fine-tuning techniques for pre-trained cross-lingual models via continual learning. arXiv preprint arXiv:2004.14218, 2020. URL: https://arxiv.org/abs/2004.14218, arXiv:2004.14218.
Ziming Liu, Eric J. Michaud, and Max Tegmark. Omnigrok: grokking beyond algorithmic data. In International Conference on Learning Representations (ICLR). 2023.
Zirui Liu, Jiayi Yuan, Hongye Jin, Shaochen Zhong, Zhaozhuo Xu, Vladimir Braverman, Beidi Chen, and Xia Hu. KIVI: a tuning-free asymmetric 2bit quantization for KV cache. In Proceedings of the 41st International Conference on Machine Learning (ICML). 2024.
Greta M. Ljung and George E. P. Box. On a measure of lack of fit in time series models. Biometrika, 65(2):297–303, 1978.
Stuart P. Lloyd. Least squares quantization in pcm. IEEE Transactions on Information Theory, 28(2):129–137, 1982.
Gabriel Loaiza-Ganem and John P. Cunningham. The continuous Bernoulli: fixing a pervasive error in variational autoencoders. In Advances in Neural Information Processing Systems (NeurIPS), volume 32. 2019. URL: https://arxiv.org/abs/1907.06845.
Francesco Locatello, Stefan Bauer, Mario Lucic, Gunnar Rätsch, Sylvain Gelly, Bernhard Schölkopf, and Olivier Bachem. Challenging common assumptions in the unsupervised learning of disentangled representations. In Proceedings of the 36th International Conference on Machine Learning (ICML). 2019. URL: https://arxiv.org/abs/1811.12359.
Jonathan Long, Evan Shelhamer, and Trevor Darrell. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2015.
H. Christopher Longuet-Higgins. A computer algorithm for reconstructing a scene from two projections. Nature, 293(5828):133–135, 1981.
David Lopez-Paz and Maxime Oquab. Revisiting classifier two-sample tests. In International Conference on Learning Representations (ICLR). 2017. URL: https://arxiv.org/abs/1610.06545.
Gary Lorden. Procedures for reacting to a change in distribution. The Annals of Mathematical Statistics, 42(6):1897–1908, 1971.
Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In International Conference on Learning Representations. 2019.
Aaron Lou, Chenlin Meng, and Stefano Ermon. Discrete diffusion modeling by estimating the ratios of the data distribution. In International Conference on Machine Learning (ICML). 2024. URL: https://arxiv.org/abs/2310.16834.
Yin Lou, Rich Caruana, Johannes Gehrke, and Giles Hooker. Accurate intelligible models with pairwise interactions. In ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD). 2013.
David G. Lowe. Distinctive image features from scale-invariant keypoints. International Journal of Computer Vision, 60(2):91–110, 2004.
Ryan Lowe, Nissan Pow, Iulian Serban, and Joelle Pineau. The Ubuntu dialogue corpus: a large dataset for research in unstructured multi-turn dialogue systems. In Proceedings of the 16th Annual Meeting of the Special Interest Group on Discourse and Dialogue (SIGDIAL), 285–294. 2015.
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch. Multi-agent actor-critic for mixed cooperative-competitive environments. In Advances in Neural Information Processing Systems 30 (NIPS). 2017.
Cheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen, Chongxuan Li, and Jun Zhu. Dpm-solver++: fast solver for guided sampling of diffusion probabilistic models. arXiv preprint arXiv:2211.01095, 2022. URL: https://arxiv.org/abs/2211.01095.
Cheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen, Chongxuan Li, and Jun Zhu. Dpm-solver: a fast ode solver for diffusion probabilistic model sampling in around 10 steps. In Advances in Neural Information Processing Systems (NeurIPS). 2022. URL: https://arxiv.org/abs/2206.00927.
Lu Lu, Pengzhan Jin, Guofei Pang, Zhongqiang Zhang, and George Em Karniadakis. Learning nonlinear operators via DeepONet based on the universal approximation theorem of operators. Nature Machine Intelligence, 3(3):218–229, 2021. doi:10.1038/s42256-021-00302-5.
Lu Lu, Xuhui Meng, Zhiping Mao, and George Em Karniadakis. DeepXDE: a deep learning library for solving differential equations. SIAM Review, 63(1):208–228, 2021. doi:10.1137/19M1274067.
Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. Fantastically ordered prompts and where to find them: overcoming few-shot prompt order sensitivity. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL), 8086–8098. 2022.
Bruce D. Lucas and Takeo Kanade. An iterative image registration technique with an application to stereo vision. In Proceedings of the 7th International Joint Conference on Artificial Intelligence (IJCAI), 674–679. 1981.
James Lucas, George Tucker, Roger Grosse, and Mohammad Norouzi. Don't blame the ELBO! A linear VAE perspective on posterior collapse. In Advances in Neural Information Processing Systems 32 (NeurIPS). 2019. URL: https://arxiv.org/abs/1911.02469.
John R. Lucas. Minds, machines and Gödel. Philosophy, 36(137):112–127, 1961.
Alexandra Sasha Luccioni, Sylvain Viguier, and Anne-Laure Ligozat. Estimating the carbon footprint of BLOOM, a 176B parameter language model. Journal of Machine Learning Research, 24(253):1–15, 2023. URL: http://jmlr.org/papers/v24/23-0069.html.
Sasha Luccioni, Yacine Jernite, and Emma Strubell. Power hungry processing: watts driving the cost of AI deployment? In ACM Conference on Fairness, Accountability, and Transparency (FAccT). 2024. URL: https://arxiv.org/abs/2311.16863.
Kristian Lum and William Isaac. To predict and serve? Significance, 13(5):14–19, October 2016. doi:10.1111/j.1740-9713.2016.00960.x.
Scott M. Lundberg, Gabriel Erion, Hugh Chen, Alex DeGrave, Jordan M. Prutkin, Bala Nair, Ronit Katz, Jonathan Himmelfarb, Nisha Bansal, and Su-In Lee. From local explanations to global understanding with explainable AI for trees. Nature Machine Intelligence, 2:56–67, 2020.
Scott M. Lundberg and Su-In Lee. A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems (NeurIPS). 2017. URL: https://arxiv.org/abs/1705.07874.
Weile Luo, Ruibo Fan, Zeyu Li, Dayou Du, Qiang Wang, and Xiaowen Chu. Benchmarking and dissecting the Nvidia Hopper GPU architecture. In 2024 IEEE International Parallel and Distributed Processing Symposium (IPDPS), 656–667. 2024. doi:10.1109/IPDPS57955.2024.00064.
Thang Luong, Hieu Pham, and Christopher D. Manning. Effective approaches to attention-based neural machine translation. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, 1412–1421. 2015. doi:10.18653/v1/D15-1166.
Timothy W. Lyons, Christopher T. Reinhard, and Noah J. Planavsky. The rise of oxygen in earth's early ocean and atmosphere. Nature, 506(7488):307–315, 2014. doi:10.1038/nature13068.
Marcos López de Prado. Advances in Financial Machine Learning. Wiley, 2018. ISBN 978-1-119-48208-6. Capitolo 7, «Cross-Validation in Finance».
Jerry Ma and Denis Yarats. On the adequacy of untuned warmup for adaptive optimization. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, 8828–8836. 2021.
Monte MacDiarmid, Timothy Maxwell, Nicholas Schiefer, Jesse Mu, Jared Kaplan, David Duvenaud, Sam Bowman, Alex Tamkin, Ethan Perez, Mrinank Sharma, Carson Denison, and Evan Hubinger. Simple probes can catch sleeper agents. Anthropic Research, 2024. URL: https://www.anthropic.com/research/probes-catch-sleeper-agents.
Ashwin Machanavajjhala, Daniel Kifer, Johannes Gehrke, and Muthuramakrishnan Venkitasubramaniam. ℓ-diversity: privacy beyond $k$-anonymity. ACM Transactions on Knowledge Discovery from Data, 1(1):3, 2007. doi:10.1145/1217299.1217302.
James MacQueen. Some methods for classification and analysis of multivariate observations. In Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability, volume 1, 281–297. 1967.
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark. Self-Refine: iterative refinement with self-feedback. In Advances in Neural Information Processing Systems (NeurIPS), volume 36, 46534–46594. 2023. URL: https://arxiv.org/abs/2303.17651.
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. Towards deep learning models resistant to adversarial attacks. In International Conference on Learning Representations (ICLR). 2018. URL: https://arxiv.org/abs/1706.06083.
Matthew V. Mahoney. Text compression as a test for artificial intelligence. In Proceedings of the Sixteenth National Conference on Artificial Intelligence (AAAI), 970. 1999.
Spyros Makridakis, Evangelos Spiliotis, and Vassilios Assimakopoulos. The m4 competition: 100,000 time series and 61 forecasting methods. International Journal of Forecasting, 36(1):54–74, 2020.
Spyros Makridakis, Evangelos Spiliotis, and Vassilios Assimakopoulos. M5 accuracy competition: results, findings, and conclusions. International Journal of Forecasting, 38(4):1346–1364, 2022.
Yury A. Malkov and Dmitry A. Yashunin. Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs. IEEE Transactions on Pattern Analysis and Machine Intelligence, 42(4):824–836, 2020. doi:10.1109/TPAMI.2018.2889473.
Sadhika Malladi, Kaifeng Lyu, Abhishek Panigrahi, and Sanjeev Arora. On the SDEs and scaling rules for adaptive gradient algorithms. In Advances in Neural Information Processing Systems 35 (NeurIPS). 2022. arXiv:2205.10287.
Massimiliano Marcellino, James H. Stock, and Mark W. Watson. A comparison of direct and iterated multistep AR methods for forecasting macroeconomic time series. Journal of Econometrics, 135(1–2):499–526, 2006. doi:10.1016/j.jeconom.2005.07.020.
Marco Marchesi. Megapixel size image creation using generative adversarial networks. arXiv preprint arXiv:1706.00082, 2017.
Mitchell P. Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz. Building a large annotated corpus of English: the penn treebank. Computational Linguistics, 19(2):313–330, 1993. URL: https://aclanthology.org/J93-2004/.
Andrei Andreevich Markov. Essai d'une recherche statistique sur le texte du roman "eugène onéguine", illustrant la liaison des épreuves en chaîne. Izvestiya Imperatorskoi Akademii Nauk (Bulletin de l'Académie Impériale des Sciences de St.-Pétersbourg), VI serie, 7(3):153–162, 1913. Traduzione inglese: "An Example of Statistical Investigation of the Text Eugene Onegin Concerning the Connection of Samples in Chains", Science in Context, 19(4):591–600, 2006, doi:10.1017/S0269889706001074.
Benjamin M. Marlin and Richard S. Zemel. Collaborative prediction and ranking with non-random missing data. In Proceedings of the Third ACM Conference on Recommender Systems (RecSys '09), 5–12. 2009. doi:10.1145/1639714.1639717.
Eric Martin and Chris Cundy. Parallelizing linear recurrent neural nets over sequence length. In International Conference on Learning Representations (ICLR). 2018.
Eric Masanet, Arman Shehabi, Nuoa Lei, Sarah Smith, and Jonathan Koomey. Recalibrating global data center energy-use estimates. Science, 367(6481):984–986, 2020.
Michaël Mathieu, Mikael Henaff, and Yann LeCun. Fast training of convolutional networks through FFTs. In International Conference on Learning Representations (ICLR). 2014. arXiv:1312.5851.
Andreas Maurer, Massimiliano Pontil, and Bernardino Romera-Paredes. The benefit of multitask representation learning. Journal of Machine Learning Research, 17(81):1–32, 2016.
Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David Forsyth, and Dan Hendrycks. HarmBench: a standardized evaluation framework for automated red teaming and robust refusal. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, 35181–35224. PMLR, 2024. URL: https://proceedings.mlr.press/v235/mazeika24a.html.
David McAllester and Karl Stratos. Formal limitations on the measurement of mutual information. In International Conference on Artificial Intelligence and Statistics (AISTATS). 2020. URL: https://arxiv.org/abs/1811.04251.
David A. McAllester. Some PAC-Bayesian theorems. Machine Learning, 37(3):355–363, 1999. doi:10.1023/A:1007618624809.
Andrew McCallum, Dayne Freitag, and Fernando C. N. Pereira. Maximum entropy Markov models for information extraction and segmentation. In Pat Langley, editor, Proceedings of the Seventeenth International Conference on Machine Learning (ICML 2000), 591–598. Morgan Kaufmann, 2000.
Andrew McCallum and Kamal Nigam. A comparison of event models for naive Bayes text classification. In Learning for Text Categorization: Papers from the 1998 AAAI Workshop, number WS-98-05 in AAAI Technical Report, 41–48. 1998. URL: https://aaai.org/papers/041-ws98-05-007/.
Levi D. McClenny and Ulisses M. Braga-Neto. Self-adaptive physics-informed neural networks. Journal of Computational Physics, 474:111722, 2023. URL: https://arxiv.org/abs/2009.04544.
Michael McCloskey and Neal J. Cohen. Catastrophic interference in connectionist networks: the sequential learning problem. Psychology of Learning and Motivation, 24:109–165, 1989.
Warren S. McCulloch and Walter Pitts. A logical calculus of the ideas immanent in nervous activity. The Bulletin of Mathematical Biophysics, 5:115–133, 1943.
Ryan McDonald, Fernando Pereira, Kiril Ribarov, and Jan Hajič. Non-projective dependency parsing using spanning tree algorithms. In Proceedings of Human Language Technology Conference and Conference on Empirical Methods in Natural Language Processing (HLT/EMNLP), 523–530. 2005.
Robert J. McEliece, Edward C. Posner, Eugene R. Rodemich, and Santosh S. Venkatesh. The capacity of the Hopfield associative memory. IEEE Transactions on Information Theory, 33(4):461–482, 1987.
Leonard A. McGee and Stanley F. Schmidt. Discovery of the Kalman filter as a practical tool for aerospace and industry. NASA Technical Memorandum 86847, NASA Ames Research Center, nov 1985.
Amy McGovern and Andrew G. Barto. Automatic discovery of subgoals in reinforcement learning using diverse density. In Proceedings of the 18th International Conference on Machine Learning (ICML). 2001. URL: https://hdl.handle.net/20.500.14394/10400.
Nick McGreivy and Ammar Hakim. Weak baselines and reporting biases lead to overoptimism in machine learning for fluid-related partial differential equations. Nature Machine Intelligence, 6(10):1256–1269, 2024. doi:10.1038/s42256-024-00897-5.
Leland McInnes, John Healy, and James Melville. Umap: uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426, 2018. URL: https://arxiv.org/abs/1802.03426.
Brandon McKinzie, Zhe Gan, Jean-Philippe Fauconnier, and others. MM1: methods, analysis & insights from multimodal LLM pre-training. arXiv preprint arXiv:2403.09611, 2024.
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Agüera y Arcas. Communication-efficient learning of deep networks from decentralized data. In Artificial Intelligence and Statistics (AISTATS). 2017. URL: https://arxiv.org/abs/1602.05629.
Quinn McNemar. Note on the sampling error of the difference between correlated proportions or percentages. Psychometrika, 12(2):153–157, 1947.
Cole Medin. Context-engineering-intro. Repository GitHub, 2025. URL: coleam00/context-engineering-intro.
Ninareh Mehrabi, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. A survey on bias and fairness in machine learning. ACM Computing Surveys, 54(6):1–35, 2021. URL: https://arxiv.org/abs/1908.09635.
Lingrui Mei, Jiayu Yao, Yuyao Ge, Yiwei Wang, Baolong Bi, Yujun Cai, Jiazhi Liu, Mingyu Li, Zhong-Zhi Li, Duzhen Zhang, Chenlin Zhou, Jiayi Mao, Tianze Xia, Jiafeng Guo, and Shenghua Liu. A survey of context engineering for large language models. arXiv preprint arXiv:2507.13334, 2025.
Lennart Meincke, Ethan Mollick, Lilach Mollick, and Dan Shapiro. Prompting Science Report 1: prompt engineering is complicated and contingent. The Wharton School, Generative AI Labs, arXiv:2503.04818, 4 marzo 2025, 2025. URL: https://arxiv.org/abs/2503.04818.
Lennart Meincke, Ethan Mollick, Lilach Mollick, and Dan Shapiro. Prompting Science Report 2: the decreasing value of chain of thought in prompting. The Wharton School, Generative AI Labs, arXiv:2506.07142, 8 giugno 2025, 2025. URL: https://arxiv.org/abs/2506.07142.
Lennart Meincke, Ethan Mollick, Lilach Mollick, and Dan Shapiro. Prompting Science Report 3: i'll pay you or i'll kill you, but will you care? The Wharton School, Generative AI Labs, arXiv:2508.00614, 1 agosto 2025, 2025. URL: https://arxiv.org/abs/2508.00614.
Luigi Federico Menabrea. Sketch of the analytical engine invented by charles babbage. Scientific Memoirs, 3:666–731, 1843. Tradotto dal francese (Bibliothèque universelle de Genève, 1842), con le note A-G, da Ada Augusta, contessa di Lovelace.
Chenlin Meng, Robin Rombach, Ruiqi Gao, Diederik P. Kingma, Stefano Ermon, Jonathan Ho, and Tim Salimans. On distillation of guided diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 14297–14306. 2023. URL: https://arxiv.org/abs/2210.03142.
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. Locating and editing factual associations in GPT. In Advances in Neural Information Processing Systems (NeurIPS). 2022. URL: https://arxiv.org/abs/2202.05262.
Fabian Mentzer, David Minnen, Eirikur Agustsson, and Michael Tschannen. Finite scalar quantization: VQ-VAE made simple. In International Conference on Learning Representations (ICLR). 2024.
Bernard Merialdo. Tagging English text with a probabilistic model. Computational Linguistics, 20(2):155–171, 1994. URL: https://aclanthology.org/J94-2001/.
William Merrill, Jackson Petty, and Ashish Sabharwal. The illusion of state in state-space models. In Proceedings of the 41st International Conference on Machine Learning (ICML), volume 235 of Proceedings of Machine Learning Research, 35492–35506. PMLR, 2024. arXiv:2404.08819.
William Merrill and Ashish Sabharwal. The parallelism tradeoff: limitations of log-precision transformers. Transactions of the Association for Computational Linguistics, 11:531–545, 2023. doi:10.1162/tacl_a_00562.
William Merrill and Ashish Sabharwal. The expressive power of transformers with chain of thought. In International Conference on Learning Representations (ICLR), 7690–7706. 2024. URL: https://openreview.net/forum?id=NjNGlPh8Wh.
Annamaria Mesaros, Toni Heittola, and Tuomas Virtanen. Metrics for polyphonic sound event detection. Applied Sciences, 6(6):162, 2016. doi:10.3390/app6060162.
Lars Mescheder, Andreas Geiger, and Sebastian Nowozin. Which training methods for GANs do actually converge? In Jennifer Dy and Andreas Krause, editors, Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Machine Learning Research, 3481–3490. PMLR, 2018. URL: https://proceedings.mlr.press/v80/mescheder18a.html.
Nicholas Metropolis. The beginning of the Monte Carlo method. Los Alamos Science, pages 125–130, 1987. Special Issue dedicated to Stanisław Ulam.
Luke Metz, Ben Poole, David Pfau, and Jascha Sohl-Dickstein. Unrolled generative adversarial networks. In International Conference on Learning Representations (ICLR). 2017. URL: https://arxiv.org/abs/1611.02163.
H. N. Mhaskar. Neural networks for optimal approximation of smooth and analytic functions. Neural Computation, 8(1):164–177, 1996. doi:10.1162/neco.1996.8.1.164.
Xupeng Miao, Chunan Shi, Jiangfei Duan, Xiaoli Xi, Dahua Lin, Bin Cui, and Zhihao Jia. SpotServe: serving generative large language models on preemptible instances. In Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS). 2024. doi:10.1145/3620665.3640411.
Paul Michel, Omer Levy, and Graham Neubig. Are sixteen heads really better than one? In Advances in Neural Information Processing Systems 32 (NeurIPS 2019), 14014–14024. 2019. URL: https://arxiv.org/abs/1905.10650.
Alessio Micheli. Neural network for graphs: a contextual constructive approach. IEEE Transactions on Neural Networks, 20(3):498–511, 2009.
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu. Mixed precision training. In International Conference on Learning Representations. 2018.
Paulius Micikevicius, Dusan Stosic, Neil Burgess, Marius Cornea, Pradeep Dubey, Richard Grisenthwaite, Sangwon Ha, Alexander Heinecke, Patrick Judd, John Kamalu, Naveen Mellempudi, Stuart Oberman, Mohammad Shoeybi, Michael Siu, and Hao Wu. FP8 formats for deep learning. arXiv preprint arXiv:2209.05433, 2022.
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. Efficient estimation of word representations in vector space. In International Conference on Learning Representations (ICLR), Workshop Track. 2013.
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S. Corrado, and Jeffrey Dean. Distributed representations of words and phrases and their compositionality. In Advances in Neural Information Processing Systems 26 (NIPS), 3111–3119. 2013.
Tomas Mikolov, Wen-tau Yih, and Geoffrey Zweig. Linguistic regularities in continuous space word representations. In Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), 746–751. 2013. URL: https://aclanthology.org/N13-1090/.
Maxim Milakov and Natalia Gimelshein. Online normalizer calculation for softmax. 2018. URL: https://arxiv.org/abs/1805.02867, arXiv:1805.02867.
Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. NeRF: representing scenes as neural radiance fields for view synthesis. In European Conference on Computer Vision (ECCV). 2020. URL: https://arxiv.org/abs/2003.08934.
Tim Miller. Explanation in artificial intelligence: insights from the social sciences. Artificial Intelligence, 267:1–38, 2019. doi:10.1016/j.artint.2018.07.007.
Beren Millidge, Alexander Tschantz, and Christopher L. Buckley. Whence the expected free energy? Neural Computation, 33(2):447–482, 2021. URL: https://arxiv.org/abs/2004.08128.
Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. Rethinking the role of demonstrations: what makes in-context learning work? In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP), 11048–11064. 2022.
Marvin Minsky and Seymour Papert. Perceptrons: An Introduction to Computational Geometry. MIT Press, 1969.
Ilya Mironov. Rényi differential privacy. In 2017 IEEE 30th Computer Security Foundations Symposium (CSF), 263–275. 2017. arXiv:1702.07476, doi:10.1109/CSF.2017.11.
Leon Mirsky. Symmetric gauge functions and unitarily invariant norms. The Quarterly Journal of Mathematics, 11(1):50–59, 1960. doi:10.1093/qmath/11.1.50.
Mehdi Mirza and Simon Osindero. Conditional generative adversarial nets. arXiv preprint arXiv:1411.1784, 2014.
Iman Mirzadeh, Keivan Alizadeh, Hooman Shahrokhi, Oncel Tuzel, Samy Bengio, and Mehrdad Farajtabar. GSM-symbolic: understanding the limitations of mathematical reasoning in large language models. In International Conference on Learning Representations (ICLR). 2025. URL: https://arxiv.org/abs/2410.05229, arXiv:2410.05229.
Asit Mishra, Jorge Albericio Latorre, Jeff Pool, Darko Stosic, Dusan Stosic, Ganesh Venkatesh, Chong Yu, and Paulius Micikevicius. Accelerating sparse deep neural networks. arXiv preprint arXiv:2104.08378, 2021.
Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. Model cards for model reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency (FAT*), 220–229. 2019. doi:10.1145/3287560.3287596.
Tom M. Mitchell. Machine Learning. McGraw-Hill Series in Computer Science. McGraw-Hill, 1997. ISBN 978-0-07-042807-2.
Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida. Spectral normalization for generative adversarial networks. In International Conference on Learning Representations (ICLR). 2018.
Volodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu. Asynchronous methods for deep reinforcement learning. In Proceedings of the 33rd International Conference on Machine Learning (ICML). 2016.
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller. Playing atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602, 2013.
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis. Human-level control through deep reinforcement learning. Nature, 518:529–533, 2015.
Mehryar Mohri, Afshin Rostamizadeh, and Ameet Talwalkar. Foundations of Machine Learning. MIT Press, Cambridge, MA, 2 edition, 2018.
Christoph Molnar. Interpretable Machine Learning: A Guide for Making Black Box Models Explainable. Edizione dell'autore, second edition, 2022. URL: https://christophm.github.io/interpretable-ml-book/.
Guido Montúfar, Razvan Pascanu, Kyunghyun Cho, and Yoshua Bengio. On the number of linear regions of deep neural networks. In Advances in Neural Information Processing Systems (NIPS), volume 27, 2924–2932. 2014.
Hans Moravec. Mind Children: The Future of Robot and Human Intelligence. Harvard University Press, 1988.
Jose G. Moreno-Torres, Troy Raeder, Rocío Alaiz-Rodríguez, Nitesh V. Chawla, and Francisco Herrera. A unifying view on dataset shift in classification. Pattern Recognition, 45(1):521–530, 2012.
Christopher Morris, Martin Ritzert, Matthias Fey, William L. Hamilton, Jan Eric Lenssen, Gaurav Rattan, and Martin Grohe. Weisfeiler and leman go neural: higher-order graph neural networks. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, 4602–4609. 2019. doi:10.1609/aaai.v33i01.33014602.
Frederick Mosteller and David L. Wallace. Inference and Disputed Authorship: The Federalist. Addison-Wesley, Reading, Massachusetts, 1964.
Saman Motamed, Laura Culp, Kevin Swersky, Priyank Jaini, and Robert Geirhos. Do generative video models understand physical principles? arXiv preprint arXiv:2501.09038, 2025. URL: https://arxiv.org/abs/2501.09038.
Ramaravind K. Mothilal, Amit Sharma, and Chenhao Tan. Explaining machine learning classifiers through diverse counterfactual explanations. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency (FAT*), 607–617. 2020. doi:10.1145/3351095.3372850.
Davoud Moulavi, Pablo A. Jaskowiak, Ricardo J. G. B. Campello, Arthur Zimek, and Jörg Sander. Density-based clustering validation. In Proceedings of the 2014 SIAM International Conference on Data Mining (SDM), 839–847. 2014.
George V. Moustakides. Optimal stopping times for detecting changes in distributions. The Annals of Statistics, 14(4):1379–1387, 1986.
Allan H. Murphy. A new vector partition of the probability score. Journal of Applied Meteorology, 12(4):595–600, 1973.
Rafael Müller, Simon Kornblith, and Geoffrey Hinton. When does label smoothing help? In Advances in Neural Information Processing Systems, volume 32. 2019.
Thomas Müller, Alex Evans, Christoph Schied, and Alexander Keller. Instant neural graphics primitives with a multiresolution hash encoding. ACM Transactions on Graphics, 41(4):102:1–102:15, 2022. URL: https://arxiv.org/abs/2201.05989.
Èlizbar A. Nadaraya. On estimating regression. Theory of Probability and Its Applications, 9(1):141–142, 1964.
Mahdi Pakdaman Naeini, Gregory F. Cooper, and Milos Hauskrecht. Obtaining well calibrated probabilities using bayesian binning. In Proceedings of the Twenty-Ninth AAAI Conference on Artificial Intelligence, 2901–2907. 2015. doi:10.1609/aaai.v29i1.9602.
Vinod Nair and Geoffrey E. Hinton. Rectified linear units improve restricted Boltzmann machines. In Proceedings of the 27th International Conference on Machine Learning (ICML). 2010.
Kaoru Nakano. Associatron—a model of associative memory. IEEE Transactions on Systems, Man, and Cybernetics, SMC-2(3):380–388, 1972. doi:10.1109/TSMC.1972.4309133.
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever. Deep double descent: where bigger models and more data hurt. In International Conference on Learning Representations (ICLR). 2020. arXiv:1912.02292.
Preetum Nakkiran, Prayaag Venkat, Sham Kakade, and Tengyu Ma. Optimal regularization can mitigate double descent. In International Conference on Learning Representations (ICLR). 2021. arXiv:2003.01897.
Eric Nalisnick, Akihiro Matsukawa, Yee Whye Teh, Dilan Gorur, and Balaji Lakshminarayanan. Do deep generative models know what they don't know? In International Conference on Learning Representations (ICLR). 2019.
Eric Nalisnick, Akihiro Matsukawa, Yee Whye Teh, and Balaji Lakshminarayanan. Detecting out-of-distribution inputs to deep generative models using typicality. arXiv preprint arXiv:1906.02994, 2019.
Hyeonuk Nam, Seong-Hu Kim, Byeong-Yun Ko, and Yong-Hwa Park. Frequency dynamic convolution: frequency-adaptive pattern recognition for sound event detection. In Interspeech, 2763–2767. 2022. doi:10.21437/Interspeech.2022-10127.
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt. Progress measures for grokking via mechanistic interpretability. In International Conference on Learning Representations (ICLR). 2023. URL: https://arxiv.org/abs/2301.05217, arXiv:2301.05217.
Neel Nanda, Andrew Lee, and Martin Wattenberg. Emergent linear representations in world models of self-supervised sequence models. arXiv preprint arXiv:2309.00941, 2023. URL: https://arxiv.org/abs/2309.00941, arXiv:2309.00941.
Arvind Narayanan. Fact checking Moravec's paradox. AI as Normal Technology, 2026. Consultato il 16 agosto 2026. URL: https://www.normaltech.ai/p/fact-checking-moravecs-paradox.
Arvind Narayanan and Vitaly Shmatikov. Robust de-anonymization of large sparse datasets. In 2008 IEEE Symposium on Security and Privacy (SP 2008), 111–125. 2008. URL: https://arxiv.org/abs/cs/0610105.
Deepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley, Mostofa Patwary, Vijay Korthikanti, Dmitri Vainbrand, Prethvi Kashinkunti, Julie Bernauer, Bryan Catanzaro, Amar Phanishayee, and Matei Zaharia. Efficient large-scale language model training on GPU clusters using Megatron-LM. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC '21), 1–15. 2021. doi:10.1145/3458817.3476209.
Milad Nasr, Nicholas Carlini, Chawin Sitawarin, Sander V. Schulhoff, Jamie Hayes, Michael Ilie, Juliette Pluto, Shuang Song, Harsh Chaudhari, Ilia Shumailov, Abhradeep Thakurta, Kai Yuanqing Xiao, Andreas Terzis, and Florian Tramèr. The attacker moves second: stronger adaptive attacks bypass defenses against LLM jailbreaks and prompt injections. 2025. URL: https://arxiv.org/abs/2510.09023, arXiv:2510.09023.
Saul B. Needleman and Christian D. Wunsch. A general method applicable to the search for similarities in the amino acid sequence of two proteins. Journal of Molecular Biology, 48(3):443–453, 1970. doi:10.1016/0022-2836(70)90057-4.
John A. Nelder and Robert W. M. Wedderburn. Generalized linear models. Journal of the Royal Statistical Society: Series A (General), 135(3):370–384, 1972. doi:10.2307/2344614.
George L. Nemhauser, Laurence A. Wolsey, and Marshall L. Fisher. An analysis of approximations for maximizing submodular set functions—I. Mathematical Programming, 14(1):265–294, 1978. doi:10.1007/BF01588971.
Yurii Nesterov and Vladimir Spokoiny. Random gradient-free minimization of convex functions. Foundations of Computational Mathematics, 17(2):527–566, 2017. doi:10.1007/s10208-015-9296-2.
Jerzy Neyman and Egon S. Pearson. On the problem of the most efficient tests of statistical hypotheses. Philosophical Transactions of the Royal Society of London A, 231:289–337, 1933.
Andrew Ng. Artificial intelligence is the new electricity. Intervento al programma MSx, Stanford Graduate School of Business; resoconto di Shana Lynch, Stanford GSB Insights, 3 2017. URL: https://www.gsb.stanford.edu/insights/andrew-ng-why-ai-new-electricity.
Andrew Y. Ng. Feature selection, L1 vs. L2 regularization, and rotational invariance. In Proceedings of the 21st International Conference on Machine Learning (ICML). 2004.
Andrew Y. Ng, Daishi Harada, and Stuart Russell. Policy invariance under reward transformations: theory and application to reward shaping. In Proceedings of the Sixteenth International Conference on Machine Learning, 278–287. Morgan Kaufmann, 1999.
Andrew Y. Ng and Michael I. Jordan. On discriminative vs. generative classifiers: a comparison of logistic regression and naive bayes. In Advances in Neural Information Processing Systems 14 (NIPS 2001), 841–848. 2001.
Andrew Y. Ng and Stuart Russell. Algorithms for inverse reinforcement learning. In Proceedings of the Seventeenth International Conference on Machine Learning, 663–670. Morgan Kaufmann, 2000.
Alex Nichol, Joshua Achiam, and John Schulman. On first-order meta-learning algorithms. arXiv preprint arXiv:1803.02999, 2018.
Alexander Quinn Nichol and Prafulla Dhariwal. Improved denoising diffusion probabilistic models. In Proceedings of the 38th International Conference on Machine Learning (ICML), 8162–8171. 2021.
John Nickolls, Ian Buck, Michael Garland, and Kevin Skadron. Scalable parallel programming with CUDA. ACM Queue, 6(2):40–53, 2008.
Alexandru Niculescu-Mizil and Rich Caruana. Predicting good probabilities with supervised learning. In Proceedings of the 22nd International Conference on Machine Learning (ICML), 625–632. 2005. doi:10.1145/1102351.1102430.
Shen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang, Jingyang Ou, Jun Hu, Jun Zhou, Yankai Lin, Ji-Rong Wen, and Chongxuan Li. Large language diffusion models. arXiv preprint arXiv:2502.09992, 2025. URL: https://arxiv.org/abs/2502.09992.
Yuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, and Jayant Kalagnanam. A time series is worth 64 words: long-term forecasting with transformers. In International Conference on Learning Representations (ICLR). 2023. URL: https://arxiv.org/abs/2211.14730.
Sophie J. Nightingale and Hany Farid. AI-synthesized faces are indistinguishable from real faces and more trustworthy. Proceedings of the National Academy of Sciences, 119(8):e2120481119, 2022. doi:10.1073/pnas.2120481119.
Erik Nijkamp, Mitch Hill, Song-Chun Zhu, and Ying Nian Wu. Learning non-convergent non-persistent short-run MCMC toward energy-based model. In Advances in Neural Information Processing Systems (NeurIPS). 2019. URL: https://arxiv.org/abs/1904.09770.
Malvina Nissim, Rik van Noord, and Rob van der Goot. Fair is better than sensational: man is to doctor as woman is to doctor. Computational Linguistics, 46(2):487–497, 2020. URL: https://aclanthology.org/2020.cl-2.7/, doi:10.1162/coli_a_00379.
Joakim Nivre. Non-projective dependency parsing in expected linear time. In Proceedings of the Joint Conference of the 47th Annual Meeting of the ACL and the 4th IJCNLP, 351–359. 2009.
Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Yoav Goldberg, Jan Hajič, Christopher D. Manning, Ryan McDonald, Slav Petrov, Sampo Pyysalo, Natalia Silveira, Reut Tsarfaty, and Daniel Zeman. Universal dependencies v1: a multilingual treebank collection. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC 2016), 1659–1666. Portorož, Slovenia, 2016. European Language Resources Association (ELRA). URL: https://aclanthology.org/L16-1262/.
David A. Nix and Andreas S. Weigend. Estimating the mean and variance of the target probability distribution. In Proceedings of the 1994 IEEE International Conference on Neural Networks (ICNN), volume 1, 55–60. 1994. doi:10.1109/ICNN.1994.374138.
Rodrigo Nogueira, Zhiying Jiang, and Jimmy Lin. Investigating the limitations of transformers with simple arithmetic tasks. arXiv preprint arXiv:2102.13019, 2021. 1st Mathematical Reasoning in General Artificial Intelligence Workshop, ICLR 2021.
Rodrigo Nogueira, Wei Yang, Kyunghyun Cho, and Jimmy Lin. Multi-stage document ranking with BERT. arXiv preprint arXiv:1910.14424, 2019.
Curtis G. Northcutt, Anish Athalye, and Jonas Mueller. Pervasive label errors in test sets destabilize machine learning benchmarks. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, volume 1. 2021. URL: https://datasets-benchmarks-proceedings.neurips.cc/paper/2021/hash/f2217062e9a397a1dca429e7d70bc6ca-Abstract-round1.html.
Curtis G. Northcutt, Lu Jiang, and Isaac L. Chuang. Confident learning: estimating uncertainty in dataset labels. Journal of Artificial Intelligence Research, 70:1373–1411, 2021.
Chris Nota and Philip S. Thomas. Is the policy gradient a gradient? In Proceedings of the 19th International Conference on Autonomous Agents and MultiAgent Systems (AAMAS). 2020. arXiv:1906.07073.
Albert B. J. Novikoff. On convergence proofs on perceptrons. In Proceedings of the Symposium on the Mathematical Theory of Automata, volume 12, 615–622. New York, 1962. Polytechnic Institute of Brooklyn. Volume XII della Microwave Research Institute Symposia Series; atti stampati nel 1963 da Polytechnic Press.
Michael T. Nygard. Release It! Design and Deploy Production-Ready Software. Pragmatic Bookshelf, 2007.
Augustus Odena, Vincent Dumoulin, and Chris Olah. Deconvolution and checkerboard artifacts. Distill, 2016.
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter. Zoom in: an introduction to circuits. Distill, 2020. URL: https://distill.pub/2020/circuits/zoom-in/.
Mikel Olazaran. A sociological study of the official history of the perceptrons controversy. Social Studies of Science, 26(3):611–659, 1996. doi:10.1177/030631296026003005.
Frans A. Oliehoek and Christopher Amato. A Concise Introduction to Decentralized POMDPs. SpringerBriefs in Intelligent Systems. Springer, 2016.
Bruno A. Olshausen and David J. Field. Emergence of simple-cell receptive field properties by learning a sparse code for natural images. Nature, 381(6583):607–609, 1996. doi:10.1038/381607a0.
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, and others. In-context learning and induction heads. Transformer Circuits Thread, 2022. URL: https://arxiv.org/abs/2209.11895.
Isaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang, Tianhao Wu, Joseph E. Gonzalez, M. Waleed Kadous, and Ion Stoica. RouteLLM: learning to route LLMs from preference data. In International Conference on Learning Representations (ICLR). 2025.
Diego Ongaro and John Ousterhout. In search of an understandable consensus algorithm. In Proceedings of the 2014 USENIX Annual Technical Conference (USENIX ATC 14), 305–319. Philadelphia, PA, 2014. USENIX Association.
Kenta Oono and Taiji Suzuki. Graph neural networks exponentially lose expressive power for node classification. In International Conference on Learning Representations (ICLR). 2020.
Sageev Oore, Ian Simon, Sander Dieleman, Douglas Eck, and Karen Simonyan. This time with feeling: learning expressive musical performance. arXiv preprint arXiv:1808.03715, 2018.
Boris N. Oreshkin, Dmitri Carpov, Nicolas Chapados, and Yoshua Bengio. N-beats: neural basis expansion analysis for interpretable time series forecasting. In International Conference on Learning Representations (ICLR). 2020. URL: https://arxiv.org/abs/1905.10437.
Laurent Orseau and Rémi Munos. Super-exponential regret for UCT, AlphaGo and variants. 2024. URL: https://arxiv.org/abs/2405.04407, arXiv:2405.04407.
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy. Deep exploration via bootstrapped DQN. In Advances in Neural Information Processing Systems, volume 29. 2016. URL: https://arxiv.org/abs/1602.04621.
Addy Osmani. Comprehension debt: the hidden cost of AI generated code. addyosmani.com, 14 marzo 2026, 2026. URL: https://addyosmani.com/blog/comprehension-debt/.
Addy Osmani. Loop engineering. addyosmani.com, 7 giugno 2026, 2026. URL: https://addyosmani.com/blog/loop-engineering/.
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, and others. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems, volume 35. 2022.
Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang, Shishir G. Patil, Ion Stoica, and Joseph E. Gonzalez. MemGPT: towards LLMs as operating systems. 2023. arXiv:2310.08560.
Ewan S. Page. Continuous inspection schemes. Biometrika, 41(1/2):100–115, 1954.
Artidoro Pagnoni, Ram Pasunuru, Pedro Rodriguez, John Nguyen, Benjamin Muller, Margaret Li, Chunting Zhou, Lili Yu, Jason Weston, Luke Zettlemoyer, Gargi Ghosh, Mike Lewis, Ari Holtzman, and Srinivasan Iyer. Byte latent transformer: patches scale better than tokens. 2024. URL: https://arxiv.org/abs/2412.09871, arXiv:2412.09871.
Arka Pal, Deep Karkhanis, Samuel Dooley, Manley Roberts, Siddartha Naidu, and Colin White. Smaug: fixing failure modes of preference optimisation with DPO-positive. arXiv preprint arXiv:2402.13228, 2024.
Andrei Paleyes, Raoul-Gabriel Urma, and Neil D. Lawrence. Challenges in deploying machine learning: a survey of case studies. ACM Computing Surveys, 55(6):1–29, 2022.
Bo Pang, Lillian Lee, and Shivakumar Vaithyanathan. Thumbs up? sentiment classification using machine learning techniques. In Proceedings of the 2002 Conference on Empirical Methods in Natural Language Processing (EMNLP), 79–86. 2002. URL: https://aclanthology.org/W02-1011/, doi:10.3115/1118693.1118704.
Arjun Panickssery, Samuel R. Bowman, and Shi Feng. LLM evaluators recognize and favor their own generations. In Advances in Neural Information Processing Systems (NeurIPS), volume 37. 2024. URL: https://proceedings.neurips.cc/paper_files/paper/2024/hash/7f1f0218e45f5414c79c0679633e47bc-Abstract-Conference.html.
George Papamakarios, Theo Pavlakou, and Iain Murray. Masked autoregressive flow for density estimation. In Advances in Neural Information Processing Systems (NeurIPS). 2017. URL: https://arxiv.org/abs/1705.07057.
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. BLEU: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics (ACL), 311–318. Philadelphia, Pennsylvania, USA, 2002. Association for Computational Linguistics. URL: https://aclanthology.org/P02-1040/, doi:10.3115/1073083.1073135.
Angshuman Parashar, Minsoo Rhu, Anurag Mukkara, Antonio Puglielli, Rangharajan Venkatesan, Brucek Khailany, Joel Emer, Stephen W. Keckler, and William J. Dally. SCNN: an accelerator for compressed-sparse convolutional neural networks. In Proceedings of the 44th Annual International Symposium on Computer Architecture (ISCA), 27–40. 2017. doi:10.1145/3079856.3080254.
Eli Pariser. The Filter Bubble: What the Internet Is Hiding from You. Penguin Press, New York, 2011.
Aaron Parisi, Yao Zhao, and Noah Fiedel. TALM: tool augmented language models. 2022. arXiv:2205.12255.
Daniel S. Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D. Cubuk, and Quoc V. Le. SpecAugment: a simple data augmentation method for automatic speech recognition. In Proceedings of Interspeech, 2613–2617. 2019.
Joon Sung Park, Joseph C. O'Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. Generative agents: interactive simulacra of human behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST). 2023.
Joon Sung Park, Carolyn Q. Zou, Jonne Kamphorst, Niles Egan, Aaron Shaw, Benjamin Mako Hill, Carrie Cai, Meredith Ringel Morris, Percy Liang, Robb Willer, and Michael S. Bernstein. LLM agents grounded in self-reports enable general-purpose simulation of individuals. 2026. Versione 3, 28 giugno 2026; la versione 1 (2024) si intitolava «Generative Agent Simulations of 1,000 People». URL: https://arxiv.org/abs/2411.10109, arXiv:2411.10109.
Jack Parker-Holder, Philip Ball, Jake Bruce, Vibhavari Dasagi, Kristian Holsheimer, Christos Kaplanis, Alexandre Moufarek, Guy Scully, Jeremy Shar, Jimmy Shi, Stephen Spencer, Jessica Yung, Michael Dennis, Sultan Kenjeyev, Shangbang Long, Vlad Mnih, Harris Chan, Maxime Gazeau, Bonnie Li, Fabio Pardo, Luyu Wang, Lei Zhang, Frederic Besse, Tim Harley, Anna Mitenkova, Jane Wang, Jeff Clune, Demis Hassabis, Raia Hadsell, Adrian Bolton, Satinder Singh, and Tim Rocktäschel. Genie 2: a large-scale foundation world model. Google DeepMind blog, 4 dicembre 2024, 2024. URL: https://deepmind.google/discover/blog/genie-2-a-large-scale-foundation-world-model/.
Jack Parker-Holder and Shlomi Fruchter. Genie 3: a new frontier for world models. Google DeepMind blog, 5 agosto 2025, 2025. URL: https://deepmind.google/discover/blog/genie-3-a-new-frontier-for-world-models/.
Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Łukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran. Image transformer. In Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Machine Learning Research, 4055–4064. PMLR, 2018.
Thomas Parr, Giovanni Pezzulo, and Karl J. Friston. Active Inference: The Free Energy Principle in Mind, Brain, and Behavior. The MIT Press, 2022. URL: https://direct.mit.edu/books/oa-monograph/5299/, doi:10.7551/mitpress/12441.001.0001.
Ankita Pasad, Ju-Chieh Chou, and Karen Livescu. Layer-wise analysis of a self-supervised speech representation model. In IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 914–921. 2021. doi:10.1109/ASRU51503.2021.9688093.
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio. On the difficulty of training recurrent neural networks. In Proceedings of the 30th International Conference on Machine Learning (ICML), 1310–1318. 2013.
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. Pytorch: an imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems, volume 32. 2019.
Pitch Patarasuk and Xin Yuan. Bandwidth optimal all-reduce algorithms for clusters of workstations. Journal of Parallel and Distributed Computing, 69(2):117–124, 2009. doi:10.1016/j.jpdc.2008.09.002.
Deepak Pathak, Pulkit Agrawal, Alexei A. Efros, and Trevor Darrell. Curiosity-driven exploration by self-supervised prediction. In International Conference on Machine Learning (ICML). 2017. URL: https://arxiv.org/abs/1705.05363.
Jaideep Pathak, Shashank Subramanian, Peter Harrington, Sanjeev Raja, Ashesh Chattopadhyay, Morteza Mardani, Thorsten Kurth, David Hall, Zongyi Li, Kamyar Azizzadenesheli, Pedram Hassanzadeh, Karthik Kashinath, and Animashree Anandkumar. FourCastNet: a global data-driven high-resolution weather model using adaptive Fourier neural operators. arXiv preprint arXiv:2202.11214, 2022.
David Patterson, Joseph Gonzalez, Urs Hölzle, Quoc Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David R. So, Maud Texier, and Jeff Dean. The carbon footprint of machine learning training will plateau, then shrink. Computer, 55(7):18–28, 2022. doi:10.1109/MC.2022.3148714.
David Patterson, Joseph Gonzalez, Quoc Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David So, Maud Texier, and Jeff Dean. Carbon emissions and large neural network training. arXiv preprint arXiv:2104.10350, 2021.
Martin Pawelczyk, Klaus Broelemann, and Gjergji Kasneci. On counterfactual explanations under predictive multiplicity. In Proceedings of the 36th Conference on Uncertainty in Artificial Intelligence (UAI), volume 124 of Proceedings of Machine Learning Research, 809–818. 2020.
Judea Pearl. The solution for the branching factor of the alpha-beta pruning algorithm and its optimality. Communications of the ACM, 25(8):559–564, 1982. doi:10.1145/358589.358616.
Judea Pearl and Dana Mackenzie. The Book of Why: The New Science of Cause and Effect. Basic Books, 2018.
Karl Pearson. On lines and planes of closest fit to systems of points in space. The London, Edinburgh, and Dublin Philosophical Magazine and Journal of Science, 2(11):559–572, 1901.
William Peebles and Saining Xie. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2023. URL: https://arxiv.org/abs/2212.09748.
Bo Peng, Eric Alcaide, Quentin Anthony, and others. RWKV: reinventing RNNs for the transformer era. In Findings of the Association for Computational Linguistics: EMNLP 2023. 2023.
Bo Peng, Daniel Goldstein, Quentin Anthony, and others. Eagle and Finch: RWKV with matrix-valued states and dynamic recurrence. In Conference on Language Modeling (COLM). 2024.
Bo Peng and others. RWKV-7 “Goose” with expressive dynamic state evolution. arXiv preprint arXiv:2503.14456, 2025.
Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, and Furu Wei. Kosmos-2: grounding multimodal large language models to the world. arXiv preprint arXiv:2306.14824, 2023.
Jeffrey Pennington, Richard Socher, and Christopher D. Manning. Glove: global vectors for word representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2014.
Lionel S. Penrose and Roger Penrose. Impossible objects: a special type of visual illusion. British Journal of Psychology, 49(1):31–33, 1958.
Roger Penrose. A generalized inverse for matrices. Mathematical Proceedings of the Cambridge Philosophical Society, 51(3):406–413, 1955. doi:10.1017/S0305004100030401.
Roger Penrose. On the cohomology of impossible figures. Structural Topology, 1991.
Juan C. Perdomo, Tijana Zrnic, Celestine Mendler-Dünner, and Moritz Hardt. Performative prediction. In Proceedings of the 37th International Conference on Machine Learning (ICML), volume 119 of Proceedings of Machine Learning Research, 7599–7609. 2020.
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving. Red teaming language models with language models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 3419–3448. Association for Computational Linguistics, 2022. doi:10.18653/v1/2022.emnlp-main.225.
Bryan Perozzi, Rami Al-Rfou, and Steven Skiena. Deepwalk: online learning of social representations. In Proceedings of the 20th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD), 701–710. 2014. URL: https://arxiv.org/abs/1403.6652.
L. Personnaz, I. Guyon, and G. Dreyfus. Information storage and retrieval in spin-glass like neural networks. Journal de Physique Lettres, 46(8):L359–L365, 1985. doi:10.1051/jphyslet:01985004608035900.
Karl Pertsch, Kyle Stachowicz, Brian Ichter, Danny Driess, Suraj Nair, Quan Vuong, Oier Mees, Chelsea Finn, and Sergey Levine. FAST: efficient action tokenization for vision-language-action models. In Proceedings of Robotics: Science and Systems (RSS). 2025. doi:10.15607/RSS.2025.XXI.012.
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. Deep contextualized word representations. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), 2227–2237. 2018. doi:10.18653/v1/N18-1202.
Aleksandar Petrov, Emanuele La Malfa, Philip H. S. Torr, and Adel Bibi. Language model tokenizers introduce unfairness between languages. In Advances in Neural Information Processing Systems (NeurIPS), volume 36, 36963–36990. 2023.
Aleksandr Petrov and Craig Macdonald. A systematic review and replicability study of BERT4Rec for sequential recommendation. In Proceedings of the 16th ACM Conference on Recommender Systems (RecSys '22), 436–447. ACM, 2022. doi:10.1145/3523227.3548487.
Jonas Pfeiffer, Naman Goyal, Xi Victoria Lin, Xian Li, James Cross, Sebastian Riedel, and Mikel Artetxe. Lifting the curse of multilinguality by pre-training modular transformers. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL), 3479–3495. 2022.
Karol J. Piczak. ESC: dataset for environmental sound classification. In Proceedings of the 23rd ACM International Conference on Multimedia, 1015–1018. ACM, 2015. doi:10.1145/2733373.2806390.
Joelle Pineau, Philippe Vincent-Lamarre, Koustuv Sinha, Vincent Larivière, Alina Beygelzimer, Florence d'Alché-Buc, Emily Fox, and Hugo Larochelle. Improving reproducibility in machine learning research (a report from the neurips 2019 reproducibility program). Journal of Machine Learning Research, 22(164):1–20, 2021. URL: https://jmlr.org/papers/v22/20-303.html.
Telmo Pires, Eva Schlinger, and Dan Garrette. How multilingual is multilingual BERT? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 4996–5001. 2019.
John C. Platt. Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods. In Advances in Large Margin Classifiers, 61–74. MIT Press, 1999.
Geoff Pleiss, Manish Raghavan, Felix Wu, Jon Kleinberg, and Kilian Q. Weinberger. On fairness and calibration. In Advances in Neural Information Processing Systems 30 (NIPS). 2017. URL: https://arxiv.org/abs/1709.02012.
R.-E. Plessix. A review of the adjoint-state method for computing the gradient of a functional with geophysical applications. Geophysical Journal International, 167(2):495–503, 2006. doi:10.1111/j.1365-246X.2006.02978.x.
Ira Pohl. Heuristic search viewed as path finding in a graph. Artificial Intelligence, 1(3–4):193–204, 1970.
Michael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y. Fu, Tri Dao, Stephen A. Baccus, and others. Hyena hierarchy: towards larger convolutional language models. In Proceedings of the 40th International Conference on Machine Learning (ICML). 2023.
Boris T. Polyak. Some methods of speeding up the convergence of iteration methods. USSR Computational Mathematics and Mathematical Physics, 4(5):1–17, 1964. doi:10.1016/0041-5553(64)90137-5.
Dean A. Pomerleau. ALVINN: an autonomous land vehicle in a neural network. In Advances in Neural Information Processing Systems (NIPS), volume 1, 305–313. 1989. URL: https://proceedings.neurips.cc/paper/1988/hash/812b4ba287f5ee0bc9d43bbf5bbe87fb-Abstract.html.
Dean A. Pomerleau. Efficient training of artificial neural networks for autonomous navigation. Neural Computation, 3(1):88–97, 1991. doi:10.1162/neco.1991.3.1.88.
Natalia Ponomareva, Hussein Hazimeh, Alex Kurakin, Zheng Xu, Carson Denison, H. Brendan McMahan, Sergei Vassilvitskii, Steve Chien, and Abhradeep Guha Thakurta. How to DP-fy ML: a practical guide to machine learning with differential privacy. Journal of Artificial Intelligence Research, 77:1113–1201, 2023. URL: https://arxiv.org/abs/2303.00654.
Ben Poole, Sherjil Ozair, Aaron van den Oord, Alexander A. Alemi, and George Tucker. On variational bounds of mutual information. In Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, 5171–5180. PMLR, 2019. URL: https://proceedings.mlr.press/v97/poole19a.html.
Tomer Porian, Mitchell Wortsman, Jenia Jitsev, Ludwig Schmidt, and Yair Carmon. Resolving discrepancies in compute-optimal scaling of language models. In Advances in Neural Information Processing Systems, volume 37, 100535–100570. Curran Associates, Inc., 2024. doi:10.52202/079017-3189.
Martin F. Porter. An algorithm for suffix stripping. Program, 14(3):130–137, 1980. doi:10.1108/eb046814.
Martin F. Porter. Snowball: a language for stemming algorithms. 2001. URL: https://snowballstem.org/texts/introduction.html.
Matt Post. A call for clarity in reporting BLEU scores. In Proceedings of the Third Conference on Machine Translation: Research Papers, 186–191. Brussels, Belgium, 2018. Association for Computational Linguistics. doi:10.18653/v1/W18-6319.
Ralph K. Potter, George A. Kopp, and Harriet C. Green. Visible Speech. D. Van Nostrand, New York, 1947.
Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra. Grokking: generalization beyond overfitting on small algorithmic datasets. In ICLR 2022 Workshop on Mathematical Reasoning and AI (MATH-AI). 2022. URL: https://arxiv.org/abs/2201.02177, arXiv:2201.02177.
Doina Precup, Richard S. Sutton, and Sanjoy Dasgupta. Off-policy temporal-difference learning with function approximation. In Proceedings of the Eighteenth International Conference on Machine Learning, 417–424. Morgan Kaufmann, 2001.
Ofir Press, Noah A. Smith, and Mike Lewis. Train short, test long: attention with linear biases enables input length extrapolation. In International Conference on Learning Representations (ICLR). 2022.
Simon J. D. Prince. Understanding Deep Learning. MIT Press, 2023. ISBN 9780262048644. la versione consultata è la bozza dell'8 febbraio 2026, su https://udlbook.com.
Hilary Putnam. Minds and machines. In Sidney Hook, editor, Dimensions of Mind, pages 148–179. New York University Press, New York, 1960.
Jorge Pérez, Pablo Barceló, and Javier Marinkovic. Attention is Turing-complete. Journal of Machine Learning Research, 22(75):1–35, 2021. URL: https://www.jmlr.org/papers/v22/20-302.html.
Ruoyu Qin, Zheming Li, Weiran He, Jialei Cui, Feng Ren, Mingxing Zhang, Yongwei Wu, Weimin Zheng, and Xinran Xu. Mooncake: trading more storage for less computation. a KVCache-centric architecture for serving LLM chatbot. In 23rd USENIX Conference on File and Storage Technologies (FAST). 2025.
Jiezhong Qiu, Yuxiao Dong, Hao Ma, Jian Li, Kuansan Wang, and Jie Tang. Network embedding as matrix factorization: unifying DeepWalk, LINE, PTE, and node2vec. In Proceedings of the Eleventh ACM International Conference on Web Search and Data Mining (WSDM), 459–467. 2018. doi:10.1145/3159652.3159706.
Joaquin Quiñonero-Candela, Masashi Sugiyama, Anton Schwaighofer, and Neil D. Lawrence, editors. Dataset Shift in Machine Learning. MIT Press, 2009.
Stephan Rabanser, Stephan Günnemann, and Zachary C. Lipton. Failing loudly: an empirical study of methods for detecting dataset shift. In Advances in Neural Information Processing Systems 32 (NeurIPS 2019), 1394–1406. 2019.
Lawrence R. Rabiner. A tutorial on hidden Markov models and selected applications in speech recognition. Proceedings of the IEEE, 77(2):257–286, 1989. doi:10.1109/5.18626.
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning (ICML). 2021.
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. Robust speech recognition via large-scale weak supervision. In Proceedings of the 40th International Conference on Machine Learning (ICML), volume 202 of Proceedings of Machine Learning Research, 28492–28518. 2023. arXiv:2212.04356.
Alec Radford, Luke Metz, and Soumith Chintala. Unsupervised representation learning with deep convolutional generative adversarial networks. In International Conference on Learning Representations (ICLR). 2016.
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving language understanding by generative pre-training. Technical Report, OpenAI, 2018.
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners. OpenAI Technical Report, 2019. URL: https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf.
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning, Stefano Ermon, and Chelsea Finn. Direct preference optimization: your language model is secretly a reward model. In Advances in Neural Information Processing Systems, volume 36. 2023. URL: https://arxiv.org/abs/2305.18290.
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67, 2020. URL: https://arxiv.org/abs/1910.10683, arXiv:1910.10683.
Maithra Raghu, Chiyuan Zhang, Jon Kleinberg, and Samy Bengio. Transfusion: understanding transfer learning for medical imaging. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019. URL: https://proceedings.neurips.cc/paper_files/paper/2019/file/eb1e78328c46506b46a4ac4a1e378b91-Paper.pdf.
Nasim Rahaman, Aristide Baratin, Devansh Arpit, Felix Draxler, Min Lin, Fred Hamprecht, Yoshua Bengio, and Aaron Courville. On the spectral bias of neural networks. In Proceedings of the 36th International Conference on Machine Learning (ICML), 5301–5310. 2019.
Afshin Rahimi, Trevor Cohn, and Timothy Baldwin. Semi-supervised user geolocation via graph convolutional networks. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics. 2018.
Maziar Raissi, Paris Perdikaris, and George Em Karniadakis. Physics-informed neural networks: a deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations. Journal of Computational Physics, 378:686–707, 2019. doi:10.1016/j.jcp.2018.10.045.
Maziar Raissi, Alireza Yazdani, and George Em Karniadakis. Hidden fluid mechanics: learning velocity and pressure fields from flow visualizations. Science, 367(6481):1026–1030, 2020. doi:10.1126/science.aaw4741.
Senthooran Rajamanoharan, Tom Lieberum, Nicolas Sonnerat, Arthur Conmy, Vikrant Varma, János Kramár, and Neel Nanda. Jumping ahead: improving reconstruction fidelity with JumpReLU sparse autoencoders. 2024. arXiv:2407.14435.
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. ZeRO: memory optimizations toward training trillion parameter models. In SC20: International Conference for High Performance Computing, Networking, Storage and Analysis. 2020.
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. SQuAD: 100,000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2383–2392. 2016. doi:10.18653/v1/D16-1264.
Prajit Ramachandran, Tom Le Paine, Pooya Khorrami, Mohammad Babaeizadeh, Shiyu Chang, Yang Zhang, Mark A. Hasegawa-Johnson, Roy H. Campbell, and Thomas S. Huang. Fast generation for convolutional autoregressive models. In International Conference on Learning Representations (ICLR), Workshop Track. 2017.
Aaditya Ramdas, Nicolás García Trillos, and Marco Cuturi. On Wasserstein two-sample testing and related families of nonparametric tests. Entropy, 19(2):47, 2017. doi:10.3390/e19020047.
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In Proceedings of the 38th International Conference on Machine Learning (ICML), 8821–8831. 2021.
Ladislav Rampášek, Michael Galkin, Vijay Prakash Dwivedi, Anh Tuan Luu, Guy Wolf, and Dominique Beaini. Recipe for a general, powerful, scalable graph transformer. In Advances in Neural Information Processing Systems (NeurIPS). 2022. URL: https://arxiv.org/abs/2205.12454.
Hubert Ramsauer, Bernhard Schäfl, Johannes Lehner, Philipp Seidl, Michael Widrich, Thomas Adler, Lukas Gruber, Markus Holzleitner, Milena Pavlović, Geir Kjetil Sandve, Victor Greiff, David Kreil, Michael Kopp, Günter Klambauer, Johannes Brandstetter, and Sepp Hochreiter. Hopfield networks is all you need. In International Conference on Learning Representations (ICLR). 2021. URL: https://arxiv.org/abs/2008.02217.
William M. Rand. Objective criteria for the evaluation of clustering methods. Journal of the American Statistical Association, 66(336):846–850, 1971. doi:10.1080/01621459.1971.10482356.
René Ranftl, Katrin Lasinger, David Hafner, Konrad Schindler, and Vladlen Koltun. Towards robust monocular depth estimation: mixing datasets for zero-shot cross-dataset transfer. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(3):1623–1637, 2022.
Marc'Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba. Sequence level training with recurrent neural networks. In International Conference on Learning Representations (ICLR). 2016.
David Raposo, Sam Ritter, Blake Richards, Timothy Lillicrap, Peter Conway Humphreys, and Adam Santoro. Mixture-of-depths: dynamically allocating compute in transformer-based language models. arXiv preprint arXiv:2404.02258, 2024.
Tabish Rashid, Mikayel Samvelyan, Christian Schroeder de Witt, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson. Qmix: monotonic value function factorisation for deep multi-agent reinforcement learning. In Proceedings of the 35th International Conference on Machine Learning (ICML). 2018.
Carl Edward Rasmussen and Christopher K. I. Williams. Gaussian Processes for Machine Learning. MIT Press, 2006.
Ali Razavi, Aäron van den Oord, and Oriol Vinyals. Generating diverse high-fidelity images with vq-vae-2. In Advances in Neural Information Processing Systems (NeurIPS), volume 32. 2019.
Esteban Real, Alok Aggarwal, Yanping Huang, and Quoc V. Le. Regularized evolution for image classifier architecture search. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, 4780–4789. 2019.
Sashank J. Reddi, Satyen Kale, and Sanjiv Kumar. On the convergence of Adam and beyond. In International Conference on Learning Representations (ICLR). 2018.
Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi. You only look once: unified, real-time object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2016.
Joseph Redmon and Ali Farhadi. YOLO9000: better, faster, stronger. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 7263–7271. 2017.
Joseph Redmon and Ali Farhadi. YOLOv3: an incremental improvement. arXiv preprint arXiv:1804.02767, 2018.
Nils Reimers and Iryna Gurevych. Sentence-BERT: sentence embeddings using siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 3982–3992. 2019. URL: https://arxiv.org/abs/1908.10084.
Christian H. Reinsch. Smoothing by spline functions. Numerische Mathematik, 10(3):177–183, 1967. doi:10.1007/BF02162161.
Liliang Ren, Yang Liu, Yadong Lu, Yelong Shen, Chen Liang, and Weizhu Chen. Samba: simple hybrid state space models for efficient unlimited context language modeling. In International Conference on Learning Representations (ICLR). 2025. arXiv:2406.07522.
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster R-CNN: towards real-time object detection with region proposal networks. In Advances in Neural Information Processing Systems, volume 28. 2015.
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu. Fastspeech 2: fast and high-quality end-to-end text to speech. In Proceedings of the 9th International Conference on Learning Representations (ICLR). 2021. URL: https://arxiv.org/abs/2006.04558.
Steffen Rendle, Christoph Freudenthaler, Zeno Gantner, and Lars Schmidt-Thieme. BPR: bayesian personalized ranking from implicit feedback. In Proceedings of the 25th Conference on Uncertainty in Artificial Intelligence, 452–461. 2009.
Steffen Rendle, Walid Krichene, Li Zhang, and John Anderson. Neural collaborative filtering vs. matrix factorization revisited. In Proceedings of the 14th ACM Conference on Recommender Systems, 240–248. 2020.
Steffen Rendle, Walid Krichene, Li Zhang, and Yehuda Koren. Revisiting the performance of iALS on item recommendation benchmarks. In Proceedings of the 16th ACM Conference on Recommender Systems (RecSys '22), 427–435. 2022. arXiv:2110.14037. doi:10.1145/3523227.3548486.
Paul Resnick, Neophytos Iacovou, Mitesh Suchak, Peter Bergstrom, and John Riedl. GroupLens: an open architecture for collaborative filtering of netnews. In Proceedings of the 1994 ACM Conference on Computer Supported Cooperative Work (CSCW '94), 175–186. 1994. doi:10.1145/192844.192905.
Craig W. Reynolds. Flocks, herds and schools: a distributed behavioral model. In Proceedings of the 14th Annual Conference on Computer Graphics and Interactive Techniques (SIGGRAPH), 25–34. 1987.
Danilo Jimenez Rezende and Shakir Mohamed. Variational inference with normalizing flows. In International Conference on Machine Learning (ICML). 2015.
Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra. Stochastic backpropagation and approximate inference in deep generative models. In Proceedings of the 31st International Conference on Machine Learning (ICML). 2014. URL: https://arxiv.org/abs/1401.4082.
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. "why should i trust you?": explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD). 2016. URL: https://arxiv.org/abs/1602.04938.
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. Anchors: high-precision model-agnostic explanations. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 32. 2018.
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. Beyond accuracy: behavioral testing of NLP models with CheckList. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), 4902–4912. 2020. doi:10.18653/v1/2020.acl-main.442.
Henry Gordon Rice. Classes of recursively enumerable sets and their decision problems. Transactions of the American Mathematical Society, 74(2):358–366, 1953.
Pierre H. Richemond, Jean-Bastien Grill, Florent Altché, Corentin Tallec, Florian Strub, Andrew Brock, Samuel Smith, Soham De, Razvan Pascanu, Bilal Piot, and Michal Valko. BYOL works even without batch statistics. arXiv preprint arXiv:2010.10241, 2020. URL: https://arxiv.org/abs/2010.10241.
Martin Riedmiller. Neural fitted Q iteration – first experiences with a data efficient neural reinforcement learning method. In João Gama, Rui Camacho, Pavel B. Brazdil, Alípio Mário Jorge, and Luís Torgo, editors, Machine Learning: ECML 2005, Lecture Notes in Computer Science, 317–328. Springer, 2005. doi:10.1007/11564096_32.
Salah Rifai, Pascal Vincent, Xavier Muller, Xavier Glorot, and Yoshua Bengio. Contractive auto-encoders: explicit invariance during feature extraction. In Proceedings of the 28th International Conference on Machine Learning (ICML), 833–840. 2011.
Jorma Rissanen. Modeling by shortest data description. Automatica, 14(5):465–471, 1978. doi:10.1016/0005-1098(78)90005-5.
A. W. Rix, J. G. Beerends, M. P. Hollier, and A. P. Hekstra. Perceptual evaluation of speech quality (PESQ): a new method for speech quality assessment of telephone networks and codecs. In IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), volume 2, 749–752. 2001. doi:10.1109/ICASSP.2001.941023.
Gareth O. Roberts and Richard L. Tweedie. Exponential convergence of Langevin distributions and their discrete approximations. Bernoulli, 2(4):341–363, 1996. doi:10.2307/3318418.
Stephen Robertson and Hugo Zaragoza. The probabilistic relevance framework: BM25 and beyond. Foundations and Trends in Information Retrieval, 3(4):333–389, 2009. doi:10.1561/1500000019.
Julia Robinson. An iterative method of solving a game. Annals of Mathematics, 54(2):296–301, 1951. doi:10.2307/1969530.
Geoffrey Roeder, Yuhuai Wu, and David Duvenaud. Sticking the landing: simple, lower-variance gradient estimators for variational inference. In Advances in Neural Information Processing Systems (NeurIPS), volume 30. 2017. URL: https://arxiv.org/abs/1703.09194.
Anna Rohrbach, Lisa Anne Hendricks, Kaylee Burns, Trevor Darrell, and Kate Saenko. Object hallucination in image captioning. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP), 4035–4045. 2018.
Michal Rolinek, Dominik Zietlow, and Georg Martius. Variational autoencoders pursue PCA directions (by accident). In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2019. URL: https://arxiv.org/abs/1812.06775.
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2022.
Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio. Fitnets: hints for thin deep nets. In International Conference on Learning Representations (ICLR). 2015. arXiv:1412.6550.
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: convolutional networks for biomedical image segmentation. In Medical Image Computing and Computer-Assisted Intervention (MICCAI). 2015.
Mihaela Rosca, Balaji Lakshminarayanan, and Shakir Mohamed. Distribution matching in variational inference. arXiv preprint arXiv:1802.06847, 2018.
Frank Rosenblatt. The perceptron: a probabilistic model for information storage and organization in the brain. Psychological Review, 65(6):386–408, 1958.
Frank Rosenblatt. Principles of Neurodynamics: Perceptrons and the Theory of Brain Mechanisms. Spartan Books, Washington, DC, 1962.
Stéphane Ross and Drew Bagnell. Efficient reductions for imitation learning. In Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics (AISTATS), volume 9 of Proceedings of Machine Learning Research, 661–668. 2010. URL: https://proceedings.mlr.press/v9/ross10a.html.
Stéphane Ross, Geoffrey J. Gordon, and J. Andrew Bagnell. A reduction of imitation learning and structured prediction to no-regret online learning. In Proceedings of the 14th International Conference on Artificial Intelligence and Statistics (AISTATS), 627–635. 2011. URL: https://proceedings.mlr.press/v15/ross11a.html.
Peter J. Rousseeuw. Silhouettes: a graphical aid to the interpretation and validation of cluster analysis. Journal of Computational and Applied Mathematics, 20:53–65, 1987. doi:10.1016/0377-0427(87)90125-7.
Donald B. Rubin. Inference and missing data. Biometrika, 63(3):581–592, 1976. URL: https://doi.org/10.1093/biomet/63.3.581.
Cynthia Rudin. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature Machine Intelligence, 1(5):206–215, 2019. URL: https://arxiv.org/abs/1811.10154.
Nikita Rudin, David Hoeller, Philipp Reist, and Marco Hutter. Learning to walk in minutes using massively parallel deep reinforcement learning. In Proceedings of the 5th Conference on Robot Learning (CoRL), volume 164 of Proceedings of Machine Learning Research, 91–100. PMLR, 2022. URL: https://arxiv.org/abs/2109.11978.
David E. Rumelhart, Geoffrey E. Hinton, and Ronald J. Williams. Learning representations by back-propagating errors. Nature, 323:533–536, 1986.
Gavin A. Rummery and Mahesan Niranjan. On-line q-learning using connectionist systems. Technical Report CUED/F-INFENG/TR 166, Cambridge University Engineering Department, 1994.
Carl Runge. Über empirische funktionen und die interpolation zwischen äquidistanten ordinaten. Zeitschrift für Mathematik und Physik, 46:224–243, 1901.
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei. Imagenet large scale visual recognition challenge. International Journal of Computer Vision, 115(3):211–252, 2015. doi:10.1007/s11263-015-0816-y.
Bertrand Russell. The Problems of Philosophy. Williams and Norgate, London, 1912. URL: https://www.gutenberg.org/ebooks/5827.
Stuart Russell and Peter Norvig. Artificial Intelligence: A Modern Approach. Pearson, fourth edition, 2020. ISBN 978-0134610993.
Phillip Rust, Jonas Pfeiffer, Ivan Vulić, Sebastian Ruder, and Iryna Gurevych. How good is your tokenizer? on the monolingual performance of multilingual language models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics (ACL), 3118–3135. 2021.
Harvey Sacks, Emanuel A. Schegloff, and Gail Jefferson. A simplest systematics for the organization of turn-taking for conversation. Language, 50(4):696–735, 1974. doi:10.2307/412243.
Marco Saerens, Patrice Latinne, and Christine Decaestecker. Adjusting the outputs of a classifier to new a priori probabilities: a simple procedure. Neural Computation, 14(1):21–41, 2002.
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J. Fleet, and Mohammad Norouzi. Photorealistic text-to-image diffusion models with deep language understanding. In Advances in Neural Information Processing Systems (NeurIPS). 2022. URL: https://arxiv.org/abs/2205.11487.
Subham Sekhar Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan, Edgar Marroquin, Justin T. Chiu, Alexander M. Rush, and Volodymyr Kuleshov. Simple and effective masked diffusion language models. In Advances in Neural Information Processing Systems (NeurIPS). 2024. URL: https://arxiv.org/abs/2406.07524.
Said E. Said and David A. Dickey. Testing for unit roots in autoregressive-moving average models of unknown order. Biometrika, 71(3):599–607, 1984. doi:10.1093/biomet/71.3.599.
Takaya Saito and Marc Rehmsmeier. The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLOS ONE, 10(3):e0118432, 2015. doi:10.1371/journal.pone.0118432.
Mehdi S. M. Sajjadi, Olivier Bachem, Mario Lucic, Olivier Bousquet, and Sylvain Gelly. Assessing generative models via precision and recall. In Advances in Neural Information Processing Systems (NeurIPS). 2018.
Ruslan Salakhutdinov, Andriy Mnih, and Geoffrey Hinton. Restricted Boltzmann machines for collaborative filtering. In Proceedings of the 24th International Conference on Machine Learning (ICML '07), 791–798. ACM, 2007. doi:10.1145/1273496.1273596.
Ruslan Salakhutdinov and Iain Murray. On the quantitative analysis of deep belief networks. In Proceedings of the 25th International Conference on Machine Learning (ICML 2008), 872–879. 2008. doi:10.1145/1390156.1390266.
Justin Salamon, Christopher Jacoby, and Juan Pablo Bello. A dataset and taxonomy for urban sound research. In Proceedings of the 22nd ACM International Conference on Multimedia, 1041–1044. ACM, 2014. doi:10.1145/2647868.2655045.
Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen. Improved techniques for training gans. In Advances in Neural Information Processing Systems, volume 29. 2016.
Tim Salimans and Jonathan Ho. Progressive distillation for fast sampling of diffusion models. In International Conference on Learning Representations (ICLR). 2022. URL: https://arxiv.org/abs/2202.00512.
Tim Salimans, Andrej Karpathy, Xi Chen, and Diederik P. Kingma. PixelCNN++: improving the PixelCNN with discretized logistic mixture likelihood and other modifications. In International Conference on Learning Representations (ICLR). 2017.
Tim Salimans and David A. Knowles. Fixed-form variational posterior approximation through stochastic linear regression. Bayesian Analysis, 8(4):837–882, 2013.
David Salinas, Valentin Flunkert, Jan Gasthaus, and Tim Januschowski. Deepar: probabilistic forecasting with autoregressive recurrent networks. International Journal of Forecasting, 36(3):1181–1191, 2020. URL: https://arxiv.org/abs/1704.04110.
Gerard Salton and Christopher Buckley. Term-weighting approaches in automatic text retrieval. Information Processing & Management, 24(5):513–523, 1988. doi:10.1016/0306-4573(88)90021-0.
Arthur L. Samuel. Some studies in machine learning using the game of checkers. IBM Journal of Research and Development, 3(3):210–229, 1959.
Alvaro Sanchez-Gonzalez, Jonathan Godwin, Tobias Pfaff, Rex Ying, Jure Leskovec, and Peter W. Battaglia. Learning to simulate complex physics with graph networks. In Proceedings of the 37th International Conference on Machine Learning. 2020.
Mark Sanderson and W. Bruce Croft. The history of information retrieval research. Proceedings of the IEEE, 100(Special Centennial Issue):1444–1451, 2012. doi:10.1109/JPROC.2012.2189916.
Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. MobileNetV2: inverted residuals and linear bottlenecks. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 4510–4520. 2018. URL: https://arxiv.org/abs/1801.04381.
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108, 2019. 5th Workshop on Energy Efficient Machine Learning and Cognitive Computing, NeurIPS 2019.
Keshav Santhanam, Omar Khattab, Jon Saad-Falcon, Christopher Potts, and Matei Zaharia. ColBERTv2: effective and efficient retrieval via lightweight late interaction. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 3715–3734. Seattle, United States, 2022. Association for Computational Linguistics. doi:10.18653/v1/2022.naacl-main.272.
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. Whose opinions do language models reflect? In Proceedings of the 40th International Conference on Machine Learning (ICML), volume 202 of Proceedings of Machine Learning Research, 29971–30004. 2023. URL: https://arxiv.org/abs/2303.17548.
Shibani Santurkar, Dimitris Tsipras, Andrew Ilyas, and Aleksander Madry. How does batch normalization help optimization? In Advances in Neural Information Processing Systems 31 (NeurIPS). 2018.
Badrul Sarwar, George Karypis, Joseph Konstan, and John Riedl. Item-based collaborative filtering recommendation algorithms. In Proceedings of the 10th International Conference on World Wide Web (WWW '01), 285–295. 2001. doi:10.1145/371920.372071.
Danilo Sato, Arif Wider, and Christoph Windheuser. Continuous delivery for machine learning. https://martinfowler.com/articles/cd4ml.html, 2019. martinfowler.com.
Axel Sauer, Dominik Lorenz, Andreas Blattmann, and Robin Rombach. Adversarial diffusion distillation. In European Conference on Computer Vision (ECCV). 2024. URL: https://arxiv.org/abs/2311.17042.
Norbert Sauer. On the density of families of sets. Journal of Combinatorial Theory, Series A, 13(1):145–147, 1972.
Franco Scarselli, Marco Gori, Ah Chung Tsoi, Markus Hagenbuchner, and Gabriele Monfardini. The graph neural network model. IEEE Transactions on Neural Networks, 20(1):61–80, 2009.
Rylan Schaeffer, Brando Miranda, and Sanmi Koyejo. Are emergent abilities of large language models a mirage? In Advances in Neural Information Processing Systems 36 (NeurIPS). 2023. URL: https://arxiv.org/abs/2304.15004, arXiv:2304.15004.
Daniel Scharstein and Richard Szeliski. A taxonomy and evaluation of dense two-frame stereo correspondence algorithms. International Journal of Computer Vision, 47(1):7–42, 2002.
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver. Prioritized experience replay. In International Conference on Learning Representations (ICLR). 2016.
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. Toolformer: language models can teach themselves to use tools. In Advances in Neural Information Processing Systems (NeurIPS), volume 36. 2023.
Imanol Schlag, Kazuki Irie, and Jürgen Schmidhuber. Linear transformers are secretly fast weight programmers. In Proceedings of the 38th International Conference on Machine Learning (ICML). 2021.
Michael Schlichtkrull, Thomas N. Kipf, Peter Bloem, Rianne van den Berg, Ivan Titov, and Max Welling. Modeling relational data with graph convolutional networks. In The Semantic Web: 15th International Conference, ESWC, 593–607. 2018.
Erhard Schmidt. Zur theorie der linearen und nichtlinearen integralgleichungen. i. teil: entwicklung willkürlicher funktionen nach systemen vorgeschriebener. Mathematische Annalen, 63:433–476, 1907. doi:10.1007/BF01449770.
Tobias Schnabel, Adith Swaminathan, Ashudeep Singh, Navin Chandak, and Thorsten Joachims. Recommendations as treatments: debiasing learning and evaluation. In International Conference on Machine Learning (ICML). 2016. URL: https://arxiv.org/abs/1602.05352.
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli. Wav2vec: unsupervised pre-training for speech recognition. In Interspeech, 3465–3469. 2019. doi:10.21437/Interspeech.2019-1873.
Isaac Jacob Schoenberg. Contributions to the problem of approximation of equidistant data by analytic functions. Quarterly of Applied Mathematics, 4(1):45–99, 1946. doi:10.1090/qam/15914.
Isaac Jacob Schoenberg. Spline functions and the problem of graduation. Proceedings of the National Academy of Sciences, 52(4):947–950, 1964. doi:10.1073/pnas.52.4.947.
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, Timothy Lillicrap, and David Silver. Mastering atari, go, chess and shogi by planning with a learned model. Nature, 588(7839):604–609, 2020. URL: https://arxiv.org/abs/1911.08265.
Bianca Schroeder, Adam Wierman, and Mor Harchol-Balter. Open versus closed: a cautionary tale. In 3rd Symposium on Networked Systems Design and Implementation (NSDI). USENIX Association, 2006.
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz. Trust region policy optimization. In Proceedings of the 32nd International Conference on Machine Learning. 2015.
John Schulman, Philipp Moritz, Sergey Levine, Michael I. Jordan, and Pieter Abbeel. High-dimensional continuous control using generalized advantage estimation. In International Conference on Learning Representations (ICLR). 2016. URL: https://arxiv.org/abs/1506.02438, arXiv:1506.02438.
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017.
Mike Schuster and Kaisuke Nakajima. Japanese and korean voice search. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 5149–5152. 2012. doi:10.1109/ICASSP.2012.6289079.
Mike Schuster and Kuldip K. Paliwal. Bidirectional recurrent neural networks. IEEE Transactions on Signal Processing, 45(11):2673–2681, 1997.
Gideon Schwarz. Estimating the dimension of a model. The Annals of Statistics, 6(2):461–464, 1978.
Bernhard Schölkopf, Ralf Herbrich, and Alex J. Smola. A generalized representer theorem. In Computational Learning Theory (COLT 2001), volume 2111 of Lecture Notes in Computer Science, 416–426. Springer, 2001. doi:10.1007/3-540-44581-1_27.
Bernhard Schölkopf, John C. Platt, John Shawe-Taylor, Alex J. Smola, and Robert C. Williamson. Estimating the support of a high-dimensional distribution. Neural Computation, 13(7):1443–1471, 2001.
Bernhard Schölkopf, Alexander Smola, and Klaus-Robert Müller. Nonlinear component analysis as a kernel eigenvalue problem. Neural Computation, 10(5):1299–1319, 1998.
Bernhard Schölkopf and Alexander J. Smola. Learning with Kernels: Support Vector Machines, Regularization, Optimization, and Beyond. MIT Press, 2002.
Johannes L. Schönberger and Jan-Michael Frahm. Structure-from-motion revisited. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2016.
Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr. Quantifying language models' sensitivity to spurious features in prompt design or: how i learned to start worrying about prompt formatting. In The Twelfth International Conference on Learning Representations (ICLR). 2024. URL: https://openreview.net/forum?id=RIu5lyNXjT.
D. Sculley, Gary Holt, Daniel Golovin, Eugene Davydov, Todd Phillips, Dietmar Ebner, Vinay Chaudhary, Michael Young, Jean-François Crespo, and Dan Dennison. Hidden technical debt in machine learning systems. In Advances in Neural Information Processing Systems (NeurIPS), volume 28. 2015.
John R. Searle. Speech Acts: An Essay in the Philosophy of Language. Cambridge University Press, Cambridge, 1969. doi:10.1017/CBO9781139173438.
Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. Grad-cam: visual explanations from deep networks via gradient-based localization. In IEEE International Conference on Computer Vision (ICCV). 2017. URL: https://arxiv.org/abs/1610.02391.
Lesia Semenova, Cynthia Rudin, and Ronald Parr. On the existence of simpler machine learning models. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency (FAccT), 1827–1858. 2022. doi:10.1145/3531146.3533232.
Rico Sennrich, Barry Haddow, and Alexandra Birch. Neural machine translation of rare words with subword units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (ACL). 2016.
Joan Serrà, David Álvarez, Vicenç Gómez, Olga Slizovskaia, José F. Núñez, and Jordi Luque. Input complexity and out-of-distribution detection with likelihood-based generative models. In International Conference on Learning Representations (ICLR). 2020.
Jay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar, Pradeep Ramani, and Tri Dao. FlashAttention-3: fast and accurate attention with asynchrony and low-precision. In Advances in Neural Information Processing Systems (NeurIPS). 2024. URL: https://arxiv.org/abs/2407.08608, arXiv:2407.08608.
Cosma Rohilla Shalizi. “Attention”, “Transformers”, in neural network “Large Language Models”. Notebook online, 2023. URL: http://bactra.org/notebooks/nn-attention-and-transformers.html.
Ohad Shamir and Naftali Tishby. Model selection and stability in k-means clustering. In Proceedings of the 21st Annual Conference on Learning Theory (COLT). 2008.
Shreya Shankar, Rolando Garcia, Joseph M. Hellerstein, and Aditya G. Parameswaran. Operationalizing machine learning: an interview study. arXiv preprint arXiv:2209.09125, 2022.
Claude E. Shannon. A mathematical theory of communication. Bell System Technical Journal, 27:379–423, 623–656, 1948.
Claude E. Shannon. Programming a computer for playing chess. Philosophical Magazine, 41(314):256–275, 1950.
Claude E. Shannon. Prediction and entropy of printed english. Bell System Technical Journal, 30(1):50–64, 1951.
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y. K. Li, Y. Wu, and Daya Guo. DeepSeekMath: pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300, 2024.
Lloyd S. Shapley. Stochastic games. Proceedings of the National Academy of Sciences, 39(10):1095–1100, 1953. doi:10.1073/pnas.39.10.1095.
Lloyd S. Shapley. Some topics in two-person games. In Melvin Dresher, Lloyd S. Shapley, and Albert William Tucker, editors, Advances in Game Theory, volume 52 of Annals of Mathematics Studies, pages 1–28. Princeton University Press, Princeton, NJ, 1964. doi:10.1515/9781400882014-002.
Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R. Johnston, Shauna Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Da Yan, Miranda Zhang, and Ethan Perez. Towards understanding sycophancy in language models. arXiv preprint arXiv:2310.13548, 2023. URL: https://arxiv.org/abs/2310.13548.
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. Self-attention with relative position representations. In Proceedings of NAACL-HLT, 464–468. 2018.
Noam Shazeer. Fast transformer decoding: one write-head is all you need. arXiv preprint arXiv:1911.02150, 2019. URL: https://arxiv.org/abs/1911.02150.
Noam Shazeer. GLU variants improve transformer. arXiv preprint arXiv:2002.05202, 2020.
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc V. Le, Geoffrey E. Hinton, and Jeff Dean. Outrageously large neural networks: the sparsely-gated mixture-of-experts layer. In International Conference on Learning Representations (ICLR). 2017.
Saharon Shelah. A combinatorial problem; stability and order for models and theories in infinitary languages. Pacific Journal of Mathematics, 1972.
Jonathan Shen, Ruoming Pang, Ron J. Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, RJ Skerry-Ryan, Rif A. Saurous, Yannis Agiomyrgiannakis, and Yonghui Wu. Natural tts synthesis by conditioning wavenet on mel spectrogram predictions. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 4779–4783. 2018. URL: https://arxiv.org/abs/1712.05884.
Kangshen Shen, John N. Crossley, and Anthony W.-C. Lun. The Nine Chapters on the Mathematical Art: Companion and Commentary. Oxford University Press, 1999.
Sheng Shen, Le Hou, Yanqi Zhou, Nan Du, Shayne Longpre, Jason Wei, Hyung Won Chung, Barret Zoph, William Fedus, Xinyun Chen, Tu Vu, Yuexin Wu, Wuyang Chen, Albert Webson, Yunxuan Li, Vincent Zhao, Hongkun Yu, Kurt Keutzer, Trevor Darrell, and Denny Zhou. Mixture-of-experts meets instruction tuning: a winning combination for large language models. In International Conference on Learning Representations (ICLR), 18858–18884. 2024. URL: https://openreview.net/forum?id=6mLjDwYte5.
Walter A. Shewhart. Economic Control of Quality of Manufactured Product. D. Van Nostrand, New York, 1931.
Jiaxin Shi, Kehang Han, Zhe Wang, Arnaud Doucet, and Michalis K. Titsias. Simplified and generalized masked diffusion for discrete data. In Advances in Neural Information Processing Systems (NeurIPS). 2024. URL: https://arxiv.org/abs/2406.04329.
Yuhui Shi and Russell Eberhart. A modified particle swarm optimizer. In 1998 IEEE International Conference on Evolutionary Computation Proceedings, 69–73. IEEE, 1998. doi:10.1109/ICEC.1998.699146.
Hidetoshi Shimodaira. Improving predictive inference under covariate shift by weighting the log-likelihood function. Journal of Statistical Planning and Inference, 90(2):227–244, 2000.
Noah Shinn, Federico Cassano, Edward Berman, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. Reflexion: language agents with verbal reinforcement learning. In Advances in Neural Information Processing Systems (NeurIPS), volume 36. 2023.
Jonathan Shlomi, Peter Battaglia, and Jean-Roch Vlimant. Graph neural networks in particle physics. Machine Learning: Science and Technology, 2(2):021001, 2021. doi:10.1088/2632-2153/abbf9a.
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. Megatron-LM: training multi-billion parameter language models using model parallelism. arXiv preprint arXiv:1909.08053, 2019.
Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov. Membership inference attacks against machine learning models. In 2017 IEEE Symposium on Security and Privacy (SP), 3–18. 2017.
Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje. Learning important features through propagating activation differences. In Proceedings of the 34th International Conference on Machine Learning (ICML), volume 70 of PMLR, 3145–3153. 2017.
Kurt Shuster, Spencer Poff, Moya Chen, Douwe Kiela, and Jason Weston. Retrieval augmentation reduces hallucination in conversation. In Findings of the Association for Computational Linguistics: EMNLP 2021, 3784–3803. Punta Cana, Dominican Republic, 2021. Association for Computational Linguistics. URL: https://aclanthology.org/2021.findings-emnlp.320/, doi:10.18653/v1/2021.findings-emnlp.320.
Robin Sibson. SLINK: an optimally efficient algorithm for the single-link cluster method. The Computer Journal, 16(1):30–34, 1973.
David Siegmund. Sequential Analysis: Tests and Confidence Intervals. Springer Series in Statistics. Springer, New York, 1985.
Benjamin H. Sigelman, Luiz André Barroso, Mike Burrows, Pat Stephenson, Manoj Plakal, Donald Beaver, Saul Jaspan, and Chandan Shanbhag. Dapper, a large-scale distributed systems tracing infrastructure. Technical Report dapper-2010-1, Google, 2010.
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis. Mastering the game of go with deep neural networks and tree search. Nature, 529:484–489, 2016.
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy Lillicrap, Karen Simonyan, and Demis Hassabis. A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play. Science, 362(6419):1140–1144, 2018. doi:10.1126/science.aar6404.
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller. Deterministic policy gradient algorithms. In Proceedings of the 31st International Conference on Machine Learning, volume 32 of Proceedings of Machine Learning Research, 387–395. 2014. URL: https://proceedings.mlr.press/v32/silver14.html.
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy Lillicrap, Fan Hui, Laurent Sifre, George van den Driessche, Thore Graepel, and Demis Hassabis. Mastering the game of go without human knowledge. Nature, 550:354–359, 2017.
Patrice Simard, Bernard Victorri, Yann LeCun, and John Denker. Tangent Prop: a formalism for specifying selected invariances in an adaptive network. In John E. Moody, Stephen J. Hanson, and Richard P. Lippmann, editors, Advances in Neural Information Processing Systems, volume 4. Morgan Kaufmann, 1992. URL: https://proceedings.neurips.cc/paper_files/paper/1991/file/65658fde58ab3c2b6e5132a39fae7cb9-Paper.pdf.
Noah Simon, Jerome Friedman, Trevor Hastie, and Robert Tibshirani. A sparse-group lasso. Journal of Computational and Graphical Statistics, 22(2):231–245, 2013.
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. Deep inside convolutional networks: visualising image classification models and saliency maps. International Conference on Learning Representations (ICLR) Workshop, 2014. URL: https://arxiv.org/abs/1312.6034.
Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. In International Conference on Learning Representations (ICLR). 2015.
Aaditya K. Singh and DJ Strouse. Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs. arXiv preprint arXiv:2402.14903, 2024.
Satinder Singh, Tommi Jaakkola, Michael L. Littman, and Csaba Szepesvári. Convergence results for single-step on-policy reinforcement-learning algorithms. Machine Learning, 38(3):287–308, 2000.
Prasann Singhal, Tanya Goyal, Jiacheng Xu, and Greg Durrett. A long way to go: investigating length correlations in RLHF. In Conference on Language Modeling (COLM). 2024. URL: https://openreview.net/forum?id=G8LaO1P0xv.
Justin Sirignano and Konstantinos Spiliopoulos. DGM: a deep learning algorithm for solving partial differential equations. Journal of Computational Physics, 375:1339–1364, 2018.
Joar Skalse and Alessandro Abate. Misspecification in inverse reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence. 2023. arXiv:2212.03201.
Joar Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, and David Krueger. Defining and characterizing reward gaming. In Advances in Neural Information Processing Systems, volume 35, 9460–9471. 2022. Il PDF degli atti e la versione arXiv si intitolano «Defining and Characterizing Reward Hacking». URL: https://arxiv.org/abs/2209.13085, doi:10.52202/068431-0687.
Dylan Slack, Sophie Hilgard, Emily Jia, Sameer Singh, and Himabindu Lakkaraju. Fooling LIME and SHAP: adversarial attacks on post hoc explanation methods. In AAAI/ACM Conference on AI, Ethics, and Society (AIES). 2020. URL: https://arxiv.org/abs/1911.02508.
Arnold W. M. Smeulders, Marcel Worring, Simone Santini, Amarnath Gupta, and Ramesh Jain. Content-based image retrieval at the end of the early years. IEEE Transactions on Pattern Analysis and Machine Intelligence, 22(12):1349–1380, 2000. doi:10.1109/34.895972.
Daniel Smilkov, Nikhil Thorat, Been Kim, Fernanda Viégas, and Martin Wattenberg. SmoothGrad: removing noise by adding noise. arXiv preprint arXiv:1706.03825, 2017. URL: https://arxiv.org/abs/1706.03825.
Andries Petrus Smit, Nathan Grinsztajn, Paul Duckworth, Thomas D. Barrett, and Arnu Pretorius. Should we be going MAD? A look at multi-agent debate strategies for LLMs. In Proceedings of the 41st International Conference on Machine Learning (ICML), volume 235 of Proceedings of Machine Learning Research, 45883–45905. 2024. URL: https://arxiv.org/abs/2311.17371.
Jimmy T.H. Smith, Andrew Warrington, and Scott W. Linderman. Simplified state space layers for sequence modeling. In International Conference on Learning Representations (ICLR). 2023.
Leslie N. Smith. Cyclical learning rates for training neural networks. In 2017 IEEE Winter Conference on Applications of Computer Vision (WACV), 464–472. 2017. doi:10.1109/WACV.2017.58.
Reid G. Smith. The contract net protocol: high-level communication and control in a distributed problem solver. IEEE Transactions on Computers, C-29(12):1104–1113, 1980. doi:10.1109/TC.1980.1675516.
Paul Smolensky. Information processing in dynamical systems: foundations of harmony theory. In David E. Rumelhart and James L. McClelland, editors, Parallel Distributed Processing: Explorations in the Microstructure of Cognition, Volume 1: Foundations, pages 194–281. MIT Press, Cambridge, MA, 1986.
Slawek Smyl. A hybrid method of exponential smoothing and recurrent neural networks for time series forecasting. International Journal of Forecasting, 36(1):75–85, 2020. doi:10.1016/j.ijforecast.2019.03.017.
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. Scaling LLM test-time compute optimally can be more effective than scaling parameters for reasoning. In International Conference on Learning Representations (ICLR), 10131–10165. 2025. URL: https://openreview.net/forum?id=4FWAwZtd2n.
Jake Snell, Kevin Swersky, and Richard S. Zemel. Prototypical networks for few-shot learning. In Advances in Neural Information Processing Systems 30 (NeurIPS). 2017. URL: https://arxiv.org/abs/1703.05175.
Jasper Snoek, Hugo Larochelle, and Ryan P. Adams. Practical bayesian optimization of machine learning algorithms. In Advances in Neural Information Processing Systems, volume 25. 2012.
Jascha Sohl-Dickstein, Eric A. Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In Proceedings of the 32nd International Conference on Machine Learning (ICML), volume 37, 2256–2265. 2015. URL: https://proceedings.mlr.press/v37/sohl-dickstein15.html.
Ray J. Solomonoff. A formal theory of inductive inference. part i. Information and Control, 7(1):1–22, 1964. doi:10.1016/S0019-9958(64)90223-2.
Kyunghwan Son, Daewoo Kim, Wan Ju Kang, David Earl Hostallero, and Yung Yi. QTRAN: learning to factorize with transformation for cooperative multi-agent reinforcement learning. In Proceedings of the 36th International Conference on Machine Learning (ICML), volume 97 of Proceedings of Machine Learning Research, 5887–5896. 2019. URL: https://arxiv.org/abs/1905.05408.
Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. In International Conference on Learning Representations (ICLR). 2021. URL: https://arxiv.org/abs/2010.02502.
Yang Song and Prafulla Dhariwal. Improved techniques for training consistency models. In International Conference on Learning Representations (ICLR). 2024. URL: https://arxiv.org/abs/2310.14189.
Yang Song, Prafulla Dhariwal, Mark Chen, and Ilya Sutskever. Consistency models. In International Conference on Machine Learning (ICML). 2023. URL: https://arxiv.org/abs/2303.01469.
Yang Song and Stefano Ermon. Generative modeling by estimating gradients of the data distribution. In Advances in Neural Information Processing Systems (NeurIPS). 2019.
Yang Song, Sahaj Garg, Jiaxin Shi, and Stefano Ermon. Sliced score matching: a scalable approach to density and score estimation. In Proceedings of the 35th Uncertainty in Artificial Intelligence Conference, volume 115 of Proceedings of Machine Learning Research, 574–584. PMLR, 2020.
Yang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. In International Conference on Learning Representations (ICLR). 2021. URL: https://arxiv.org/abs/2011.13456.
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro. The implicit bias of gradient descent on separable data. Journal of Machine Learning Research, 19(70):1–57, 2018.
Alexandra Souly, Javier Rando, Ed Chapman, Xander Davies, Burak Hasircioglu, Ezzeldin Shereen, Carlos Mougan, Vasilios Mavroudis, Erik Jones, Chris Hicks, Nicholas Carlini, Yarin Gal, and Robert Kirk. Poisoning attacks on LLMs require a near-constant number of poison samples. arXiv preprint arXiv:2510.07192, 2025. URL: https://arxiv.org/abs/2510.07192.
C. Spearman. “general intelligence,” objectively determined and measured. The American Journal of Psychology, 15(2):201–292, 1904.
Elizabeth S. Spelke and Katherine D. Kinzler. Core knowledge. Developmental Science, 10(1):89–96, 2007. doi:10.1111/j.1467-7687.2007.00569.x.
Roland Sprague. Über mathematische kampfspiele. Tôhoku Mathematical Journal, 41:438–444, 1936. Vol. 41, annata 1935-36.
Zayne Sprague, Fangcong Yin, Juan Diego Rodriguez, Dongwei Jiang, Manya Wadhwa, Prasann Singhal, Xinyu Zhao, Xi Ye, Kyle Mahowald, and Greg Durrett. To CoT or not to CoT? chain-of-thought helps mainly on math and symbolic reasoning. In International Conference on Learning Representations (ICLR). 2025. URL: https://arxiv.org/abs/2409.12183.
Richard Sproat and Navdeep Jaitly. RNN approaches to text normalization: a challenge. arXiv preprint arXiv:1611.00068, 2016.
Karen Spärck Jones. A statistical interpretation of term specificity and its application in retrieval. Journal of Documentation, 28(1):11–21, 1972. doi:10.1108/eb026526.
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. Dropout: a simple way to prevent neural networks from overfitting. Journal of Machine Learning Research, 15:1929–1958, 2014.
Felix Stahlberg and Bill Byrne. On NMT search errors and model errors: cat got your tongue? In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 3356–3362. Hong Kong, China, 2019. Association for Computational Linguistics. doi:10.18653/v1/D19-1331.
Trevor Standley, Amir R. Zamir, Dawn Chen, Leonidas Guibas, Jitendra Malik, and Silvio Savarese. Which tasks should be learned together in multi-task learning? In Proceedings of the 37th International Conference on Machine Learning (ICML), volume 119 of Proceedings of Machine Learning Research, 9120–9132. 2020.
George Stein, Jesse C. Cresswell, Rasa Hosseinzadeh, Yi Sui, Brendan Leigh Ross, Valentin Villecroze, Zhaoyan Liu, Anthony L. Caterini, J. Eric T. Taylor, and Gabriel Loaiza-Ganem. Exposing flaws of generative model evaluation metrics and their unfair treatment of diffusion models. In Advances in Neural Information Processing Systems, volume 36, 3732–3784. 2023. URL: https://arxiv.org/abs/2306.04675.
Peter Steinberger. You shouldn't be prompting coding agents anymore. You should be designing loops that prompt your agents. Post su X, 7 giugno 2026, 2026. URL: https://x.com/steipete/status/2063697162748260627.
Ingo Steinwart. On the influence of the kernel on the consistency of support vector machines. Journal of Machine Learning Research, 2:67–93, 2001.
S. S. Stevens, J. Volkmann, and E. B. Newman. A scale for the measurement of the psychological magnitude pitch. The Journal of the Acoustical Society of America, 8(3):185–190, 1937. doi:10.1121/1.1915893.
Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M. Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano. Learning to summarize from human feedback. In Advances in Neural Information Processing Systems, volume 33, 3008–3021. 2020.
Tanya Stivers, N. J. Enfield, Penelope Brown, Christina Englert, Makoto Hayashi, Trine Heinemann, Gertie Hoymann, Federico Rossano, Jan Peter de Ruiter, Kyung-Eun Yoon, and Stephen C. Levinson. Universals and cultural variation in turn-taking in conversation. Proceedings of the National Academy of Sciences, 106(26):10587–10592, 2009. doi:10.1073/pnas.0903616106.
Jonathan M. Stokes, Kevin Yang, Kyle Swanson, Wengong Jin, Andres Cubillos-Ruiz, Nina M. Donghia, Craig R. MacNair, Shawn French, Lindsey A. Carfrae, Zohar Bloom-Ackermann, Victoria M. Tran, Anush Chiappino-Pepe, Ahmed H. Badran, Ian W. Andrews, Emma J. Chory, George M. Church, Eric D. Brown, Tommi S. Jaakkola, Regina Barzilay, and James J. Collins. A deep learning approach to antibiotic discovery. Cell, 180(4):688–702.e13, 2020. doi:10.1016/j.cell.2020.01.021.
Charles J. Stone. Consistent nonparametric regression. The Annals of Statistics, 5(4):595–620, 1977.
Charles J. Stone. Optimal global rates of convergence for nonparametric regression. The Annals of Statistics, 10(4):1040–1053, 1982. doi:10.1214/aos/1176345969.
Charles J. Stone. Additive regression and other nonparametric models. The Annals of Statistics, 13(2):689–705, 1985.
Philip J. Stone, Dexter C. Dunphy, Marshall S. Smith, and Daniel M. Ogilvie. The General Inquirer: A Computer Approach to Content Analysis. MIT Press, 1966.
Marilyn Strathern. `improving ratings': audit in the British University system. European Review, 5(3):305–321, 1997. doi:10.1002/(SICI)1234-981X(199707)5:3<305::AID-EURO184>3.0.CO;2-4.
Alexander L. Strehl and Michael L. Littman. An analysis of model-based interval estimation for Markov decision processes. Journal of Computer and System Sciences, 74(8):1309–1331, 2008. doi:10.1016/j.jcss.2007.08.009.
Carolin Strobl, Anne-Laure Boulesteix, Achim Zeileis, and Torsten Hothorn. Bias in random forest variable importance measures: illustrations, sources and a solution. BMC Bioinformatics, 2007. URL: https://doi.org/10.1186/1471-2105-8-25.
Emma Strubell, Ananya Ganesh, and Andrew McCallum. Energy and policy considerations for deep learning in nlp. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 3645–3650. 2019.
Pascal Sturmfels, Scott Lundberg, and Su-In Lee. Visualizing the impact of feature attribution baselines. Distill, 2020. URL: https://distill.pub/2020/attribution-baselines/.
Thomas Stützle and Marco Dorigo. A short convergence proof for a class of ant colony optimization algorithms. IEEE Transactions on Evolutionary Computation, 6(4):358–365, 2002. doi:10.1109/TEVC.2002.802444.
Thomas Stützle and Holger H. Hoos. MAX–MIN ant system. Future Generation Computer Systems, 16(8):889–914, 2000. doi:10.1016/S0167-739X(00)00043-1.
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. RoFormer: enhanced transformer with rotary position embedding. Neurocomputing, 568:127063, 2024.
Sainbayar Sukhbaatar, Arthur Szlam, Jason Weston, and Rob Fergus. End-to-end memory networks. In Advances in Neural Information Processing Systems, volume 28, 2440–2448. 2015.
Fei Sun, Jun Liu, Jian Wu, Changhua Pei, Xiao Lin, Wenwu Ou, and Peng Jiang. BERT4Rec: sequential recommendation with bidirectional encoder representations from transformer. In Proceedings of the 28th ACM International Conference on Information and Knowledge Management (CIKM), 1441–1450. 2019. doi:10.1145/3357384.3357895.
Mingjie Sun, Zhuang Liu, Anna Bair, and J. Zico Kolter. A simple and effective pruning approach for large language models. In International Conference on Learning Representations (ICLR). 2024.
Yutao Sun, Li Dong, Shaohan Huang, Shuming Ma, Yuqing Xia, Jilong Xue, Jianyong Wang, and Furu Wei. Retentive network: a successor to transformer for large language models. arXiv preprint arXiv:2307.08621, 2023.
Zekun Sun and Chaz Firestone. The dark room problem. Trends in Cognitive Sciences, 24(5):346–348, 2020. doi:10.1016/j.tics.2020.02.006.
Zhiqing Sun, Zhi-Hong Deng, Jian-Yun Nie, and Jian Tang. Rotate: knowledge graph embedding by relational rotation in complex space. In International Conference on Learning Representations. 2019.
Zhiqing Sun, Shikhar Vashishth, Soumya Sanyal, Partha Talukdar, and Yiming Yang. A re-evaluation of knowledge graph completion methods. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 5516–5522. 2020. doi:10.18653/v1/2020.acl-main.489.
Mukund Sundararajan and Amir Najmi. The many Shapley values for model explanation. In Proceedings of the 37th International Conference on Machine Learning (ICML), volume 119 of Proceedings of Machine Learning Research, 9269–9278. 2020. URL: https://proceedings.mlr.press/v119/sundararajan20b.html.
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. Axiomatic attribution for deep networks. In International Conference on Machine Learning (ICML). 2017. URL: https://arxiv.org/abs/1703.01365.
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vinicius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z. Leibo, Karl Tuyls, and Thore Graepel. Value-decomposition networks for cooperative multi-agent learning based on team reward. In Proceedings of the 17th International Conference on Autonomous Agents and MultiAgent Systems (AAMAS), 2085–2087. 2018.
Rao Surapaneni, Miku Jha, Michael Vakoc, and Todd Segal. Announcing the Agent2Agent protocol (A2A). Google Developers Blog, apr 2025. URL: https://developers.googleblog.com/en/a2a-a-new-era-of-agent-interoperability/.
Harini Suresh and John Guttag. A framework for understanding sources of harm throughout the machine learning life cycle. In Equity and Access in Algorithms, Mechanisms, and Optimization, EAAMO '21, 1–9. New York, NY, USA, 2021. ACM. doi:10.1145/3465416.3483305.
Ilya Sutskever. An observation on generalization. Intervento al workshop \emph Large Language Models and Transformers, Simons Institute for the Theory of Computing, Berkeley, 8 2023. Registrazione video con trascrizione; non esiste un articolo di accompagnamento.
Ilya Sutskever and Tijmen Tieleman. On the convergence properties of contrastive divergence. In Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics (AISTATS), volume 9 of Proceedings of Machine Learning Research, 789–795. 2010. URL: https://proceedings.mlr.press/v9/sutskever10a.html.
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. Sequence to sequence learning with neural networks. In Advances in Neural Information Processing Systems, volume 27. 2014.
Richard S. Sutton. Learning to predict by the methods of temporal differences. Machine Learning, 3(1):9–44, 1988.
Richard S. Sutton. Integrated architectures for learning, planning, and reacting based on approximating dynamic programming. In Proceedings of the Seventh International Conference on Machine Learning, 216–224. Morgan Kaufmann, 1990.
Richard S. Sutton. Dyna, an integrated architecture for learning, planning, and reacting. ACM SIGART Bulletin, 2(4):160–163, 1991. doi:10.1145/122344.122377.
Richard S. Sutton and Andrew G. Barto. Reinforcement Learning: An Introduction. MIT Press, second edition, 2018.
Richard S. Sutton, David McAllester, Satinder Singh, and Yishay Mansour. Policy gradient methods for reinforcement learning with function approximation. In Advances in Neural Information Processing Systems, volume 12. 2000.
Richard S. Sutton, Doina Precup, and Satinder Singh. Between mdps and semi-mdps: a framework for temporal abstraction in reinforcement learning. Artificial Intelligence, 112(1-2):181–211, 1999. doi:10.1016/S0004-3702(99)00052-1.
Mirac Suzgun and Adam Tauman Kalai. Meta-prompting: enhancing language models with task-agnostic scaffolding. arXiv preprint arXiv:2401.12954, 2024.
Latanya Sweeney. Simple demographics often identify people uniquely. Data Privacy Working Paper 3, Carnegie Mellon University, Pittsburgh, 2000. URL: https://dataprivacylab.org/projects/identifiability/paper1.pdf.
Latanya Sweeney. K-anonymity: a model for protecting privacy. International Journal of Uncertainty, Fuzziness and Knowledge-Based Systems, 10(5):557–570, 2002.
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. Going deeper with convolutions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2015.
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna. Rethinking the inception architecture for computer vision. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2818–2826. 2016.
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. Intriguing properties of neural networks. In International Conference on Learning Representations (ICLR). 2014.
Richard Szeliski. Computer Vision: Algorithms and Applications. Springer, second edition, 2022. URL: https://szeliski.org/Book/.
Cees H. Taal, Richard C. Hendriks, Richard Heusdens, and Jesper Jensen. An algorithm for intelligibility prediction of time–frequency weighted noisy speech. IEEE Transactions on Audio, Speech, and Language Processing, 19(7):2125–2136, 2011. doi:10.1109/TASL.2011.2114881.
Esteban G. Tabak and Cristina V. Turner. A family of nonparametric density estimation algorithms. Communications on Pure and Applied Mathematics, 66(2):145–164, 2013.
Esteban G. Tabak and Eric Vanden-Eijnden. Density estimation by dual ascent of the log-likelihood. Communications in Mathematical Sciences, 8(1):217–233, 2010.
Zhi Rui Tam, Cheng-Kuang Wu, Yi-Lin Tsai, Chieh-Yen Lin, Hung-yi Lee, and Yun-Nung Chen. Let me speak freely? A study on the impact of format restrictions on performance of large language models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track (EMNLP), 1218–1236. Association for Computational Linguistics, 2024. URL: https://aclanthology.org/2024.emnlp-industry.91/.
Ming Tan. Multi-agent reinforcement learning: independent vs. cooperative agents. In Machine Learning Proceedings 1993 (ICML), 330–337. 1993. doi:10.1016/B978-1-55860-307-3.50049-6.
Mingxing Tan and Quoc V. Le. Efficientnet: rethinking model scaling for convolutional neural networks. In Proceedings of the 36th International Conference on Machine Learning (ICML). 2019.
Matthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil, Nithin Raghavan, Utkarsh Singhal, Ravi Ramamoorthi, Jonathan T. Barron, and Ren Ng. Fourier features let networks learn high frequency functions in low dimensional domains. In Advances in Neural Information Processing Systems (NeurIPS). 2020. URL: https://arxiv.org/abs/2006.10739.
Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler. Long range arena: a benchmark for efficient transformers. In International Conference on Learning Representations (ICLR). 2021.
Sean J. Taylor and Benjamin Letham. Forecasting at scale. The American Statistician, 72(1):37–45, 2018. doi:10.1080/00031305.2017.1380080.
Wilson L. Taylor. «cloze procedure»: a new tool for measuring readability. Journalism Quarterly, 30(4):415–433, 1953. doi:10.1177/107769905303000401.
Zachary Teed and Jia Deng. RAFT: recurrent all-pairs field transforms for optical flow. In European Conference on Computer Vision (ECCV). 2020. URL: https://arxiv.org/abs/2003.12039.
Matus Telgarsky. Benefits of depth in neural networks. In 29th Annual Conference on Learning Theory (COLT). 2016.
Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, Hoagy Cunningham, Nicholas L. Turner, Callum McDougall, Monte MacDiarmid, C. Daniel Freeman, Theodore R. Sumers, Edward Rees, Joshua Batson, Adam Jermyn, Shan Carter, Chris Olah, and Tom Henighan. Scaling monosemanticity: extracting interpretable features from claude 3 sonnet. Transformer Circuits Thread, Anthropic, 2024. URL: https://transformer-circuits.pub/2024/scaling-monosemanticity/.
Joshua B. Tenenbaum, Vin de Silva, and John C. Langford. A global geometric framework for nonlinear dimensionality reduction. Science, 290(5500):2319–2323, 2000.
Gerald Tesauro. Temporal difference learning and TD-Gammon. Communications of the ACM, 38(3):58–68, 1995. doi:10.1145/203330.203343.
Lucien Tesnière. Éléments de syntaxe structurale. Klincksieck, Paris, 1959.
Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. BEIR: a heterogeneous benchmark for zero-shot evaluation of information retrieval models. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, volume 1. 2021. URL: https://datasets-benchmarks-proceedings.neurips.cc/paper/2021/hash/65b9eea6e1cc6bb9f0cd2a47751a186f-Abstract-round2.html.
Lucas Theis and Matthias Bethge. Generative image modeling using spatial LSTMs. In Advances in Neural Information Processing Systems (NeurIPS), volume 28. 2015.
Lucas Theis, Aäron van den Oord, and Matthias Bethge. A note on the evaluation of generative models. In International Conference on Learning Representations (ICLR). 2016.
Ken Thompson. Regular expression search algorithm. Communications of the ACM, 11(6):419–422, 1968. doi:10.1145/363347.363387.
William R. Thompson. On the likelihood that one unknown probability exceeds another in view of the evidence of two samples. Biometrika, 25(3–4):285–294, 1933.
Tristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh, Adina Williams, Douwe Kiela, and Candace Ross. Winoground: probing vision and language models for visio-linguistic compositionality. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 5238–5248. 2022.
Sunil Thulasidasan, Gopinath Chennupati, Jeff A. Bilmes, Tanmoy Bhattacharya, and Sarah Michalak. On mixup training: improved calibration and predictive uncertainty for deep neural networks. In Advances in Neural Information Processing Systems (NeurIPS). 2019.
Yuandong Tian, Xinlei Chen, and Surya Ganguli. Understanding self-supervised learning dynamics without contrastive pairs. In Proceedings of the 38th International Conference on Machine Learning (ICML), 10268–10278. 2021.
Robert Tibshirani, Guenther Walther, and Trevor Hastie. Estimating the number of clusters in a data set via the gap statistic. Journal of the Royal Statistical Society: Series B, 63(2):411–423, 2001.
Tijmen Tieleman. Training restricted Boltzmann machines using approximations to the likelihood gradient. In Proceedings of the 25th International Conference on Machine Learning (ICML). 2008. URL: https://dl.acm.org/doi/10.1145/1390156.1390290, doi:10.1145/1390156.1390290.
Tijmen Tieleman and Geoffrey Hinton. Lecture 6.5—rmsprop: divide the gradient by a running average of its recent magnitude. COURSERA: Neural Networks for Machine Learning, 2012.
Philippe Tillet. Introducing Triton: open-source GPU programming for neural networks. OpenAI, luglio 2021, 2021. URL: https://openai.com/index/triton/.
Philippe Tillet, H. T. Kung, and David Cox. Triton: an intermediate language and compiler for tiled neural network computations. In Proceedings of the 3rd ACM SIGPLAN International Workshop on Machine Learning and Programming Languages (MAPL), 10–19. 2019.
Michael E. Tipping and Christopher M. Bishop. Probabilistic principal component analysis. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 61(3):611–622, 1999. doi:10.1111/1467-9868.00196.
Michalis K. Titsias. Variational learning of inducing variables in sparse gaussian processes. In Proceedings of the Twelfth International Conference on Artificial Intelligence and Statistics (AISTATS), volume 5 of Proceedings of Machine Learning Research, 567–574. 2009.
Alexander Tong, Kilian Fatras, Nikolay Malkin, Guillaume Huguet, Yanlei Zhang, Jarrid Rector-Brooks, Guy Wolf, and Yoshua Bengio. Improving and generalizing flow-based generative models with minibatch optimal transport. Transactions on Machine Learning Research, 2024. URL: https://openreview.net/forum?id=CD9Snc73AW.
Shengbang Tong, Zhuang Liu, Yuexiang Zhai, Yi Ma, Yann LeCun, and Saining Xie. Eyes wide shut? exploring the visual shortcomings of multimodal LLMs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 9568–9578. 2024.
Jake Topping, Francesco Di Giovanni, Benjamin Paul Chamberlain, Xiaowen Dong, and Michael M. Bronstein. Understanding over-squashing and bottlenecks on graphs via curvature. In International Conference on Learning Representations. 2022.
Kristina Toutanova and Danqi Chen. Observed versus latent features for knowledge base and text inference. In Proceedings of the 3rd Workshop on Continuous Vector Space Models and their Compositionality. 2015.
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou. Training data-efficient image transformers & distillation through attention. In Proceedings of the 38th International Conference on Machine Learning (ICML), 10347–10357. 2021.
Hugo Touvron, Andrea Vedaldi, Matthijs Douze, and Hervé Jégou. Fixing the train-test resolution discrepancy. In Advances in Neural Information Processing Systems 32 (NeurIPS). 2019.
Ke Tran, Arianna Bisazza, and Christof Monz. The importance of being recurrent for modeling hierarchical structure. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP), 4731–4736. 2018.
Lloyd N. Trefethen. Approximation Theory and Approximation Practice. SIAM, Philadelphia, 2013.
Bill Triggs, Philip F. McLauchlan, Richard I. Hartley, and Andrew W. Fitzgibbon. Bundle adjustment: a modern synthesis. In Bill Triggs, Andrew Zisserman, and Richard Szeliski, editors, Vision Algorithms: Theory and Practice, volume 1883 of Lecture Notes in Computer Science, 298–372. Springer, 2000. doi:10.1007/3-540-44480-7_21.
Théo Trouillon, Johannes Welbl, Sebastian Riedel, Éric Gaussier, and Guillaume Bouchard. Complex embeddings for simple link prediction. In Proceedings of the 33rd International Conference on Machine Learning. 2016.
Michael Tschannen, Josip Djolonga, Paul K. Rubenstein, Sylvain Gelly, and Mario Lucic. On mutual information maximization for representation learning. In International Conference on Learning Representations (ICLR). 2020. URL: https://arxiv.org/abs/1907.13625.
Maria Tsimpoukelli, Jacob Menick, Serkan Cabi, S. M. Ali Eslami, Oriol Vinyals, and Felix Hill. Multimodal few-shot learning with frozen language models. In Advances in Neural Information Processing Systems (NeurIPS). 2021.
Dimitris Tsipras, Shibani Santurkar, Logan Engstrom, Alexander Turner, and Aleksander Madry. Robustness may be at odds with accuracy. In International Conference on Learning Representations (ICLR). 2019. URL: https://openreview.net/forum?id=SyxAb30cY7.
John N. Tsitsiklis. Asynchronous stochastic approximation and Q-learning. Machine Learning, 16(3):185–202, 1994.
John N. Tsitsiklis. On the convergence of optimistic policy iteration. Journal of Machine Learning Research, 3:59–72, 2002.
John N. Tsitsiklis and Benjamin Van Roy. An analysis of temporal-difference learning with function approximation. IEEE Transactions on Automatic Control, 42(5):674–690, 1997. doi:10.1109/9.580874.
Kathryn Tunyasuvunakool, Jonas Adler, Zachary Wu, Tim Green, Michal Zielinski, Augustin Žídek, Alex Bridgland, Andrew Cowie, Clemens Meyer, Agata Laydon, Sameer Velankar, Gerard J. Kleywegt, Alex Bateman, Richard Evans, Alexander Pritzel, Michael Figurnov, Olaf Ronneberger, Russ Bates, Simon A. A. Kohl, Anna Potapenko, Andrew J. Ballard, Bernardino Romera-Paredes, Stanislav Nikolov, Rishub Jain, Ellen Clancy, David Reiman, Stig Petersen, Andrew W. Senior, Koray Kavukcuoglu, Ewan Birney, Pushmeet Kohli, John Jumper, and Demis Hassabis. Highly accurate protein structure prediction for the human proteome. Nature, 596(7873):590–596, 2021. doi:10.1038/s41586-021-03828-1.
Alan M. Turing. On computable numbers, with an application to the Entscheidungsproblem. Proceedings of the London Mathematical Society, 42:230–265, 1936. Correzione in vol. 43 (1937), pp. 544–546.
Alan M. Turing. Computing machinery and intelligence. Mind, 59(236):433–460, 1950.
Miles Turpin, Julian Michael, Ethan Perez, and Samuel R. Bowman. Language models don't always say what they think: unfaithful explanations in chain-of-thought prompting. In Advances in Neural Information Processing Systems (NeurIPS), volume 36. 2023. URL: https://arxiv.org/abs/2305.04388.
Jan Tönshoff, Martin Ritzert, Eran Rosenbluth, and Martin Grohe. Where did the gap go? reassessing the long-range graph benchmark. Transactions on Machine Learning Research, 2024. URL: https://openreview.net/forum?id=Nm0WX86sKv, arXiv:2309.00367.
Benigno Uria, Iain Murray, and Hugo Larochelle. A deep and tractable density estimator. In Proceedings of the 31st International Conference on Machine Learning (ICML), volume 32 of Proceedings of Machine Learning Research, 467–475. 2014.
Keyon Vafa, Justin Y. Chen, Ashesh Rambachan, Jon Kleinberg, and Sendhil Mullainathan. Evaluating the world model implicit in a generative model. arXiv preprint arXiv:2406.03689, 2024. URL: https://arxiv.org/abs/2406.03689, arXiv:2406.03689.
Dani Valevski, Yaniv Leviathan, Moab Arar, and Shlomi Fruchter. Diffusion models are real-time game engines. In International Conference on Learning Representations (ICLR). 2025. URL: https://arxiv.org/abs/2408.14837.
Leslie G. Valiant. A theory of the learnable. Communications of the ACM, 27(11):1134–1142, 1984.
Rianne van den Berg, Thomas N. Kipf, and Max Welling. Graph convolutional matrix completion. arXiv preprint arXiv:1706.02263, 2017.
Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. Wavenet: a generative model for raw audio. arXiv preprint arXiv:1609.03499, 2016. URL: https://arxiv.org/abs/1609.03499.
Aäron van den Oord, Nal Kalchbrenner, and Koray Kavukcuoglu. Pixel recurrent neural networks. In International Conference on Machine Learning (ICML). 2016.
Aäron van den Oord, Nal Kalchbrenner, Oriol Vinyals, Lasse Espeholt, Alex Graves, and Koray Kavukcuoglu. Conditional image generation with PixelCNN decoders. In Advances in Neural Information Processing Systems (NeurIPS). 2016.
Aäron van den Oord, Yazhe Li, Igor Babuschkin, Karen Simonyan, Oriol Vinyals, Koray Kavukcuoglu, George van den Driessche, Edward Lockhart, Luis C. Cobo, Florian Stimberg, Norman Casagrande, Dominik Grewe, Seb Noury, Sander Dieleman, Erich Elsen, Nal Kalchbrenner, Heiga Zen, Alex Graves, Helen King, Tom Walters, Dan Belov, and Demis Hassabis. Parallel WaveNet: fast high-fidelity speech synthesis. In Proceedings of the 35th International Conference on Machine Learning. 2018.
Aäron van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018.
Aäron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu. Neural discrete representation learning. In Advances in Neural Information Processing Systems (NeurIPS), volume 30. 2017.
Laurens van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. Journal of Machine Learning Research, 9:2579–2605, 2008.
Hado van Hasselt. Double q-learning. In Advances in Neural Information Processing Systems, volume 23, 2613–2621. 2010.
Hado van Hasselt, Arthur Guez, and David Silver. Deep reinforcement learning with double q-learning. In Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence. 2016.
Vladimir N. Vapnik. The Nature of Statistical Learning Theory. Springer, New York, 1995.
Vladimir N. Vapnik and Alexey Ya. Chervonenkis. On the uniform convergence of relative frequencies of events to their probabilities. Theory of Probability and Its Applications, 16(2):264–280, 1971.
Vladimir N. Vapnik and Alexander Ya. Lerner. Pattern recognition using generalized portrait method. Automation and Remote Control, 24(6):774–780, 1963.
Sudhir Varma and Richard Simon. Bias in error estimation when using cross-validation for model selection. BMC Bioinformatics, 7:91, 2006. doi:10.1186/1471-2105-7-91.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems, volume 30. 2017.
Andrea Vattani. K-means requires exponentially many iterations even in the plane. Discrete & Computational Geometry, 45(4):596–616, 2011.
Petar Veličković, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio. Graph attention networks. In International Conference on Learning Representations (ICLR). 2018. URL: https://arxiv.org/abs/1710.10903.
Alexander Sasha Vezhnevets, Simon Osindero, Tom Schaul, Nicolas Heess, Max Jaderberg, David Silver, and Koray Kavukcuoglu. Feudal networks for hierarchical reinforcement learning. In Proceedings of the 34th International Conference on Machine Learning (ICML). 2017. URL: https://arxiv.org/abs/1703.01161.
Pablo Villalobos, Anson Ho, Jaime Sevilla, Tamay Besiroglu, Lennart Heim, and Marius Hobbhahn. Position: will we run out of data? Limits of LLM scaling based on human-generated data. In Proceedings of the 41st International Conference on Machine Learning (ICML), volume 235 of Proceedings of Machine Learning Research, 49523–49544. 2024. URL: https://proceedings.mlr.press/v235/villalobos24a.html.
James Vincent. How three French students used borrowed code to put the first AI portrait in Christie's. The Verge, October 2018. URL: https://www.theverge.com/2018/10/23/18013190/ai-art-portrait-auction-christies-belamy-obvious-robbie-barrat-gans.
Pascal Vincent. A connection between score matching and denoising autoencoders. Neural Computation, 2011. URL: https://direct.mit.edu/neco/article/23/7/1661/7677, doi:10.1162/NECO_a_00142.
Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol. Extracting and composing robust features with denoising autoencoders. In Proceedings of the 25th International Conference on Machine Learning (ICML), 1096–1103. 2008.
Nguyen Xuan Vinh, Julien Epps, and James Bailey. Information theoretic measures for clusterings comparison: variants, properties, normalization and correction for chance. Journal of Machine Learning Research, 11:2837–2854, 2010.
Oriol Vinyals, Igor Babuschkin, Wojciech M. Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H. Choi, Richard Powell, Timo Ewalds, Petko Georgiev, Junhyuk Oh, Dan Horgan, Manuel Kroiss, Ivo Danihelka, Aja Huang, Laurent Sifre, Trevor Cai, John P. Agapiou, Max Jaderberg, Alexander Sasha Vezhnevets, Rémi Leblond, Tobias Pohlen, Valentin Dalibard, David Budden, Yury Sulsky, James Molloy, Tom Le Paine, Çaglar Gülçehre, Ziyu Wang, Tobias Pfaff, Yuhuai Wu, Roman Ring, Dani Yogatama, Dario Wünsch, Katrina McKinney, Oliver Smith, Tom Schaul, Timothy P. Lillicrap, Koray Kavukcuoglu, Demis Hassabis, Chris Apps, and David Silver. Grandmaster level in starcraft ii using multi-agent reinforcement learning. Nature, 575:350–354, 2019.
Oriol Vinyals, Charles Blundell, Timothy Lillicrap, Koray Kavukcuoglu, and Daan Wierstra. Matching networks for one shot learning. In Advances in Neural Information Processing Systems 29 (NeurIPS), 3630–3638. 2016. URL: https://arxiv.org/abs/1606.04080.
Oriol Vinyals and Quoc V. Le. A neural conversational model. 2015. ICML Deep Learning Workshop. URL: https://arxiv.org/abs/1506.05869, arXiv:1506.05869.
Andrew J. Viterbi. Error bounds for convolutional codes and an asymptotically optimum decoding algorithm. IEEE Transactions on Information Theory, 13(2):260–269, 1967. doi:10.1109/TIT.1967.1054010.
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. Analyzing multi-head self-attention: specialized heads do the heavy lifting, the rest can be pruned. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 5797–5808. 2019. URL: https://aclanthology.org/P19-1580/.
Vasily Volkov. Better performance at lower occupancy. GPU Technology Conference (GTC 2010), 2010. Slide. URL: https://www.nvidia.com/content/GTC-2010/pdfs/2238_GTC2010.pdf.
Sandra Wachter, Brent Mittelstadt, and Chris Russell. Counterfactual explanations without opening the black box: automated decisions and the gdpr. Harvard Journal of Law & Technology, 31(2):841–887, 2017. URL: https://arxiv.org/abs/1711.00399.
Robert A. Wagner and Michael J. Fischer. The string-to-string correction problem. Journal of the ACM, 21(1):168–173, 1974. doi:10.1145/321796.321811.
Roger Waleffe, Wonmin Byeon, Duncan Riach, Brandon Norick, Vijay Korthikanti, Tri Dao, Albert Gu, Ali Hatamizadeh, Sudhakar Singh, Deepak Narayanan, Garvit Kulshreshtha, Vartika Singh, Jared Casper, Jan Kautz, Mohammad Shoeybi, and Bryan Catanzaro. An empirical study of Mamba-based language models. arXiv preprint arXiv:2406.07887, 2024.
Marilyn A. Walker, Diane J. Litman, Candace A. Kamm, and Alicia Abella. PARADISE: a framework for evaluating spoken dialogue agents. In Proceedings of the 35th Annual Meeting of the Association for Computational Linguistics (ACL), 271–280. 1997.
Bram Wallace, Meihua Dang, Rafael Rafailov, Linqi Zhou, Aaron Lou, Senthil Purushwalkam, Stefano Ermon, Caiming Xiong, Shafiq Joty, and Nikhil Naik. Diffusion model alignment using direct preference optimization. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2024. URL: https://arxiv.org/abs/2311.12908.
Chris S. Wallace and David M. Boulton. An information measure for classification. The Computer Journal, 11(2):185–194, 1968. doi:10.1093/comjnl/11.2.185.
Eric Wallace, Kai Xiao, Reimar Leike, Lilian Weng, Johannes Heidecke, and Alex Beutel. The instruction hierarchy: training LLMs to prioritize privileged instructions. arXiv preprint arXiv:2404.13208, 2024. URL: https://arxiv.org/abs/2404.13208.
Chengyi Wang, Sanyuan Chen, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, Lei He, Sheng Zhao, and Furu Wei. Neural codec language models are zero-shot text to speech synthesizers. arXiv preprint arXiv:2301.02111, 2023.
Ke Alexander Wang, Geoff Pleiss, Jacob R. Gardner, Stephen Tyree, Kilian Q. Weinberger, and Andrew Gordon Wilson. Exact Gaussian processes on a million data points. In Advances in Neural Information Processing Systems 32. 2019.
Kevin Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt. Interpretability in the wild: a circuit for indirect object identification in GPT-2 small. In International Conference on Learning Representations (ICLR). 2023. URL: https://arxiv.org/abs/2211.00593.
Peng Wang, Shuai Bai, Sinan Tan, and others. Qwen2-VL: enhancing vision-language model's perception of the world at any resolution. arXiv preprint arXiv:2409.12191, 2024.
Shuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii, and Jerome Revaud. DUSt3R: geometric 3D vision made easy. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 20697–20709. 2024. doi:10.1109/CVPR52733.2024.01956.
Sifan Wang, Shyam Sankaran, and Paris Perdikaris. Respecting causality for training physics-informed neural networks. Computer Methods in Applied Mechanics and Engineering, 421:116813, 2024. URL: https://arxiv.org/abs/2203.07404.
Sifan Wang, Yujun Teng, and Paris Perdikaris. Understanding and mitigating gradient flow pathologies in physics-informed neural networks. SIAM Journal on Scientific Computing, 43(5):A3055–A3081, 2021. doi:10.1137/20M1318043.
Sifan Wang, Hanwen Wang, and Paris Perdikaris. Learning the solution operator of parametric partial differential equations with physics-informed DeepONets. Science Advances, 7(40):eabi8605, 2021. doi:10.1126/sciadv.abi8605.
Sifan Wang, Xinling Yu, and Paris Perdikaris. When and why PINNs fail to train: a neural tangent kernel perspective. Journal of Computational Physics, 449:110768, 2022. URL: https://arxiv.org/abs/2007.14527.
Tongzhou Wang and Phillip Isola. Understanding contrastive representation learning through alignment and uniformity on the hypersphere. In Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, 9929–9939. PMLR, 2020. URL: https://proceedings.mlr.press/v119/wang20k.html.
Xiang Wang, Xiangnan He, Yixin Cao, Meng Liu, and Tat-Seng Chua. KGAT: knowledge graph attention network for recommendation. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (KDD '19), 950–958. ACM, 2019. doi:10.1145/3292500.3330989.
Xiang Wang, Xiangnan He, Meng Wang, Fuli Feng, and Tat-Seng Chua. Neural graph collaborative filtering. In Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval, 165–174. 2019.
Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He. Non-local neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 7794–7803. 2018.
Xinlong Wang, Xiaosong Zhang, Zhengxiong Luo, Quan Sun, Yufeng Cui, Jinsheng Wang, Fan Zhang, Yueze Wang, Zhen Li, Qiying Yu, Yingli Zhao, Yulong Ao, Xuebin Min, Tao Li, Boya Wu, Bo Zhao, Bowen Zhang, Liangdong Wang, Guang Liu, Zheqi He, Xi Yang, Jingjing Liu, Yonghua Lin, Tiejun Huang, and Zhongyuan Wang. Emu3: next-token prediction is all you need. arXiv preprint arXiv:2409.18869, 2024.
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. Self-consistency improves chain of thought reasoning in language models. In International Conference on Learning Representations (ICLR). 2023. URL: https://arxiv.org/abs/2203.11171.
Yun Wang, Juncheng Li, and Florian Metze. A comparison of five multiple instance learning pooling functions for sound event detection with weak labeling. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 31–35. 2019.
Zhou Wang, Alan C. Bovik, Hamid R. Sheikh, and Eero P. Simoncelli. Image quality assessment: from error visibility to structural similarity. IEEE Transactions on Image Processing, 13(4):600–612, 2004. doi:10.1109/TIP.2003.819861.
Ziyu Wang, Tom Schaul, Matteo Hessel, Hado van Hasselt, Marc Lanctot, and Nando de Freitas. Dueling network architectures for deep reinforcement learning. In Proceedings of the 33rd International Conference on Machine Learning (ICML). 2016.
Stanley L. Warner. Randomized response: a survey technique for eliminating evasive answer bias. Journal of the American Statistical Association, 60(309):63–69, 1965. doi:10.1080/01621459.1965.10480775.
Alex Warstadt, Leshem Choshen, Aaron Mueller, Adina Williams, Ethan Wilcox, and Chengxu Zhuang. Call for papers – the BabyLM challenge: sample-efficient pretraining on a developmentally plausible corpus. arXiv preprint arXiv:2301.11796, 2023.
Ronald L. Wasserstein and Nicole A. Lazar. The ASA's statement on p-values: context, process, and purpose. The American Statistician, 70(2):129–133, 2016.
Christopher J. C. H. Watkins. Learning from Delayed Rewards. PhD thesis, University of Cambridge, 1989.
Christopher J. C. H. Watkins and Peter Dayan. Q-learning. Machine Learning, 8:279–292, 1992.
Geoffrey S. Watson. Smooth regression analysis. Sankhyā: The Indian Journal of Statistics, Series A, 26(4):359–372, 1964.
Mark Weber, Giacomo Domeniconi, Jie Chen, Daniel Karl I. Weidele, Claudio Bellei, Tom Robinson, and Charles E. Leiserson. Anti-money laundering in bitcoin: experimenting with graph convolutional networks for financial forensics. In KDD '19 Workshop on Anomaly Detection in Finance. 2019. arXiv:1908.02591.
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. Jailbroken: how does LLM safety training fail? In Advances in Neural Information Processing Systems, volume 36. 2023. URL: https://arxiv.org/abs/2307.02483.
Greg C. G. Wei and Martin A. Tanner. A monte carlo implementation of the EM algorithm and the poor man's data augmentation algorithms. Journal of the American Statistical Association, 85(411):699–704, 1990. doi:10.1080/01621459.1990.10474930.
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. Emergent abilities of large language models. Transactions on Machine Learning Research (TMLR), 2022. URL: https://arxiv.org/abs/2206.07682, arXiv:2206.07682.
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems, volume 35. 2022. URL: https://arxiv.org/abs/2201.11903.
Jerry Wei, Jason Wei, Yi Tay, Dustin Tran, Albert Webson, Yifeng Lu, Xinyun Chen, Hanxiao Liu, Da Huang, Denny Zhou, and Tengyu Ma. Larger language models do in-context learning differently. arXiv preprint arXiv:2303.03846, 2023. URL: https://arxiv.org/abs/2303.03846.
Joseph Weizenbaum. ELIZA — a computer program for the study of natural language communication between man and machine. Communications of the ACM, 9(1):36–45, 1966.
Joseph Weizenbaum. Computer Power and Human Reason: From Judgment to Calculation. W. H. Freeman and Company, San Francisco, 1976.
Max Welling and Yee Whye Teh. Bayesian learning via stochastic gradient langevin dynamics. In Proceedings of the 28th International Conference on Machine Learning (ICML). 2011. URL: https://icml.cc/2011/papers/398_icmlpaper.pdf.
Li K. Wenliang and Heishiro Kanagawa. Blindness of score-based methods to isolated components and mixing proportions. arXiv preprint arXiv:2008.10087, 2020.
Paul J. Werbos. Beyond Regression: New Tools for Prediction and Analysis in the Behavioral Sciences. PhD thesis, Harvard University, 1974.
Paul J. Werbos. Applications of advances in nonlinear sensitivity analysis. In System Modeling and Optimization, volume 38 of Lecture Notes in Control and Information Sciences, pages 762–770. Springer, 1982.
Jason Weston, Sumit Chopra, and Antoine Bordes. Memory networks. arXiv preprint arXiv:1410.3916, 2014. International Conference on Learning Representations, 2015.
Rasmus Widing. PRP: product requirement prompts. Repository GitHub \texttt Wirasm/prp (già \texttt Wirasm/PRPs-agentic-eng), primo commit 11 giugno 2025, 2025. URL: Wirasm/prp.
Bernard Widrow and Marcian E. Hoff. Adaptive switching circuits. In IRE WESCON Convention Record, Part 4, 96–104. 1960.
Sarah Wiegreffe and Yuval Pinter. Attention is not not explanation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 11–20. 2019. URL: https://arxiv.org/abs/1908.04626.
Brandon T. Willard and Rémi Louf. Efficient guided generation for large language models. arXiv preprint arXiv:2307.09702, 2023.
Ronald J. Williams. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine Learning, 8:229–256, 1992.
Ronald J. Williams and David Zipser. A learning algorithm for continually running fully recurrent neural networks. Neural Computation, 1(2):270–280, 1989.
Samuel Williams, Andrew Waterman, and David Patterson. Roofline: an insightful visual performance model for multicore architectures. Communications of the ACM, 52(4):65–76, 2009.
Simon Willison. Prompt injection attacks against GPT-3. Blog personale, 12 settembre 2022, 2022. URL: https://simonwillison.net/2022/Sep/12/prompt-injection/.
Simon Willison. The lethal trifecta for AI agents: private data, untrusted content, and external communication. Blog personale, 16 giugno 2025, 2025. URL: https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/.
Shmuel Winograd. Arithmetic Complexity of Computations. Number 33 in CBMS-NSF Regional Conference Series in Applied Mathematics. SIAM, Philadelphia, 1980.
Patrick H. Winston. Learning: support vector machines (lezione 16). MIT 6.034 Artificial Intelligence, Fall 2010, MIT OpenCourseWare, 2010. URL: https://ocw.mit.edu/courses/6-034-artificial-intelligence-fall-2010/resources/lecture-16-learning-support-vector-machines/.
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. Transformers: state-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, 38–45. Association for Computational Linguistics, 2020. URL: https://aclanthology.org/2020.emnlp-demos.6.
David H. Wolpert. Stacked generalization. Neural Networks, 5(2):241–259, 1992. URL: https://doi.org/10.1016/S0893-6080(05)80023-1.
Danny Wood, Tingting Mu, Andrew M. Webb, Henry W. J. Reeve, Mikel Luján, and Gavin Brown. A unified theory of diversity in ensemble learning. Journal of Machine Learning Research, 24(359):1–49, 2023.
Chengyue Wu, Xiaokang Chen, Zhiyu Wu, Yiyang Ma, Xingchao Liu, Zizheng Pan, Wen Liu, Zhenda Xie, Xingkai Yu, Chong Ruan, and Ping Luo. Janus: decoupling visual encoding for unified multimodal understanding and generation. arXiv preprint arXiv:2410.13848, 2024.
Chengyue Wu, Hao Zhang, Shuchen Xue, Zhijian Liu, Shizhe Diao, Ligeng Zhu, Ping Luo, Song Han, and Enze Xie. Fast-dLLM: training-free acceleration of diffusion LLM by enabling KV cache and parallel decoding. arXiv preprint arXiv:2505.22618, 2025.
Felix Wu, Amauri H. Souza, Tianyi Zhang, Christopher Fifty, Tao Yu, and Kilian Q. Weinberger. Simplifying graph convolutional networks. In Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, 6861–6871. PMLR, 2019. arXiv:1902.07153.
Guanlong Wu, Zheng Zhang, Yao Zhang, Weili Wang, Jianyu Niu, Ye Wu, and Yinqian Zhang. I know what you asked: prompt leakage via KV-cache sharing in multi-tenant LLM serving. In Proceedings of the Network and Distributed System Security Symposium (NDSS 2025). San Diego, CA, USA, 2025. Internet Society. doi:10.14722/ndss.2025.241772.
Haixu Wu, Jiehui Xu, Jianmin Wang, and Mingsheng Long. Autoformer: decomposition transformers with auto-correlation for long-term series forecasting. In Advances in Neural Information Processing Systems (NeurIPS). 2021. URL: https://arxiv.org/abs/2106.13008.
Jeff Wu, Long Ouyang, Daniel M. Ziegler, Nisan Stiennon, Ryan Lowe, Jan Leike, and Paul Christiano. Recursively summarizing books with human feedback. arXiv preprint arXiv:2109.10862, 2021.
Junfeng Wu, Yi Jiang, Chuofan Ma, Yuliang Liu, Hengshuang Zhao, Zehuan Yuan, Song Bai, and Xiang Bai. Liquid: language models are scalable and unified multi-modal generators. arXiv preprint arXiv:2412.04332, 2024.
Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang Zhu, Li Jiang, Xiaoyun Zhang, Shaokun Zhang, Jiale Liu, Ahmed Hassan Awadallah, Ryen W. White, Doug Burger, and Chi Wang. AutoGen: enabling next-gen LLM applications via multi-agent conversations. In Conference on Language Modeling (COLM). 2024.
Shijie Wu and Mark Dredze. Beto, bentz, becas: the surprising cross-lingual effectiveness of bert. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 833–844. 2019.
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, and others. Google's neural machine translation system: bridging the gap between human and machine translation. arXiv preprint arXiv:1609.08144, 2016.
Yuhuai Wu, Elman Mansimov, Shun Liao, Alec Radford, and John Schulman. OpenAI baselines: ACKTR & A2C. OpenAI Blog, 2017. URL: https://openai.com/index/openai-baselines-acktr-a2c/.
Zonghan Wu, Shirui Pan, Fengwen Chen, Guodong Long, Chengqi Zhang, and Philip S. Yu. A comprehensive survey on graph neural networks. IEEE Transactions on Neural Networks and Learning Systems, 32(1):4–24, 2021. URL: https://arxiv.org/abs/1901.00596.
Wm. A. Wulf and Sally A. McKee. Hitting the memory wall: implications of the obvious. ACM SIGARCH Computer Architecture News, 23(1):20–24, 1995. doi:10.1145/216585.216588.
Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yiwen Ding, Boyang Hong, Ming Zhang, Junzhe Wang, Senjie Jin, Enyu Zhou, and others. The rise and potential of large language model based agents: a survey. arXiv preprint arXiv:2309.07864, 2023.
Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han. Smoothquant: accurate and efficient post-training quantization for large language models. In Proceedings of the 40th International Conference on Machine Learning (ICML), volume 202 of Proceedings of Machine Learning Research, 38087–38099. 2023.
Jianwen Xie, Yang Lu, Song-Chun Zhu, and Ying Nian Wu. A theory of generative ConvNet. In Proceedings of the 33rd International Conference on Machine Learning, volume 48 of Proceedings of Machine Learning Research, 2635–2644. PMLR, 2016.
Jinheng Xie, Weijia Mao, Zechen Bai, David Junhao Zhang, Weihao Wang, Kevin Qinghong Lin, Yuchao Gu, Zhijie Chen, Zhenheng Yang, and Mike Zheng Shou. Show-o: one single transformer to unify multimodal understanding and generation. In International Conference on Learning Representations (ICLR). 2025.
Derrick Xin, Behrooz Ghorbani, Ankush Garg, Orhan Firat, and Justin Gilmer. Do current multi-task optimization methods in deep learning even help? In Advances in Neural Information Processing Systems (NeurIPS), volume 35. 2022.
Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tie-Yan Liu. On layer normalization in the transformer architecture. In Proceedings of the 37th International Conference on Machine Learning (ICML), volume 119 of Proceedings of Machine Learning Research, 10524–10533. 2020.
Binfeng Xu, Zhiyuan Peng, Bowen Lei, Subhabrata Mukherjee, Yuchen Liu, and Dongkuan Xu. ReWOO: decoupling reasoning from observations for efficient augmented language models. arXiv preprint arXiv:2305.18323, 2023.
Keyulu Xu, Weihua Hu, Jure Leskovec, and Stefanie Jegelka. How powerful are graph neural networks? In International Conference on Learning Representations (ICLR). 2019. URL: https://arxiv.org/abs/1810.00826.
Keyulu Xu, Chengtao Li, Yonglong Tian, Tomohiro Sonobe, Ken-ichi Kawarabayashi, and Stefanie Jegelka. Representation learning on graphs with jumping knowledge networks. In Proceedings of the 35th International Conference on Machine Learning. 2018.
Shusheng Xu, Wei Fu, Jiaxuan Gao, Wenjie Ye, Weilin Liu, Zhiyu Mei, Guangju Wang, Chao Yu, and Yi Wu. Is DPO superior to PPO for LLM alignment? A comprehensive study. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, 54983–54998. PMLR, 2024. URL: https://proceedings.mlr.press/v235/xu24h.html.
Zhi-Qin John Xu, Yaoyu Zhang, Tao Luo, Yanyang Xiao, and Zheng Ma. Frequency Principle: Fourier analysis sheds light on deep neural networks. Communications in Computational Physics, 28(5):1746–1767, 2020. doi:10.4208/cicp.OA-2020-0085.
Linting Xue, Aditya Barua, Noah Constant, Rami Al-Rfou, Sharan Narang, Mihir Kale, Adam Roberts, and Colin Raffel. ByT5: towards a token-free future with pre-trained byte-to-byte models. Transactions of the Association for Computational Linguistics, 10:291–306, 2022.
Shi-Qi Yan, Jia-Chen Gu, Yun Zhu, and Zhen-Hua Ling. Corrective retrieval augmented generation. 2024. arXiv:2401.15884.
Bishan Yang, Wen-tau Yih, Xiaodong He, Jianfeng Gao, and Li Deng. Embedding entities and relations for learning and inference in knowledge bases. In International Conference on Learning Representations. 2015.
Chengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu, Quoc V. Le, Denny Zhou, and Xinyun Chen. Large language models as optimizers. In International Conference on Learning Representations (ICLR). 2024. arXiv:2309.03409.
Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao. Depth anything: unleashing the power of large-scale unlabeled data. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2024. URL: https://arxiv.org/abs/2401.10891.
Sherry Yang, Yilun Du, Seyed Kamyar Seyed Ghasemipour, Jonathan Tompson, Leslie Kaelbling, Dale Schuurmans, and Pieter Abbeel. Learning interactive real-world simulators. In International Conference on Learning Representations (ICLR). 2024. URL: https://arxiv.org/abs/2310.06114, arXiv:2310.06114.
Shu-wen Yang, Po-Han Chi, Yung-Sung Chuang, Cheng-I Jeff Lai, Kushal Lakhotia, Yist Y. Lin, Andy T. Liu, Jiatong Shi, Xuankai Chang, Guan-Ting Lin, Tzu-Hsien Huang, Wei-Cheng Tseng, Ko-tik Lee, Da-Rong Liu, Zili Huang, Shuyan Dong, Shang-Wen Li, Shinji Watanabe, Abdelrahman Mohamed, and Hung-yi Lee. SUPERB: speech processing universal PERformance benchmark. In Interspeech, 1194–1198. 2021. doi:10.21437/Interspeech.2021-1775.
Songlin Yang, Jan Kautz, and Ali Hatamizadeh. Gated delta networks: improving Mamba2 with delta rule. In International Conference on Learning Representations (ICLR). 2025. arXiv:2412.06464.
Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda, and Yoon Kim. Gated linear attention transformers with hardware-efficient training. In Proceedings of the 41st International Conference on Machine Learning (ICML). 2024.
Songlin Yang, Bailin Wang, Yu Zhang, Yikang Shen, and Yoon Kim. Parallelizing linear transformers with the delta rule over sequence length. In Advances in Neural Information Processing Systems (NeurIPS). 2024.
Shunyu Yao, Noah Shinn, Pedram Razavi, and Karthik Narasimhan. τ-bench: a benchmark for tool-agent-user interaction in real-world domains. arXiv preprint arXiv:2406.12045, 2024.
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan. Tree of thoughts: deliberate problem solving with large language models. In Advances in Neural Information Processing Systems (NeurIPS), volume 36. 2023.
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. ReAct: synergizing reasoning and acting in language models. In International Conference on Learning Representations (ICLR). 2023.
Xinyang Yi, Ji Yang, Lichan Hong, Derek Zhiyuan Cheng, Lukasz Heldt, Aditee Kumthekar, Zhe Zhao, Li Wei, and Ed Chi. Sampling-bias-corrected neural modeling for large corpus item recommendations. In Proceedings of the 13th ACM Conference on Recommender Systems (RecSys '19), 269–277. 2019. doi:10.1145/3298689.3346996.
Penghang Yin, Jiancheng Lyu, Shuai Zhang, Stanley Osher, Yingyong Qi, and Jack Xin. Understanding straight-through estimator in training activation quantized neural nets. In International Conference on Learning Representations (ICLR). 2019. arXiv:1903.05662.
Shukang Yin, Chaoyou Fu, Sirui Zhao, Tong Xu, Hao Wang, Dianbo Sui, Yunhang Shen, Ke Li, Xing Sun, and Enhong Chen. Woodpecker: hallucination correction for multimodal large language models. arXiv preprint arXiv:2310.16045, 2023.
Tianwei Yin, Michaël Gharbi, Taesung Park, Richard Zhang, Eli Shechtman, Frédo Durand, and William T. Freeman. Improved distribution matching distillation for fast image synthesis. In Advances in Neural Information Processing Systems, volume 37. 2024. URL: https://arxiv.org/abs/2405.14867.
Tianwei Yin, Michaël Gharbi, Richard Zhang, Eli Shechtman, Frédo Durand, William T. Freeman, and Taesung Park. One-step diffusion with distribution matching distillation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2024. URL: https://arxiv.org/abs/2311.18828.
Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng, Guolin Ke, Di He, Yanming Shen, and Tie-Yan Liu. Do transformers really perform bad for graph representation? In Advances in Neural Information Processing Systems (NeurIPS). 2021. URL: https://arxiv.org/abs/2106.05234.
Rex Ying, Ruining He, Kaifeng Chen, Pong Eksombatchai, William L. Hamilton, and Jure Leskovec. Graph convolutional neural networks for web-scale recommender systems. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 974–983. 2018.
Zheng-Xin Yong, Cristina Menghini, and Stephen H. Bach. Low-resource languages jailbreak GPT-4. In NeurIPS Workshop on Socially Responsible Language Modelling Research (SoLaR). 2023. URL: https://arxiv.org/abs/2310.02446, arXiv:2310.02446.
Jason Yosinski, Jeff Clune, Yoshua Bengio, and Hod Lipson. How transferable are features in deep neural networks? In Advances in Neural Information Processing Systems, volume 27. 2014.
Laurent Younes. On the convergence of Markovian stochastic algorithms with rapidly decreasing ergodicity rates. Stochastics and Stochastic Reports, 65(3-4):177–228, 1999.
Alex Yu, Vickie Ye, Matthew Tancik, and Angjoo Kanazawa. pixelNeRF: neural radiance fields from one or few images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 4576–4585. 2021. doi:10.1109/CVPR46437.2021.00455.
Chao Yu, Akash Velu, Eugene Vinitsky, Jiaxuan Gao, Yu Wang, Alexandre Bayen, and Yi Wu. The surprising effectiveness of ppo in cooperative, multi-agent games. In Advances in Neural Information Processing Systems 35, Datasets and Benchmarks Track (NeurIPS). 2022.
Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun. Orca: a distributed serving system for Transformer-Based generative models. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22), 521–538. Carlsbad, CA, 2022. USENIX Association. URL: https://www.usenix.org/conference/osdi22/presentation/yu.
Lijun Yu, José Lezama, Nitesh B. Gundavarapu, Luca Versari, Kihyuk Sohn, David Minnen, Yong Cheng, Vighnesh Birodkar, Agrim Gupta, Xiuye Gu, Alexander G. Hauptmann, Boqing Gong, Ming-Hsuan Yang, Irfan Essa, David A. Ross, and Lu Jiang. Language model beats diffusion – tokenizer is key to visual generation. In International Conference on Learning Representations (ICLR). 2024.
Lili Yu, Dániel Simig, Colin Flaherty, Armen Aghajanyan, Luke Zettlemoyer, and Mike Lewis. MEGABYTE: predicting million-byte sequences with multiscale transformers. In Advances in Neural Information Processing Systems (NeurIPS). 2023. arXiv:2305.07185.
Qifan Yu, Juncheng Li, Longhui Wei, Liang Pang, Wentao Ye, Bosheng Qin, Siliang Tang, Qi Tian, and Yueting Zhuang. HalluciDoctor: mitigating hallucinatory toxicity in visual instruction data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2024.
Tianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine, Karol Hausman, and Chelsea Finn. Gradient surgery for multi-task learning. In Advances in Neural Information Processing Systems (NeurIPS). 2020. URL: https://arxiv.org/abs/2001.06782.
Ming Yuan and Yi Lin. Model selection and estimation in regression with grouped variables. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 68(1):49–67, 2006.
Yang Yue, Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang, Yang Yue, Shiji Song, and Gao Huang. Does reinforcement learning really incentivize reasoning capacity in LLMs beyond the base model? In Advances in Neural Information Processing Systems (NeurIPS). 2025. URL: https://arxiv.org/abs/2504.13837.
Mert Yuksekgonul, Federico Bianchi, Pratyusha Kalluri, Dan Jurafsky, and James Zou. When and why vision-language models behave like bags-of-words, and what to do about it? In International Conference on Learning Representations (ICLR). 2023.
Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo. CutMix: regularization strategy to train strong classifiers with localizable features. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2019. arXiv:1905.04899.
Bilal Yurdakul and Joshua Naranjo. Statistical properties of the population stability index. Journal of Risk Model Validation, 14(4):89–100, 2020. doi:10.21314/JRMV.2020.227.
Bianca Zadrozny and Charles Elkan. Transforming classifier scores into accurate multiclass probability estimates. In Proceedings of the Eighth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD), 694–699. 2002. doi:10.1145/775047.775151.
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontañón, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, and Amr Ahmed. Big bird: transformers for longer sequences. In Advances in Neural Information Processing Systems, volume 33. 2020.
Jure Zbontar, Li Jing, Ishan Misra, Yann LeCun, and Stéphane Deny. Barlow twins: self-supervised learning via redundancy reduction. In International Conference on Machine Learning (ICML). 2021. URL: https://arxiv.org/abs/2103.03230.
Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi. SoundStream: an end-to-end neural audio codec. IEEE/ACM Transactions on Audio, Speech, and Language Processing (TASLP), 30:495–507, 2022.
Matthew D. Zeiler. ADADELTA: an adaptive learning rate method. arXiv preprint arXiv:1212.5701, 2012.
Matthew D. Zeiler and Rob Fergus. Visualizing and understanding convolutional networks. In European Conference on Computer Vision (ECCV). 2014.
Heiga Zen, Andrew Senior, and Mike Schuster. Statistical parametric speech synthesis using deep neural networks. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 7962–7966. 2013.
Heiga Zen, Keiichi Tokuda, and Alan W. Black. Statistical parametric speech synthesis. Speech Communication, 51(11):1039–1064, 2009.
Ailing Zeng, Muxi Chen, Lei Zhang, and Qiang Xu. Are transformers effective for time series forecasting? In Proceedings of the AAAI Conference on Artificial Intelligence. 2023. URL: https://arxiv.org/abs/2205.13504.
Shuangfei Zhai, Ruixiang Zhang, Preetum Nakkiran, David Berthelot, Jiatao Gu, Huangjie Zheng, Tianrong Chen, Miguel Angel Bautista, Navdeep Jaitly, and Josh Susskind. Normalizing flows are capable generative models. In Proceedings of the 42nd International Conference on Machine Learning (ICML), volume 267 of Proceedings of Machine Learning Research. 2025.
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid loss for language image pre-training. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2023. arXiv:2303.15343.
Biao Zhang and Rico Sennrich. Root mean square layer normalization. In Advances in Neural Information Processing Systems 32 (NeurIPS 2019). 2019.
Brian Hu Zhang, Blake Lemoine, and Margaret Mitchell. Mitigating unwanted biases with adversarial learning. In AAAI/ACM Conference on AI, Ethics, and Society (AIES). 2018. URL: https://arxiv.org/abs/1801.07593.
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals. Understanding deep learning requires rethinking generalization. In International Conference on Learning Representations (ICLR). 2017. URL: https://arxiv.org/abs/1611.03530.
Han Zhang, Ian Goodfellow, Dimitris Metaxas, and Augustus Odena. Self-attention generative adversarial networks. In Proceedings of the 36th International Conference on Machine Learning (ICML), volume 97 of PMLR. 2019. arXiv:1805.08318.
Hanlin Zhang, Benjamin L. Edelman, Danilo Francati, Daniele Venturi, Giuseppe Ateniese, and Boaz Barak. Watermarks in the sand: impossibility of strong watermarking for generative models. arXiv preprint arXiv:2311.04378, 2023.
Hongyi Zhang, Moustapha Cissé, Yann N. Dauphin, and David Lopez-Paz. Mixup: beyond empirical risk minimization. In International Conference on Learning Representations. 2018.
Kai Zhang, Gernot Riegler, Noah Snavely, and Vladlen Koltun. NeRF++: analyzing and improving neural radiance fields. arXiv preprint arXiv:2010.07492, 2020.
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. In IEEE/CVF International Conference on Computer Vision (ICCV). 2023. arXiv:2302.05543.
Michael Zhang, Kush Bhatia, Hermann Kumbong, and Christopher Ré. The hedgehog & the porcupine: expressive linear attentions with softmax mimicry. In International Conference on Learning Representations (ICLR). 2024. arXiv:2402.04347.
Qinsheng Zhang and Yongxin Chen. Fast sampling of diffusion models with exponential integrator. In International Conference on Learning Representations (ICLR). 2023. URL: https://arxiv.org/abs/2204.13902.
Richard Zhang. Making convolutional networks shift-invariant again. In International Conference on Machine Learning (ICML), volume 97 of PMLR, 7324–7334. 2019.
Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 586–595. 2018. doi:10.1109/CVPR.2018.00068.
Shengjia Zhao, Jiaming Song, and Stefano Ermon. Towards a deeper understanding of variational autoencoding models. arXiv preprint arXiv:1702.08658, 2017.
Wenliang Zhao, Lujia Bai, Yongming Rao, Jie Zhou, and Jiwen Lu. UniPC: a unified predictor-corrector framework for fast sampling of diffusion models. In Advances in Neural Information Processing Systems, volume 36. 2023. URL: https://arxiv.org/abs/2302.04867.
Yanli Zhao, Andrew Gu, Rohan Varma, Liang Luo, Chien-Chin Huang, Min Xu, Less Wright, Hamid Shojanazeri, Myle Ott, Sam Shleifer, Alban Desmaison, Can Balioglu, Pritam Damania, Bernard Nguyen, Geeta Chauhan, Yuchen Hao, Ajit Mathews, and Shen Li. PyTorch FSDP: experiences on scaling fully sharded data parallel. Proceedings of the VLDB Endowment, 2023.
Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. Calibrate before use: improving few-shot performance of language models. In Proceedings of the 38th International Conference on Machine Learning (ICML), 12697–12706. 2021.
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging LLM-as-a-judge with MT-bench and chatbot arena. In Advances in Neural Information Processing Systems (NeurIPS), volume 36. 2023.
Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, and Ying Sheng. SGLang: efficient execution of structured language model programs. In Advances in Neural Information Processing Systems, volume 37. 2024.
Mingqian Zheng, Jiaxin Pei, Lajanugen Logeswaran, Moontae Lee, and David Jurgens. When “a helpful assistant” is not really helpful: personas in system prompts do not improve performances of large language models. In Findings of the Association for Computational Linguistics: EMNLP 2024. 2024. arXiv:2311.10054.
Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang. DistServe: disaggregating prefill and decoding for goodput-optimized large language model serving. In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI), 193–210. 2024.
Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, and Omer Levy. LIMA: less is more for alignment. In Advances in Neural Information Processing Systems, volume 36, 55006–55021. 2023. doi:10.52202/075280-2400.
Chunting Zhou, Lili Yu, Arun Babu, Kushal Tirumala, Michihiro Yasunaga, Leonid Shamis, Jacob Kahn, Xuezhe Ma, Luke Zettlemoyer, and Omer Levy. Transfusion: predict the next token and diffuse images with one multi-modal model. arXiv preprint arXiv:2408.11039, 2024.
Haoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang, Jianxin Li, Hui Xiong, and Wancai Zhang. Informer: beyond efficient transformer for long sequence time-series forecasting. In Proceedings of the AAAI Conference on Artificial Intelligence. 2021. URL: https://arxiv.org/abs/2012.07436.
Tian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang, Liang Sun, and Rong Jin. Fedformer: frequency enhanced decomposed transformer for long-term series forecasting. In Proceedings of the 39th International Conference on Machine Learning (ICML). 2022. URL: https://arxiv.org/abs/2201.12740.
Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, and Jimmy Ba. Large language models are human-level prompt engineers. In International Conference on Learning Representations (ICLR). 2023. arXiv:2211.01910.
Yunhong Zhou, Dennis Wilkinson, Robert Schreiber, and Rong Pan. Large-scale parallel collaborative filtering for the netflix prize. In Rudolf Fleischer and Jinhui Xu, editors, Algorithmic Aspects in Information and Management (AAIM 2008), volume 5034 of Lecture Notes in Computer Science, 337–348. Berlin, Heidelberg, 2008. Springer. doi:10.1007/978-3-540-68880-8_32.
Jiong Zhu, Yujun Yan, Lingxiao Zhao, Mark Heimann, Leman Akoglu, and Danai Koutra. Beyond homophily in graph neural networks: current limitations and effective designs. In Advances in Neural Information Processing Systems, volume 33. 2020. arXiv:2006.11468.
Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A. Efros. Unpaired image-to-image translation using cycle-consistent adversarial networks. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2223–2232. 2017.
Ligeng Zhu, Zhijian Liu, and Song Han. Deep leakage from gradients. In Advances in Neural Information Processing Systems (NeurIPS), volume 32. 2019.
Brian D. Ziebart, Andrew Maas, J. Andrew Bagnell, and Anind K. Dey. Maximum entropy inverse reinforcement learning. In Proceedings of the Twenty-Third AAAI Conference on Artificial Intelligence, 1433–1438. 2008.
Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. Fine-tuning language models from human preferences. arXiv preprint arXiv:1909.08593, 2019.
Martin Zinkevich. Online convex programming and generalized infinitesimal gradient ascent. In Proceedings of the 20th International Conference on Machine Learning (ICML), 928–936. 2003.
Martin Zinkevich, Michael Johanson, Michael Bowling, and Carmelo Piccione. Regret minimization in games with incomplete information. In Advances in Neural Information Processing Systems 20, 1729–1736. 2008.
Barret Zoph, Irwan Bello, Sameer Kumar, Nan Du, Yanping Huang, Jeff Dean, Noam Shazeer, and William Fedus. ST-MoE: designing stable and transferable sparse expert models. arXiv preprint arXiv:2202.08906, 2022.
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter, and Matt Fredrikson. Universal and transferable adversarial attacks on aligned language models. Preprint arXiv:2307.15043, 2023. URL: https://arxiv.org/abs/2307.15043.
Hui Zou and Trevor Hastie. Regularization and variable selection via the elastic net. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 67(2):301–320, 2005.
Vilém Zouhar, Clara Meister, Juan Luis Gastaldi, Li Du, Tim Vieira, Mrinmaya Sachan, and Ryan Cotterell. A formal perspective on byte-pair encoding. In Findings of the Association for Computational Linguistics: ACL 2023, 598–614. 2023. doi:10.18653/v1/2023.findings-acl.38.
Bernt Øksendal. Stochastic Differential Equations: An Introduction with Applications. Springer, Berlin, 6 edition, 2003.
Erik Štrumbelj and Igor Kononenko. An efficient explanation of individual classifications using game theory. Journal of Machine Learning Research, 11:1–18, 2010.
Erik Štrumbelj and Igor Kononenko. Explaining prediction models and individual predictions with feature contributions. Knowledge and Information Systems, 41(3):647–665, 2014.
Anthropic. Introducing the model context protocol. November 2024. Pubblicato il 25 novembre 2024. Consultato il 2 ottobre 2026. URL: https://www.anthropic.com/news/model-context-protocol.
Anthropic. Effective context engineering for AI agents. Anthropic Engineering, 29 settembre 2025, 2025. URL: https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents.
Anthropic. Tracing the thoughts of a large language model. 2025. Il resoconto divulgativo dello studio \emph On the Biology of a Large Language Model. URL: https://www.anthropic.com/research/tracing-thoughts-language-model.
Center for AI Safety. Statement on AI risk. Dichiarazione pubblica, 30 maggio 2023, May 2023. Consultato il 2 ottobre 2026. URL: https://safe.ai/work/statement-on-ai-extinction-risk.
Chameleon Team. Chameleon: mixed-modal early-fusion foundation models. arXiv preprint arXiv:2405.09818, 2024.
DAIR.AI. Prompt engineering guide. Guida online, https://www.promptingguide.ai, 2024. consultata nell'agosto 2026.
DeepSeek-AI. DeepSeek-V2: a strong, economical, and efficient mixture-of-experts language model. arXiv preprint arXiv:2405.04434, 2024. URL: https://arxiv.org/abs/2405.04434.
DeepSeek-AI. DeepSeek-V3 technical report. arXiv preprint arXiv:2412.19437, 2024.
DiffusionGemma Team. DiffusionGemma technical report. arXiv preprint arXiv:2608.00146, 2026.
Equal Employment Opportunity Commission and Civil Service Commission and Department of Labor and Department of Justice. Uniform guidelines on employee selection procedures (1978). 29 C.F.R. Part 1607; 43 Fed. Reg. 38295 (Aug. 25, 1978), 1978. La regola dei quattro quinti è al § 1607.4(D). URL: https://www.law.cornell.edu/cfr/text/29/1607.4.
European Data Protection Board. Opinion 28/2024 on certain data protection aspects related to the processing of personal data in the context of AI models. Parere adottato ai sensi dell'art. 64, par. 2, GDPR, December 2024. Adottato il 17 dicembre 2024. URL: https://www.edpb.europa.eu/system/files/documents/2024-12/edpb_opinion_202428_ai-models_en.pdf.
European Parliament and Council. Regulation (eu) 2024/1689 laying down harmonised rules on artificial intelligence (artificial intelligence act). Official Journal of the European Union, 2024. URL: https://eur-lex.europa.eu/eli/reg/2024/1689/oj.
Foundation for Intelligent Physical Agents. FIPA communicative act library specification. Standard SC00037J, Foundation for Intelligent Physical Agents, 2002. URL: http://www.fipa.org/specs/fipa00037/SC00037J.pdf.
International Energy Agency. Energy and AI. World Energy Outlook Special Report, International Energy Agency, Paris, April 2025. URL: https://www.iea.org/reports/energy-and-ai.
ITU-R. Method for the subjective assessment of intermediate quality level of audio systems. Recommendation ITU-R BS.1534-3, International Telecommunication Union, Radiocommunication Sector, October 2015.
Kimi Team. Kimi Linear: an expressive, efficient attention architecture. arXiv preprint arXiv:2510.26692, 2025.
LangChain. Context engineering for agents. Blog di LangChain, 2 luglio 2025, 2025. URL: https://www.langchain.com/blog/context-engineering-for-agents.
Leonardo Pisano. Il Liber abbaci di Leonardo Pisano. Volume 1 of Scritti di Leonardo Pisano matematico del secolo decimoterzo. Tipografia delle Scienze Matematiche e Fisiche, Roma, 1857. Seconda redazione (1228) del testo del 1202; il problema dei conigli alle pp. 283–284.
MiniMax. Why did MiniMax M2 end up as a full attention model? Articolo sul blog di Hugging Face, 30 ottobre 2025, 2025. URL: https://huggingface.co/blog/MiniMax-AI/why-did-m2-end-up-as-a-full-attention-model.
MiniMax. MiniMax-01: scaling foundation models with lightning attention. arXiv preprint arXiv:2501.08313, 2025.
Model Context Protocol. Model context protocol specification, versione 2025-06-18. 2025. Consultata il 2 ottobre 2026; la versione corrente, 2026-07-28, conserva i punti citati. URL: https://modelcontextprotocol.io/specification/2025-06-18.
OpenAI. Learning to reason with LLMs. Annuncio di o1, 12 settembre 2024, 2024. URL: https://openai.com/index/learning-to-reason-with-llms/.
OpenAI. Why SWE-bench verified no longer measures frontier coding capabilities. February 2026. Pubblicato il 23 febbraio 2026. Consultato il 2 ottobre 2026. URL: https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/.
OWASP Gen AI Security Project. LLM01:2025 prompt injection. OWASP Top 10 for LLM Applications 2025, 2025. Consultato il 2 ottobre 2026. URL: https://genai.owasp.org/llmrisk/llm01-prompt-injection/.
Parlamento europeo e Consiglio dell'Unione europea. Regolamento (ue) 2016/679 del parlamento europeo e del consiglio del 27 aprile 2016 relativo alla protezione delle persone fisiche con riguardo al trattamento dei dati personali, nonché alla libera circolazione di tali dati e che abroga la direttiva 95/46/ce (regolamento generale sulla protezione dei dati). Gazzetta ufficiale dell'Unione europea, L 119, 4.5.2016, pp. 1–88, 2016. Applicabile dal 25 maggio 2018 (art. 99). URL: https://eur-lex.europa.eu/eli/reg/2016/679/oj/ita.
Parlamento europeo e Consiglio dell'Unione europea. Regolamento (ue) 2026/1744 del parlamento europeo e del consiglio dell'8 luglio 2026 che modifica i regolamenti (ue) 2024/1689, (ue) 2018/1139 e (ue) 2023/1230 per quanto riguarda la semplificazione dell'attuazione di regole armonizzate sull'intelligenza artificiale (omnibus digitale sull'ia). Gazzetta ufficiale dell'Unione europea, L, 2026/1744, 24.7.2026, 2026. In vigore dal 27 luglio 2026 (art. 4). URL: https://eur-lex.europa.eu/eli/reg/2026/1744/oj/ita.
Qwen Team. Qwen3-Next-80B-A3B-Instruct. Scheda del modello su Hugging Face, creata il 9 settembre 2025, 2025. URL: https://huggingface.co/Qwen/Qwen3-Next-80B-A3B-Instruct.