Publications
Journal papers, conference papers, and technical reports.
Journal Papers
- Akari Asai, Jacqueline He, Rulin Shao, Weijia Shi, Amanpreet Singh, Joseph Chee Chang, Kyle Lo, Luca Soldaini, Sergey Feldman, et al., “Synthesizing scientific literature with retrieval-augmented language models,” Nature 650: 857–863, 2026.
- Katherine Cohen, Andy Rapoport, Elsa Friis, Shannon Hill, Sergey Feldman, and Jessica Schleider, “The Alongside Digital Wellness Program for Youth: Longitudinal Pre-Post Outcomes Study,” JMIR Formative Research 9:e73180, 2025.
- Katherine A. Keith, Sergey Feldman, David Jurgens, Jonathan Bragg, and Rohit Bhattacharya, “RCT Rejection Sampling for Causal Estimation Evaluation,” Transactions on Machine Learning Research, 2023.
- The Childhood Acute Illness and Nutrition (CHAIN) Network, including Sergey Feldman, “Characterising paediatric mortality during and after acute illness in Sub-Saharan Africa and South Asia: a secondary analysis of the CHAIN cohort using a machine learning approach,” EClinicalMedicine 57:101838, 2023.
- Rodney Kinney, Chloe Anastasiades, Russell Authur, Iz Beltagy, Jonathan Bragg, Alexandra Buraczynski, Sergey Feldman, et al., “The Semantic Scholar Open Data Platform,” arXiv preprint, 2023.
- Michael Cafarella, Michael Anderson, Iz Beltagy, Arie Cattan, Sarah Chasins, Ido Dagan, Doug Downey, Oren Etzioni, Sergey Feldman, Tian Gao, Tom Hope, Kexin Huang, Sophie Johnson, Daniel King, Kyle Lo, Yuze Lou, Matthew Shapiro, Dinghao Shen, Shivashankar Subramanian, Lucy Lu Wang, Yuning Wang, Yitong Wang, Daniel S. Weld, Jenny Vo-Phamhi, Anna Zeng, and Jiayun Zou, “Infrastructure for rapid open knowledge network development”, AI Magazine 43: 59–68, 2022.
- Sean MacAvaney, Sergey Feldman, Nazli Goharian, Doug Downey, and Arman Cohan, “ABNIRML: Analyzing the Behavior of Neural IR Models,” Transactions of the Association for Computational Linguistics, 2022.
- Sergey Feldman, Waleed Ammar, Kyle Lo, Elly Trepman, Madeleine van Zuylen, and Oren Etzioni, “Quantifying Sex Bias in Clinical Studies at Scale With Automated Data Extraction,” JAMA Network Open, 2019.
- Sergey Feldman, Maya R. Gupta, and Bela A. Frigyik, “Revisiting Stein’s Paradox: Multi-Task Averaging,” Journal of Machine Learning Research, 2014.
- Eric K. Garcia, Sergey Feldman, Maya R. Gupta, and Santosh Srivastava, “Completely Lazy Learning,” IEEE Trans. on Knowledge and Data Engineering, 2010.
- Vagisha Sharma, Jimmy K. Eng, Sergey Feldman, Priska von Haller, Michael J. MacCoss, and William S. Noble, “Precursor Charge State Prediction for Electron Transfer Dissociation Tandem Mass Spectra,” Journal of Proteome Research, 2010.
Conference Papers & Preprints
- Nishant Balepur, Malachi Hamada, Varsha Kishore, Sergey Feldman, Amanpreet Singh, Pao Siangliulue, Joseph Chee Chang, Eunsol Choi, Jordan Lee Boyd-Graber, and Aakanksha Naik, “Language Models Don’t Know What You Want: Evaluating Personalization in Deep Research Needs Real Users,” ACL, 2026.
- Jonathan Bragg, Mike D’Arcy, Nishant Balepur, Dan Bareket, Bhavana Dalvi Mishra, Sergey Feldman, et al., “AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite,” ICLR, 2026.
- Nishant Balepur, Malachi Hamada, Varsha Kishore, Sergey Feldman, et al., “DRACULA: Hunting for the Actions Users Want Deep Research Agents to Execute,” arXiv preprint arXiv:2604.23815, 2026.
- Jena D. Hwang, Varsha Kishore, Amanpreet Singh, Dany Haddad, Aakanksha Naik, Malachi Hamada, Jonathan Bragg, Mike D’Arcy, Daniel S. Weld, Lucy Lu Wang, Doug Downey, and Sergey Feldman, “Deep Research, Shallow Evaluation: A Case Study in Meta-Evaluation for Long-Form QA Benchmarks,” arXiv preprint arXiv:2603.06942, 2026.
- Dany Haddad, Dan Bareket, Joseph Chee Chang, Jay DeYoung, Jena D. Hwang, Uri Katz, Mark Polak, Sangho Suh, Harshit Surana, Aryeh Tiktinsky, Shriya Atmakuri, Jonathan Bragg, Mike D’Arcy, Sergey Feldman, et al., “Understanding Usage and Engagement in AI-Powered Scientific Research Tools: The Asta Interaction Dataset,” arXiv preprint arXiv:2602.23335, 2026.
- Amanpreet Singh, Joseph Chee Chang, Dany Haddad, Aakanksha Naik, Jena D. Hwang, Rodney Kinney, Daniel S. Weld, Doug Downey, and Sergey Feldman, “Ai2 Scholar QA: Organized Literature Synthesis with Attribution,” ACL System Demonstrations, 2025.
- John Giorgi, Amanpreet Singh, Doug Downey, Sergey Feldman, and Lucy Wang, “TOPICAL: TOPIC Pages AutomagicaLly,” NAACL System Demonstrations, 2024.
- Monica Munnangi, Sergey Feldman, Byron C. Wallace, Silvio Amir, Tom Hope, and Aakanksha Naik, “On-the-fly Definition Augmentation of LLMs for Biomedical NER,” NAACL, 2024.
- Amanpreet Singh, Mike D’Arcy, Arman Cohan, Doug Downey, and Sergey Feldman, “SciRepEval: A Multi-Format Benchmark for Scientific Document Representations,” EMNLP, 2023.
- Yuze Lou, Bailey Kuehl, Erin Bransom, Sergey Feldman, Aakanksha Naik, and Doug Downey, “S2abEL: A Dataset for Entity Linking from Scientific Tables,” EMNLP, 2023.
- Aakanksha Naik, Sravanthi Parasa, Sergey Feldman, Lucy Lu Wang, and Tom Hope, “Literature-Augmented Clinical Outcome Prediction,” NAACL Findings, 2022.
- Shaurya Rohatgi, Doug Downey, Daniel King, and Sergey Feldman, “S2AMP: A High-Coverage Dataset of Scholarly Mentorship Inferred from Publications,” JCDL, 2022.
- Shivashankar Subramanian, Daniel King, Doug Downey, and Sergey Feldman, “S2AND: A Benchmark and Evaluation System for Author Name Disambiguation,” JCDL, 2021.
- Asia J. Biega, Fernando Diaz, Michael D. Ekstrand, Sergey Feldman, Sebastian Kohlmeier, “Overview of the TREC 2020 Fair Ranking Track,” TREC 2020.
- Sean MacAvaney, Andrew Yates, Sergey Feldman, Doug Downey, Arman Cohan, and Nazli Goharian, “Simplified Data Wrangling with ir_datasets,” SIGIR, 2021.
- Arman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey, and Daniel S. Weld, “SPECTER: Document-level Representation Learning using Citation-informed Transformers,” ACL, 2020.
- Chandra Bhagavatula, Sergey Feldman, Russell Power, and Waleed Ammar, “Content-Based Citation Recommendation,” NAACL-HLT, 2018.
- Waleed Ammar, Dirk Groeneveld, Chandra Bhagavatula, Iz Beltagy, Miles Crawford, Doug Downey, Jason Dunkelberger, Ahmed Elgohary, Sergey Feldman, Vu Ha, Rodney Kinney, Sebastian Kohlmeier, Kyle Lo, Tyler Murray, Hsu-Han Ooi, Matthew Peters, Joanna Power, Sam Skjonsberg, Lucy Lu Wang, Chris Wilhelm, Zheng Yuan, Madeleine van Zuylen, and Oren Etzioni, “Construction of the Literature Graph in Semantic Scholar,” NAACL-HLT, 2018.
- Sergey Feldman, Maya R. Gupta, and Bela A. Frigyik, “Multi-Task Averaging,” NIPS, 2012.
- Luca Cazzanti, Sergey Feldman, Maya R. Gupta, and Michael Gabbay, “Multi-Task Regularization of Generative Similarity Models,” Lecture Notes in Computer Science, 2011.
- Sergey Feldman, Marius A. Marin, Mari Ostendorf, and Maya R. Gupta, “Part-of-Speech Histogram Features for Genre Classification of Text,” IEEE ICASSP, 2009.
- Marius A. Marin, Sergey Feldman, Mari Ostendorf, and Maya R. Gupta, “Filtering Web Text to Match Target Genres,” IEEE ICASSP, 2009.
- Sergey Feldman, Marius A. Marin, Julie Medero, and Mari Ostendorf, “Classifying Factored Genres with Part-of-Speech Histograms,” NAACL-HLT, 2009.
Theses & Technical Reports
- Jeremy Li, Alex Rubinsteyn, Sergey Feldman, et al., “Scientific computing in the age of agentic AI: an exploratory field report,” OpenAI exploratory field report, 2026.
- Sergey Feldman, Kyle Lo, and Waleed Ammar, “Citation Count Analysis for Papers with Preprints,” 2018.
- Sergey Feldman, “Multi-Task Averaging: Theory and Practice,” University of Washington PhD Thesis, 2012.
- Sergey Feldman, Barbara Frewen, Michael J. MacCoss, and Maya R. Gupta, “Filtering Tandem Mass Spectra for Quality,” University of Washington Dept. of Electrical Engineering Technical Report UWEETR-2012-0001, 2012.