
Talk Title: Bringing the data close to the compute at Harvard Dataverse
Talk Abstract: Dataverse is a generalist research data repository open source software designed to support the FAIR principles (Findability, Accessibility, Interoperability, and Reuse of digital assets) and to be compliant with virtually any metadata format. With more than 100 installation around the world, Harvard Dataverse represents the most important one with about half million datasets covering many fields of science. With data becoming increasingly large and machine learning and AI equally popular also outside the hard sciences, we decided to exploit the versatility and cost effectiveness of the MOC infrastructure in terms of high performance computing (NERC) and storage (NESE). Moreover, being MOC an integrated infrastructure, it is possible to run computing on large data without moving data over the Internet.
In this talk we present a proof of concept of this approach were the Dataverse software is integrated with NESE and NERC realizing the goal of computing directly on large data.
Bio: Stefano M. Iacus is the Director of Data Science and Product Research at the Institute for Quantitative Social Science, Harvard University. He is also the Managing Director of the Dataverse Project and a member of the executive committee of the OpenDP Project. The Data Science Services and the Data Acquistion & Archiving teams at IQSS also refer to him. Iacus is also an affiliate faculty of the Kempner Institute for the Study of Natural and Artificial Intelligence at Harvard. Before landing at IQSS, he was full professor of statistics at the University of Milan, Italy. He has also served as officer at the Joint Research Centre of the European Commission (2019-2022). Member of the R Core Team for the development of the R statistical environment from 1999 till 2014 and is now a member of the R Foundation for Statistical Computing. He founded two startup companies in the fields of social media analysis and quantitative finance.