
Talk Title: The Benefits of Sharing and Collaboration in the MOC
Talk Abstract: Issac Newton wrote “If I have seen further it is by standing on the shoulders of Giants” which captures the benefits of our community culture and natural pedagology. We do this all the time with writing, speaking, publishing and sharing knowledge. When it come to computing, however, even when it happens, it is cumbersome, tedious, and inefficient. If we find interesting data sets, we have to copy them, understand the format, do our own cleaning, check for updates, pay for the compute and manage intermediate values, code updates, and so on. Go one place to find data another to find code, excuse parts on many different platforms, mange different accounts and passwords, and it seems to continually get more complicated.
Our vision for the MOC-A is to foster sharing and collaboration, making it efficient, inexpensive, with a fair and balanced burden. If Alice spends D dollars in the MOC-A to produce a high quality dataset S which is then used by N people, the (N-1)*D was saved by sharing. If P programs accessed this dataset at roughly the same time, smart community caching can save the cost of C-1 expensive disk fetches. We can think of many ways of sharing the cost burden and many interesting research challenges. Of particular note are those related to naming, provenance, and reproducility. In this talk, we argue that as the MOC-A grows to support a large number of researchers and groups often using each other tools and data, it can also serve as a testbed for exploring the benefits of sharing and collaborating. Indeed, the rest of the talks in this session are initial steps of this vision.
Semantics maybe an enabler. Data workflows or pipelines are frequently repeated and overlapping. Scheduling and resource allocation of pipelines are more efficient when with knowledge of congestion, when and what might be loading the same data, when serializing/deserializing can be avoided especially when sharing intermediate results. It is easier to find data when the format, schema, values, and governance rules are have clear generally understood semantics and standards.
One first step has already been identified by the research community embrassing the FAIR data princple (Findable, Accessible, Interoperable, and Re-useable). Along with the demand for experiments to be Essentially Reproducable, we can call this the FAIRER principle. Of course, compute is needed in each part of this principle. This talk presents a vision of a semantic cloud, some benefits and challenges and why the MOC-A is the right starting place.
Bio: Larry Rudolph is a Senior Research Scientist and VP at Two Sigma Investments LLP, a Visiting Scholar at NYU, Affiliate off CSAIL MIT, Academic Board Member of the Mass Open Cloud, and CTO of ReDigi. After receiving his PhD from NYU Courant Institute, he has been on the faculty of CMU, Hebrew University and MIT as well as a project manager at VMWare.