NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

Towards a NoOps Model for WLCG

https://doi.org/https://doi.org/10.1051/epjconf/202024507024

Gardner, Robert; Bryant, Lincoln; Stephen, Judith; Vukotic, Ilija; Weaver, Christopher; Wu, Wenjing (November 2020, 24th International Conference on Computing in High Energy and Nuclear Physics (CHEP 2019))
null (Ed.)
One of the most costly factors in providing a global computing infrastructure such as the WLCG is the human effort in deployment, integration, and operation of the distributed services supporting collaborative computing, data sharing and delivery, and analysis of extreme scale datasets. Furthermore, the time required to roll out global software updates, introduce new service components, or prototype novel systems requiring coordinated deployments across multiple facilities is often increased by communication latencies, staff availability, and in many cases expertise required for operations of bespoke services. While the WLCG (and distributed systems implemented throughout HEP) is a global service platform, it lacks the capability and flexibility of a modern platform-as-a-service including continuous integration/continuous delivery (CI/CD) methods, development-operations capabilities (DevOps, where developers assume a more direct role in the actual production infrastructure), and automation. Most importantly, tooling which reduces required training, bespoke service expertise, and the operational effort throughout the infrastructure, most notably at the resource endpoints (sites), is entirely absent in the current model. In this paper, we explore ideas and questions around potential NoOps models in this context: what is realistic given organizational policies and constraints? How should operational responsibility be organized across teams and facilities? What are the technical gaps? What are the social and cybersecurity challenges? Conversely what advantages does a NoOps model deliver for innovation and for accelerating the pace of delivery of new services needed for the HL-LHC era? We will describe initial work along these lines in the context of providing a data delivery network supporting IRIS-HEP DOMA R&D.
more » « less
Full Text Available
Managing Privilege and Access on Federated Edge Platforms

https://doi.org/10.1145/3332186.3332234

Breen, Joe; Bryant, Lincoln; Chen, Jiahui; Ford, Emerson; Gardner, Robert W.; Glupker, Gage; Griffith, Skyler; Kulbertis, Ben; McKee, Shawn; Pierce, Rose; et al (January 2019, Proceedings of the Practice and Experience in Advanced Research Computing on Rise of the Machines (learning))

Full Text Available
Developing Edge Services for Federated Infrastructure Using MiniSLATE

https://doi.org/10.1145/3332186.3332236

Breen, Joe; Bryant, Lincoln; Chen, Jiahui; Ford, Emerson; Gardner, Robert W.; Glupker, Gage; Griffith, Skyler; Kulbertis, Ben; McKee, Shawn; Pierce, Rose; et al (January 2019, Proceedings of the Practice and Experience in Advanced Research Computing on Rise of the Machines (Learning))

Full Text Available
Building the SLATE Platform

https://doi.org/10.1145/3219104.3219144

Breen, Joe; McKee, Shawn; Riedel, Benedikt; Stidd, Jason; Truong, Luan; Vukotic, Ilija; Bryant, Lincoln; Carcassi, Gabriele; Chen, Jiahui; Gardner, Robert W.; et al (July 2018, Proceedings of the Practice and Experience on Advanced Research Computing)

We describe progress on building the SLATE (Services Layer at the Edge) platform. The high level goal of SLATE is to facilitate creation of multi-institutional science computing systems by augmenting the canonical Science DMZ pattern with a generic, "programmable", secure and trusted underlayment platform. This platform permits hosting of advanced container-centric services needed for higher-level capabilities such as data transfer nodes, software and data caches, workflow services and science gateway components. SLATE uses best-of-breed data center virtualization and containerization components, and where available, software defined networking, to enable distributed automation of deployment and service lifecycle management tasks by domain experts. As such it will simplify creation of scalable platforms that connect research teams, institutions and resources to accelerate science while reducing operational costs and development cycle times.
more » « less
Full Text Available
SLATE and the Mobility of Capability

Gardner, R; Breen, J; Bryant, L; McKee, S (October 2017, Science Gateways 2017)

SLATE (Services Layer at the Edge) is a new project that, when complete, will implement “cyberinfrastructure as code” by augmenting the canonical Science DMZ pattern with a generic, programmable, secure and trusted underlayment platform. This platform will host advanced container-centric services needed for higher-level capabilities such as data transfer nodes, software and data caches, workflow services and science gateway components. SLATE will use best-of-breed data center virtualization components, and where available, software defined networking, to enable distributed automation of deployment and service lifecycle management tasks by domain experts. As such it will simplify creation of scalable platforms that connect research teams, institutions and resources to accelerate science while reducing operational costs and development cycle times. Since SLATE will be designed to require only commodity components for its functional layers, its potential for building distributed systems should extend across all data center types and scales, thus enabling creation of ubiquitous, science-driven cyberinfrastructure. By providing automation and programmatic interfaces to distributed HPC backends and other cyberinfrastructure resources, SLATE will amplify the reach of science gateways and therefore the domain communities they support.
more » « less
Full Text Available

Search for: All records