Attention:The NSF Public Access Repository (PAR) system and access will be unavailable from 11:00 PM ET on Thursday, August 13 until 12:00 AM ET on Friday, August 14 due to maintenance. We apologize for the inconvenience.


Search for: All records

Creators/Authors contains: "Heinrich, G"

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

  1. Training modern LLMs is extremely resource intensive, and customizing them for various deployment scenarios characterized by limited compute and memory resources through repeated training is impractical. In this paper, we introduce Flextron, a network architecture and post-training model optimization framework supporting flexible model deployment. The Flextron architecture utilizes a nested elastic structure to rapidly adapt to specific user-defined latency and accuracy targets during inference with no additional fine-tuning required. It is also input-adaptive, and can automatically route tokens through its sub-networks for improved performance and efficiency. We present a sample-efficient training method and associated routing algorithms for systematically transforming an existing trained LLM into a Flextron model. We evaluate Flextron on the GPT-3 and LLama-2 family of LLMs, and demonstrate superior performance over multiple end-to-end trained variants and other state-of-the-art elastic networks, all with a single pretraining run that consumes a mere 7.63% tokens compared to original pretraining. 
    more » « less
  2. Baer, Howard; Barklow, Timothy; Behnke, Ties; Belomestnykh, Sergey; Berger, Martin; de_Blas, Jorge; Braathen, Johannes; Durieux, Gauthier; Demarteau, Marcel; Faus-Golfe, Angeles (Ed.)
    Abstract In this paper we review the physics opportunities at linear$$\mathrm{e}^{+}\mathrm{e}^{-} $$ e + e colliders with a special focus on high centre-of-mass energies and beam polarisation, take a fresh look at the various accelerator technologies available or under development and, for the first time, discuss how a facility first equipped with a technology that is mature today could be upgraded with technologies of tomorrow to reach much higher energies and/or luminosities. In addition, we discuss detectors, alternative collider modes, as well as opportunities for beyond-collider experiments and R&D facilities as part of a linear collider facility (LCF). The material of this paper supports all plans for$$\mathrm{e}^{+}\mathrm{e}^{-} $$ e + e linear colliders and the additional opportunities they offer, independently of technology choice or proposed site, as well as R&D for advanced accelerator technologies. This joint perspective on the physics goals, early technologies and upgrade strategies has been developed by the LCVision team based on an initial discussion at LCWS2024 in Tokyo and a follow-up at the LCVision Community Event at CERN in January 2025. It heavily builds on decades of achievements of the global linear collider community, in particular in the context of CLIC and ILC. 
    more » « less
    Free, publicly-accessible full text available March 2, 2027