- Home
- Search Results
- Page 1 of 1
Search for: All records
-
Total Resources2
- Resource Type
-
0000100001000000
- More
- Availability
-
11
- Author / Contributor
- Filter by Author / Creator
-
-
Heinrich, G (2)
-
Abramowicz, H (1)
-
Adli, E (1)
-
Alharthi, F (1)
-
Almanza-Soto, M (1)
-
Altakach, M M (1)
-
Altmannshofer, W (1)
-
Ampudia Castelazo, S (1)
-
Angal-Kalinin, D (1)
-
Anguiano, J (1)
-
Appleby, R B (1)
-
Apsimon, O (1)
-
Arbey, A (1)
-
Arco, F (1)
-
Arquero, O (1)
-
Aryshev, A (1)
-
Asai, S (1)
-
Attié, D (1)
-
Avila-Jimenez, J L (1)
-
Baer, H (1)
-
- Filter by Editor
-
-
Baer, Howard (1)
-
Barklow, Timothy (1)
-
Behnke, Ties (1)
-
Belomestnykh, Sergey (1)
-
Berger, Martin (1)
-
Braathen, Johannes (1)
-
Demarteau, Marcel (1)
-
Durieux, Gauthier (1)
-
Faus-Golfe, Angeles (1)
-
Foster, Brian (1)
-
Gessner, Spencer (1)
-
Gori, Stefania (1)
-
Heinemeyer, Sven (1)
-
Irles, Adrian (1)
-
Ishino, Masaya (1)
-
Jeans, Daniel (1)
-
Kaabi, Walid (1)
-
Kanemura, Shinya (1)
-
Kilian, Wolfgang (1)
-
Koppenburg, Patrick (1)
-
-
Have feedback or suggestions for a way to improve these results?
!
Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher.
Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?
Some links on this page may take you to non-federal websites. Their policies may differ from this site.
-
Training modern LLMs is extremely resource intensive, and customizing them for various deployment scenarios characterized by limited compute and memory resources through repeated training is impractical. In this paper, we introduce Flextron, a network architecture and post-training model optimization framework supporting flexible model deployment. The Flextron architecture utilizes a nested elastic structure to rapidly adapt to specific user-defined latency and accuracy targets during inference with no additional fine-tuning required. It is also input-adaptive, and can automatically route tokens through its sub-networks for improved performance and efficiency. We present a sample-efficient training method and associated routing algorithms for systematically transforming an existing trained LLM into a Flextron model. We evaluate Flextron on the GPT-3 and LLama-2 family of LLMs, and demonstrate superior performance over multiple end-to-end trained variants and other state-of-the-art elastic networks, all with a single pretraining run that consumes a mere 7.63% tokens compared to original pretraining.more » « less
-
Abramowicz, H; Adli, E; Alharthi, F; Almanza-Soto, M; Altakach, M M; Altmannshofer, W; Ampudia Castelazo, S; Angal-Kalinin, D; Anguiano, J; Appleby, R B; et al (, The European Physical Journal Special Topics)Baer, Howard; Barklow, Timothy; Behnke, Ties; Belomestnykh, Sergey; Berger, Martin; de_Blas, Jorge; Braathen, Johannes; Durieux, Gauthier; Demarteau, Marcel; Faus-Golfe, Angeles (Ed.)Abstract In this paper we review the physics opportunities at linear$$\mathrm{e}^{+}\mathrm{e}^{-} $$ colliders with a special focus on high centre-of-mass energies and beam polarisation, take a fresh look at the various accelerator technologies available or under development and, for the first time, discuss how a facility first equipped with a technology that is mature today could be upgraded with technologies of tomorrow to reach much higher energies and/or luminosities. In addition, we discuss detectors, alternative collider modes, as well as opportunities for beyond-collider experiments and R&D facilities as part of a linear collider facility (LCF). The material of this paper supports all plans for$$\mathrm{e}^{+}\mathrm{e}^{-} $$ linear colliders and the additional opportunities they offer, independently of technology choice or proposed site, as well as R&D for advanced accelerator technologies. This joint perspective on the physics goals, early technologies and upgrade strategies has been developed by the LCVision team based on an initial discussion at LCWS2024 in Tokyo and a follow-up at the LCVision Community Event at CERN in January 2025. It heavily builds on decades of achievements of the global linear collider community, in particular in the context of CLIC and ILC.more » « lessFree, publicly-accessible full text available March 2, 2027
An official website of the United States government

Full Text Available