Conceptual design is the foundational stage of a design process that translates ill-defined design problems into low-fidelity design concepts and prototypes through design search, creation, and integration. In this stage, product shape design is one of the most paramount aspects. When applying deep learning-based methods to product shape design, two major challenges exist: (1) design data exhibit in multiple modalities and (2) an increasing demand for creativity. With recent advances in deep learning of cross-modal tasks (DLCMTs), which can transfer one design modality to another, we see opportunities to develop artificial intelligence (AI) to assist the design of product shapes in a new paradigm. In this paper, we conduct a systematic review of the retrieval, generation, and manipulation methods for DLCMT that involve three cross-modal types: text-to-3D shape, text-to-sketch, and sketch-to-3D shape. The review identifies 50 articles from a pool of 1341 papers in the fields of computer graphics, computer vision, and engineering design. We review (1) state-of-the-art DLCMT methods that can be applied to product shape design and (2) identify the key challenges, such as lack of consideration of engineering performance in the early design phase that need to be addressed when applying DLCMT methods. In the end, we discuss the potential solutions to these challenges and propose a list of research questions that point to future directions of data-driven conceptual design.
more »
« less
Deep Learning of Cross-Modal Tasks for Conceptual Design of Engineered Products: A Review
Abstract Conceptual design is the foundational stage of a design process, translating ill-defined design problems to low-fidelity design concepts and prototypes. While deep learning approaches are widely applied in later design stages for design automation, we see fewer attempts in conceptual design for three reasons: 1) the data in this stage exhibit multiple modalities: natural language, sketches, and 3D shapes, and these modalities are challenging to represent in deep learning methods; 2) it requires knowledge from a larger source of inspiration instead of focusing on a single design task; and 3) it requires translating designers’ intent and feedback, and hence needs more interaction with designers and/or users. With recent advances in deep learning of cross-modal tasks (DLCMT) and the availability of large cross-modal datasets, we see opportunities to apply these learning methods to the conceptual design of product shapes. In this paper, we review 30 recent journal articles and conference papers across computer graphics, computer vision, and engineering design fields that involve DLCMT of three modalities: natural language, sketches, and 3D shapes. Based on the review, we identify the challenges and opportunities of utilizing DLCMT in 3D shape concepts generation, from which we propose a list of research questions pointing to future research directions.
more »
« less
- Award ID(s):
- 2207408
- PAR ID:
- 10466552
- Publisher / Repository:
- American Society of Mechanical Engineers
- Date Published:
- ISBN:
- 978-0-7918-8626-7
- Format(s):
- Medium: X
- Location:
- St. Louis, Missouri, USA
- Sponsoring Org:
- National Science Foundation
More Like this
-
-
Conceptual diagrams are used extensively to understand abstract relationships, explain complex ideas, and solve difficult problems. To illustrate concepts effectively, experts find appropriate visual representations and translate concepts into concrete shapes. This translation step is not supported explicitly by current diagramming tools. This paper investigates how domain experts create conceptual diagrams via semi-structured interviews with 18 participants from diverse backgrounds. Our participants create, adapt, and reuse visual representations using both sketches and digital tools. However, they had trouble using current diagramming tools to transition from sketches and reuse components from earlier diagrams. Our participants also expressed frustration with the slow feedback cycles and barriers to automation of their tools. Based on these results, we suggest four opportunities of diagramming tools — exploration support, representation salience, live engagement, and vocabulary correspondence — that together enable a natural diagramming experience. Finally, we discuss possibilities to leverage recent research advances to develop natural diagramming tools.more » « less
-
Modern cities generate vast streams of urban dynamics data reflecting mobility demand, environmental conditions, and traffic patterns. The value of these data lies not only in individual modalities but in their integration—urban signals are highly interdependent, with changes in one modality often influencing others. Consequently, predicting any single urban dynamic requires information from multiple interrelated sources. Although numerous methods—ranging from deep learning models to recent LLM-based approaches—have been proposed, most are limited in scope. They either focus on single-modality prediction, rely on rigid model designs that lack flexibility, or overlook inter-modal dependencies. As a result, they struggle to adapt to dynamic urban conditions and suffer from degraded predictive performance across modalities. In this paper, we propose UniLLM, a unified large language model for multi-modal urban dynamics prediction. At its core, UniLLM introduces a Unified Cross-Modal Alignment Module that transforms heterogeneous urban data into latent representations while preserving modality-specific patterns and capturing cross-modal correlations through a contrastive learning objective. To support dynamic adaptation across tasks and modalities, we design a Routing-Aware Prompting Mechanism that learns soft prompts based on task context and modality semantics. Furthermore, a Multi-Modal Memory-Guided Adaptive Algorithm employs replay-based gradient coordination and Frank–Wolfe optimization to mitigate cross-modal catastrophic forgetting during fine-tuning. Extensive experiments across multiple cities and urban modalities demonstrate that UniLLM consistently outperforms state-of-the-art baselines. These results highlight UniLLM's potential as a flexible and robust forecasting model for real-world, multi-modal urban environments.more » « less
-
Authoring realistic haptic textures typically requires low-level parameter tuning and repeated trial-and-error, limiting speed, transparency, and creative reach. We present a language-driven authoring system that turns natural-language prompts into multimodal textures: two coordinated haptic channels—sliding vibrations via force/speed-conditioned autoregressive (AR) models and tapping transients—and a text-prompted visual preview from a diffusion model. A shared, language-aligned latent links modalities so a single prompt yields semantically consistent haptic and visual signals; designers can write goals (e.g., “gritty but cushioned surface,” “smooth and hard metal surface”) and immediately see and feel the result through a 3D haptic device. To verify that the learned latent encodes perceptually meaningful structure, we conduct an anchor-referenced, attribute-wise evaluation for roughness, slipperiness, and hardness. Participant ratings are projected to the interpretable line between two real-material references, revealing consistent trends—asperity effects in roughness, compliance in hardness, and surface-film influence in slipperiness. A human subject study further indicates coherent cross-modal experience and low effort for prompt-based iteration. The results show that language can serve as a practical control modality for texture authoring: prompts reliably steer material semantics across haptic and visual channels, enabling a prompt-first, designer oriented workflow that replaces manual parameter tuning with interpretable, text-guided refinement.more » « less
-
The ability to quickly learn a new task with minimal instruction - known as few-shot learning - is a central aspect of intelligent agents. Classical few-shot benchmarks make use of few-shot samples from a single modality, but such samples may not be sufficient to characterize an entire concept class. In contrast, humans use cross-modal information to learn new concepts efficiently. In this work, we demonstrate that one can indeed build a better visual dog classifier by reading about dogs and listening to them bark. To do so, we exploit the fact that recent multimodal foundation models such as CLIP are inherently cross-modal, mapping different modalities to the same representation space. Specifically, we propose a simple cross-modal adaptation approach that learns from few-shot examples spanning different modalities. By repurposing class names as additional one-shot training samples, we achieve SOTA results with an embarrassingly simple linear classifier for vision-language adaptation. Furthermore, we show that our approach can benefit existing methods such as prefix tuning, adapters, and classifier ensembling. Finally, to explore other modalities beyond vision and language, we construct the first (to our knowledge) audiovisual few-shot benchmark and use cross-modal training to improve the performance of both image and audio classification.more » « less
An official website of the United States government

