NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

Dolfin: Diffusion Layout Transformers without Autoencoder

Wang, Yilin; Chen, Zeyuan; Zhong, Liangjun; Ding, Zheng; Sha, Zhizhou; Tu, Zhuowen (September 2024, ECCV)

Full Text Available
TokenCompose: Grounding Diffusion with Token-level Supervision

Wang, Zirui; Sha, Zhizhou; Ding, Zheng; Wang, Yilin; Tu, Zhuowen (June 2024, Proceedings)

Full Text Available
Patched Denoising Diffusion Models For High-Resolution Image Synthesis

Ding, Zheng; Zhang, Mengqi; Wu, Jiajun Wu; Tu, Zhuowen (May 2024, ICLR)

Full Text Available
Bayesian Diffusion Models for 3D Shape Reconstruction

Xu, Haiyang; Lei, Yu; Chen, Zeyuan; Zhang, Xiang; Zhao, Yue; Wang, Yilin; Tu, Zhuowen (June 2024, IEEE)

Full Text Available
MasQCLIP for Open-Vocabulary Universal Image Segmentation

Xu, Xin; Xiong, Tianyi; Ding, Zheng; Tu, Zhuowen (October 2023, IEEE)

Full Text Available
Uni-3D: A Universal Model for Panoptic 3D Scene Reconstruction

Zhang, Xiang; Chen, Zeyuan; Wei, Fangyin; Tu, Zhuowen (October 2023, IEEE)

Full Text Available
Instance Segmentation with Mask-supervised Polygonal Regression Transformers

Lazarow, Justin; Xu, Weijian; Tu, Zhuowen (June 2022, IEEE Computer Society Conference on Computer Vision and Pattern Recognition)

In this paper, we present an end-to-end instance segmentation method that regresses a polygonal boundary for each object instance. This sparse, vectorized boundary representation for objects, while attractive in many downstream computer vision tasks, quickly runs into issues of parity that need to be addressed: parity in supervision and parity in performance when compared to existing pixel-based methods. This is due in part to object instances being annotated with ground-truth in the form of polygonal boundaries or segmentation masks, yet being evaluated in a convenient manner using only segmentation masks. Our method, BoundaryFormer, is a Transformer based architecture that directly predicts polygons yet uses instance mask segmentations as the ground-truth supervision for computing the loss. We achieve this by developing an end-to-end differentiable model that solely relies on supervision within the mask space through differentiable rasterization. BoundaryFormer matches or surpasses the Mask R-CNN method in terms of instance segmentation quality on both COCO and Cityscapes while exhibiting significantly better transferability across datasets.
more » « less
Full Text Available
Instance Segmentation with Mask-supervised Polygonal Regression Transformers

Lazarow, Justin; Xu, Weijian; Tu, Zhuowen (June 2022, IEEE Conference on Computer Vision and Pattern Recognition)

Full Text Available
Co-Scale Conv-Attentional Image Transformers

https://doi.org/10.1109/ICCV48922.2021.00983

Xu, Weijian; Xu, Yifan; Chang, Tyler; Tu, Zhuowen (October 2021, International Conference on Computer Vision)

Full Text Available
Convolutions and Self-Attention: Re-interpreting Relative Positions in Pre-trained Language Models

https://doi.org/10.18653/v1/2021.acl-long.333

Chang, Tyler A.; Xu, Yifan; Xu, Weijian; Tu, Zhuowen (August 2021, Proceedings of the conference Association for Computational Linguistics Meeting)

Full Text Available

« Prev Next »

Search for: All records