Harvesting information from captions for weakly supervised semantic segmentation

Publikation: Beiträge in SammelwerkenAufsätze in KonferenzbändenForschungbegutachtet

Standard

Harvesting information from captions for weakly supervised semantic segmentation. / Sawatzky, Johann; Banerjee, Debayan; Gall, Juergen.
2019 International Conference on Computer Vision Workshops: ICCV 2019 : proceedings : 27 October-2 November 2019, Seoul, Korea. Piscataway: Institute of Electrical and Electronics Engineers Inc., 2019. S. 4481-4490 9022140 (IEEE International Conference on Computer Vision workshops; Band 2019).

Publikation: Beiträge in SammelwerkenAufsätze in KonferenzbändenForschungbegutachtet

Harvard

Sawatzky, J, Banerjee, D & Gall, J 2019, Harvesting information from captions for weakly supervised semantic segmentation. in 2019 International Conference on Computer Vision Workshops: ICCV 2019 : proceedings : 27 October-2 November 2019, Seoul, Korea., 9022140, IEEE International Conference on Computer Vision workshops, Bd. 2019, Institute of Electrical and Electronics Engineers Inc., Piscataway, S. 4481-4490, 17th IEEE/CVF International Conference on Computer Vision Workshop - ICCVW 2019, Seoul, Südkorea, 27.10.19. https://doi.org/10.1109/ICCVW.2019.00549

APA

Sawatzky, J., Banerjee, D., & Gall, J. (2019). Harvesting information from captions for weakly supervised semantic segmentation. In 2019 International Conference on Computer Vision Workshops: ICCV 2019 : proceedings : 27 October-2 November 2019, Seoul, Korea (S. 4481-4490). Artikel 9022140 (IEEE International Conference on Computer Vision workshops; Band 2019). Institute of Electrical and Electronics Engineers Inc.. https://doi.org/10.1109/ICCVW.2019.00549

Vancouver

Sawatzky J, Banerjee D, Gall J. Harvesting information from captions for weakly supervised semantic segmentation. in 2019 International Conference on Computer Vision Workshops: ICCV 2019 : proceedings : 27 October-2 November 2019, Seoul, Korea. Piscataway: Institute of Electrical and Electronics Engineers Inc. 2019. S. 4481-4490. 9022140. (IEEE International Conference on Computer Vision workshops). doi: 10.1109/ICCVW.2019.00549

Bibtex

@inbook{13c2379a3a944f5bacd91e0409b3aeca,
title = "Harvesting information from captions for weakly supervised semantic segmentation",
abstract = "Since acquiring pixel-wise annotations for training convolutional neural networks for semantic image segmentation is time-consuming, weakly supervised approaches that only require class tags have been proposed. In this work, we propose another form of supervision, namely image captions as they can be found on the Internet. These captions have two advantages. They do not require additional curation as it is the case for the clean class tags used by current weakly supervised approaches and they provide textual context for the classes present in an image. To leverage such textual context, we deploy a multi-modal network that learns a joint embedding of the visual representation of the image and the textual representation of the caption. The network estimates text activation maps (TAMs) for class names as well as compound concepts, i.e. combinations of nouns and their attributes. The TAMs of compound concepts describing classes of interest substantially improve the quality of the estimated class activation maps which are then used to train a network for semantic segmentation. We evaluate our method on the COCO dataset where it achieves state of the art results for weakly supervised image segmentation.",
keywords = "Multimodal learning, Semantic segmentation, Weakly supervised learning, Weakly supervised semantic segmentation, Informatics",
author = "Johann Sawatzky and Debayan Banerjee and Juergen Gall",
note = "Publisher Copyright: {\textcopyright} 2019 IEEE.; 17th IEEE/CVF International Conference on Computer Vision Workshop - ICCVW 2019, ICCVW 2019 ; Conference date: 27-10-2019 Through 28-10-2019",
year = "2019",
month = oct,
doi = "10.1109/ICCVW.2019.00549",
language = "English",
isbn = "978-1-7281-5024-6",
series = "IEEE International Conference on Computer Vision workshops",
publisher = "Institute of Electrical and Electronics Engineers Inc.",
pages = "4481--4490",
booktitle = "2019 International Conference on Computer Vision Workshops",
address = "United States",
url = "https://iccv2019.thecvf.com/",

}

RIS

TY - CHAP

T1 - Harvesting information from captions for weakly supervised semantic segmentation

AU - Sawatzky, Johann

AU - Banerjee, Debayan

AU - Gall, Juergen

N1 - Conference code: 17

PY - 2019/10

Y1 - 2019/10

N2 - Since acquiring pixel-wise annotations for training convolutional neural networks for semantic image segmentation is time-consuming, weakly supervised approaches that only require class tags have been proposed. In this work, we propose another form of supervision, namely image captions as they can be found on the Internet. These captions have two advantages. They do not require additional curation as it is the case for the clean class tags used by current weakly supervised approaches and they provide textual context for the classes present in an image. To leverage such textual context, we deploy a multi-modal network that learns a joint embedding of the visual representation of the image and the textual representation of the caption. The network estimates text activation maps (TAMs) for class names as well as compound concepts, i.e. combinations of nouns and their attributes. The TAMs of compound concepts describing classes of interest substantially improve the quality of the estimated class activation maps which are then used to train a network for semantic segmentation. We evaluate our method on the COCO dataset where it achieves state of the art results for weakly supervised image segmentation.

AB - Since acquiring pixel-wise annotations for training convolutional neural networks for semantic image segmentation is time-consuming, weakly supervised approaches that only require class tags have been proposed. In this work, we propose another form of supervision, namely image captions as they can be found on the Internet. These captions have two advantages. They do not require additional curation as it is the case for the clean class tags used by current weakly supervised approaches and they provide textual context for the classes present in an image. To leverage such textual context, we deploy a multi-modal network that learns a joint embedding of the visual representation of the image and the textual representation of the caption. The network estimates text activation maps (TAMs) for class names as well as compound concepts, i.e. combinations of nouns and their attributes. The TAMs of compound concepts describing classes of interest substantially improve the quality of the estimated class activation maps which are then used to train a network for semantic segmentation. We evaluate our method on the COCO dataset where it achieves state of the art results for weakly supervised image segmentation.

KW - Multimodal learning

KW - Semantic segmentation

KW - Weakly supervised learning

KW - Weakly supervised semantic segmentation

KW - Informatics

UR - http://www.scopus.com/inward/record.url?scp=85082499279&partnerID=8YFLogxK

U2 - 10.1109/ICCVW.2019.00549

DO - 10.1109/ICCVW.2019.00549

M3 - Article in conference proceedings

AN - SCOPUS:85082499279

SN - 978-1-7281-5024-6

T3 - IEEE International Conference on Computer Vision workshops

SP - 4481

EP - 4490

BT - 2019 International Conference on Computer Vision Workshops

PB - Institute of Electrical and Electronics Engineers Inc.

CY - Piscataway

T2 - 17th IEEE/CVF International Conference on Computer Vision Workshop - ICCVW 2019

Y2 - 27 October 2019 through 28 October 2019

ER -

DOI

Zuletzt angesehen

Publikationen

  1. Multilingual disambiguation of named entities using linked data
  2. An antisaturating adaptive preaction and a slide surface to achieve soft landing control for electromagnetic actuators
  3. Adaptive control of the nonlinear dynamic behavior of the cantilever-sample system of an atomic force microscope
  4. One tool to rule? – A field experimental longitudinal study on the costs and benefits of mobile device usage in public agencies
  5. Performance of methods to select landscape metrics for modelling species richness
  6. "to expose, to show, to demonstrate, to inform, to offer. Artistic Practices around 1990"
  7. Initial evidence for a systematic link between core values and emotional experiences in environmental situations
  8. Introduction to the basics of life cycle sustainability assessment focusing on the UNEP/SETAC Life Cycle Initiative LCSA framework
  9. Multibody simulations of distributed flight arrays for Industry 4.0 applications
  10. Introduction
  11. Towards a Heuristic for Scheduling Offshore Installation Processes
  12. Score-Informed Analysis of Tuning, Intonation, Pitch Modulation, and Dynamics in Jazz Solos
  13. Assessment of occupational exertion and strain in laboratory- and real occupational environments
  14. Online Network Impedance Identification with Wave-Package and Inter-Harmonic Signals
  15. Managing the grazing landscape
  16. War isn't hell, it's entertainment
  17. Considerations on establishing prevention reporting at the national level in Germany
  18. Der Medienmanager - Unternehmer im Unternehmen
  19. The impact of digital transformation on the retailing value chain
  20. Credit constraints and margins of import
  21. Tailoring of residual stresses by specific use of defined prestress during laser shock peening
  22. Visualizers versus verbalizers
  23. Erich und die Übersetzer
  24. Multi-use of Community Energy Storage
  25. Chemistry of POPs in the Atmosphere
  26. Wir sind ihr
  27. 'Climate neutral' is a lie - abandon it as a goal
  28. Investigation On The Influence Of Remanufacturing On Production Planning And Control – A Systematic Literature Review
  29. Mythos als Aufklärung