Harvesting information from captions for weakly supervised semantic segmentation

Johann Sawatzky; Debayan Banerjee; Juergen Gall

doi:10.1109/ICCVW.2019.00549

Harvesting information from captions for weakly supervised semantic segmentation

Publikation: Beiträge in Sammelwerken › Aufsätze in Konferenzbänden › Forschung › begutachtet

Authors

Johann Sawatzky
Debayan Banerjee
Juergen Gall

Since acquiring pixel-wise annotations for training convolutional neural networks for semantic image segmentation is time-consuming, weakly supervised approaches that only require class tags have been proposed. In this work, we propose another form of supervision, namely image captions as they can be found on the Internet. These captions have two advantages. They do not require additional curation as it is the case for the clean class tags used by current weakly supervised approaches and they provide textual context for the classes present in an image. To leverage such textual context, we deploy a multi-modal network that learns a joint embedding of the visual representation of the image and the textual representation of the caption. The network estimates text activation maps (TAMs) for class names as well as compound concepts, i.e. combinations of nouns and their attributes. The TAMs of compound concepts describing classes of interest substantially improve the quality of the estimated class activation maps which are then used to train a network for semantic segmentation. We evaluate our method on the COCO dataset where it achieves state of the art results for weakly supervised image segmentation.

Originalsprache	Englisch
Titel	2019 International Conference on Computer Vision Workshops : ICCV 2019 : proceedings : 27 October-2 November 2019, Seoul, Korea
Anzahl der Seiten	10
Erscheinungsort	Piscataway
Verlag	Institute of Electrical and Electronics Engineers Inc.
Erscheinungsdatum	10.2019
Seiten	4481-4490
Aufsatznummer	9022140
ISBN (Print)	978-1-7281-5024-6
ISBN (elektronisch)	978-1-7281-5023-9
DOIs	https://doi.org/10.1109/ICCVW.2019.00549
Publikationsstatus	Erschienen - 10.2019
Extern publiziert	Ja
Veranstaltung	17th IEEE/CVF International Conference on Computer Vision Workshop - ICCVW 2019 - Seoul, Südkorea Dauer: 27.10.2019 → 28.10.2019 Konferenznummer: 17 https://iccv2019.thecvf.com/

Bibliographische Notiz

Publisher Copyright:
© 2019 IEEE.

Fachgebiete

Informatik

Weitere Publikationen dieser Person(en)

ShortPathQA: A Dataset for Controllable Fusion of Large Language Models with Knowledge Graphs

Salnikov, M., Sakhovskiy, A., Nikishina, I., Usmanova, A., Kraft, A., Möller, C., Banerjee, D., Huang, J., Jiang, L., Abdullah, R., Yan, X., Tutubalina, E., Usbeck, R. & Panchenko, A., 2026, Natural Language Processing and Information Systems: 30th International Conference on Applications of Natural Language to Information Systems, NLDB 2025, Proceedings. Ichise, R. (Hrsg.). Springer Science and Business Media Deutschland, S. 95-110 16 S. (Lecture Notes in Computer Science; Band 15836 LNCS).

Publikation: Beiträge in Sammelwerken › Aufsätze in Konferenzbänden › Forschung › begutachtet

Automating SPARQL Query Translations between DBpedia and Wikidata

Bartels, M. C., Banerjee, D. & Usbeck, R., 14.07.2025, Linking Meaning: Semantic Technologies Shaping the Future of AI: Cover 74617 Proceedings of the 21st International Conference on Semantic Systems, 3-5 September 2025, Vienna, Austria. Spahiu, B., Vahdati, S., Salatino, A., Pellegrini, T. & Havur, G. (Hrsg.). IOS Press BV, S. 176-193 18 S. (Studies on the Semantic Web; Band 62).

Publikation: Beiträge in Sammelwerken › Aufsätze in Konferenzbänden › Forschung

Best Practices in AI and Data Science Models Evaluation

Banerjee, D., Taffa, T. A. & Usbeck, R., 2025, 55. Jahrestagung der Gesellschaft für Informatik, INFORMATIK 2025: The Wide Open - Offenheit von Source bis Science, Potsdam, Germany, September 16-19, 2025. Lucke, U., Stieglitz, S., Uebernickel, F., Lamprecht, A.-L. & Klein, M. (Hrsg.). Gesellschaft für Informatik, Bonn, Band P-366. S. 1211-1219 9 S. (LNI).

Publikation: Beiträge in Sammelwerken › Aufsätze in Konferenzbänden › Forschung › begutachtet

DBLPLink 2.0 -- An Entity Linker for the DBLP Scholarly Knowledge Graph

Banerjee, D., Taffa, T. A. & Usbeck, R., 30.07.2025

Publikation: Andere wissenschaftliche Beiträge › Andere › Forschung

HySQA: Hybrid Scholarly Question Answering

Taffa, T., Banerjee, D., Assabie, Y. & Usbeck, R., 26.08.2025, Linking Meaning: Semantic Technologies Shaping the Future of AI: Proceedings of the 21st International Conference on Semantic Systems, 3-5 September 2025, Vienna, Austria. Spahiu, B., Vahdati, S., Salatino, A., Pellegrini, T. & Havur, G. (Hrsg.). Amsterdam: IOS Press BV, S. 247-263 17 S. (Studies on the Semantic Web; Band 62).

Publikation: Beiträge in Sammelwerken › Aufsätze in Konferenzbänden › Forschung › begutachtet

DOI

https://doi.org/10.1109/ICCVW.2019.00549
Endgültige, publizierte Fassung

Harvesting information from captions for weakly supervised semantic segmentation

Authors

Bibliographische Notiz

Fachgebiete

Weitere Publikationen dieser Person(en)

ShortPathQA: A Dataset for Controllable Fusion of Large Language Models with Knowledge Graphs

Automating SPARQL Query Translations between DBpedia and Wikidata

Best Practices in AI and Data Science Models Evaluation

DBLPLink 2.0 -- An Entity Linker for the DBLP Scholarly Knowledge Graph

HySQA: Hybrid Scholarly Question Answering

DOI

Zuletzt angesehen

Aktivitäten

Publikationen