New community

Subscribe to the gold package and get unlimited access to Shamra Academy

Learning to Ground Visual Objects for Visual Dialog

تعلم الكائنات المرئية الأرضية للحوار المرئي

296 0 0 0.0 ( 0 )

Download Cite

Added by Association for Computation Linguistics مقالة

Publication date 2021

fields Artificial Intelligence

and research's language is English

Created by Shamra Editor

visit our facebook page

‎Shamra Academia - شمرا أكاديميا‎

Ask ChatGPT about the research

Abstract in Arabic Abstract in English

الحوار المرئي صعبا لأنه يحتاج إلى الإجابة على سلسلة من الأسئلة المتماسكة بناء على فهم البيئة المرئية. كيفية الأرض الكائنات المرئية ذات الصلة هي واحدة من المشاكل الرئيسية. تستخدم الدراسات السابقة السؤال والتاريخ للحضور في الصورة وتحقيق أداء مرضي، في حين أن هذه الطرق ليست كافية لتحديد الكائنات المرئية ذات الصلة دون أي إرشادات. يحظر التأريض غير المناسب للكائنات المرئية أداء نماذج الحوار المرئي. في هذه الورقة، نقترح نهجا جديدا لتعلم الكائنات المرئية البرية للحوار المرئي، والذي يستخدم آلية تأريض كائنات مرئية جديدة حيث يتم استخدام كل من التوزيعات السابقة والخلفية على الكائنات المرئية لتسهيل التأريض البصرية. على وجه التحديد، يتم استنتاج التوزيع الخلفي على الكائنات المرئية من كل من السياق (التاريخ والأسئلة) والأجوبة، وتضمن التأريض المناسب للأشياء المرئية أثناء عملية التدريب. في هذه الأثناء، يتم استخدام توزيع مسبق، الذي يستنتج من السياق فقط، لتقريب التوزيع الخلفي بحيث يمكن أن تكون الكائنات المرئية المناسبة هي التأريض حتى بدون إجابات أثناء عملية الاستدلال. النتائج التجريبية على مجموعة بيانات V0.9 و V1.0 Visdial تثبت أن نهجنا يحسن النماذج القوية السابقة في كل من الإعدادات الإدارية والتمييزية من خلال هامش هامش.

Visual dialog is challenging since it needs to answer a series of coherent questions based on understanding the visual environment. How to ground related visual objects is one of the key problems. Previous studies utilize the question and history to attend to the image and achieve satisfactory performance, while these methods are not sufficient to locate related visual objects without any guidance. The inappropriate grounding of visual objects prohibits the performance of visual dialog models. In this paper, we propose a novel approach to Learn to Ground visual objects for visual dialog, which employs a novel visual objects grounding mechanism where both prior and posterior distributions over visual objects are used to facilitate visual objects grounding. Specifically, a posterior distribution over visual objects is inferred from both context (history and questions) and answers, and it ensures the appropriate grounding of visual objects during the training process. Meanwhile, a prior distribution, which is inferred from context only, is used to approximate the posterior distribution so that appropriate visual objects can be grounding even without answers during the inference process. Experimental results on the VisDial v0.9 and v1.0 datasets demonstrate that our approach improves the previous strong models in both generative and discriminative settings by a significant margin.

References used

https://aclanthology.org/

rate research

Enhancing Visual Dialog Questioner with Entity-based Strategy Learning and Augmented Guesser

250 - Association for Computation Linguistics 2021 مقالة

Considering the importance of building a good Visual Dialog (VD) Questioner, many researchers study the topic under a Q-Bot-A-Bot image-guessing game setting, where the Questioner needs to raise a series of questions to collect information of an undi sclosed image. Despite progress has been made in Supervised Learning (SL) and Reinforcement Learning (RL), issues still exist. Firstly, previous methods do not provide explicit and effective guidance for Questioner to generate visually related and informative questions. Secondly, the effect of RL is hampered by an incompetent component, i.e., the Guesser, who makes image predictions based on the generated dialogs and assigns rewards accordingly. To enhance VD Questioner: 1) we propose a Related entity enhanced Questioner (ReeQ) that generates questions under the guidance of related entities and learns entity-based questioning strategy from human dialogs; 2) we propose an Augmented Guesser that is strong and is optimized for VD especially. Experimental results on the VisDial v1.0 dataset show that our approach achieves state-of-the-art performance on both image-guessing task and question diversity. Human study further verifies that our model generates more visually related, informative and coherent questions.

enhancing visual dialog visual dialog questioner تعزيز مربع الحوار البصري سؤال الحوار المرئي صناعة حمض الفوسفور

Region under Discussion for visual dialog

229 - Association for Computation Linguistics 2021 مقالة

Visual Dialog is assumed to require the dialog history to generate correct responses during a dialog. However, it is not clear from previous work how dialog history is needed for visual dialog. In this paper we define what it means for a visual quest ion to require dialog history and we release a subset of the Guesswhat?! questions for which their dialog history completely changes their responses. We propose a novel interpretable representation that visually grounds dialog history: the Region under Discussion. It constrains the image's spatial features according to a semantic representation of the history inspired by the information structure notion of Question under Discussion.We evaluate the architecture on task-specific multimodal models and the visual transformer model LXMERT.

dialog history تاريخ الحوار صناعة حمض الفوسفور

Probing Contextual Language Models for Common Ground with Visual Representations

608 - Association for Computation Linguistics 2021 مقالة

The success of large-scale contextual language models has attracted great interest in probing what is encoded in their representations. In this work, we consider a new question: to what extent contextual representations of concrete nouns are aligned with corresponding visual representations? We design a probing model that evaluates how effective are text-only representations in distinguishing between matching and non-matching visual representations. Our findings show that language representations alone provide a strong signal for retrieving image patches from the correct object categories. Moreover, they are effective in retrieving specific instances of image patches; textual context plays an important role in this process. Visually grounded language models slightly outperform text-only language models in instance retrieval, but greatly under-perform humans. We hope our analyses inspire future research in understanding and improving the visual capabilities of language models.

يمنع الانجراف الدلالي contextual language models صناعة حمض الفوسفور

Reasoning Visual Dialog with Sparse Graph Learning and Knowledge Transfer

405 - Association for Computation Linguistics 2021 مقالة

Visual dialog is a task of answering a sequence of questions grounded in an image using the previous dialog history as context. In this paper, we study how to address two fundamental challenges for this task: (1) reasoning over underlying semantic st ructures among dialog rounds and (2) identifying several appropriate answers to the given question. To address these challenges, we propose a Sparse Graph Learning (SGL) method to formulate visual dialog as a graph structure learning task. SGL infers inherently sparse dialog structures by incorporating binary and score edges and leveraging a new structural loss function. Next, we introduce a Knowledge Transfer (KT) method that extracts the answer predictions from the teacher model and uses them as pseudo labels. We propose KT to remedy the shortcomings of single ground-truth labels, which severely limit the ability of a model to obtain multiple reasonable answers. As a result, our proposed model significantly improves reasoning capability compared to baseline methods and outperforms the state-of-the-art approaches on the VisDial v1.0 dataset. The source code is available at https://github.com/gicheonkang/SGLKT-VisDial.

sparse graph learning graph learning الرسم البياني المتفرق يتعلم الرسم البياني تعلم صناعة حمض الفوسفور

Learning to Select Question-Relevant Relations for Visual Question Answering

403 - Association for Computation Linguistics 2021 مقالة

Previous existing visual question answering (VQA) systems commonly use graph neural networks(GNNs) to extract visual relationships such as semantic relations or spatial relations. However, studies that use GNNs typically ignore the importance of each relation and simply concatenate outputs from multiple relation encoders. In this paper, we propose a novel layer architecture that fuses multiple visual relations through an attention mechanism to address this issue. Specifically, we develop a model that uses question embedding and joint embedding of the encoders to obtain dynamic attention weights with regard to the type of questions. Using the learnable attention weights, the proposed model can efficiently use the necessary visual relation features for a given question. Experimental results on the VQA 2.0 dataset demonstrate that the proposed model outperforms existing graph attention network-based architectures. Additionally, we visualize the attention weight and show that the proposed model assigns a higher weight to relations that are more relevant to the question.

select question-relevant relations حدد العلاقات ذات الصلة بالسؤال صناعة حمض الفوسفور

يمكنك البدء بجني المال وتحقيق ربح مادي من أبحاثك العلمية، المزيد

Learning to Ground Visual Objects for Visual Dialog

تعلم الكائنات المرئية الأرضية للحوار المرئي

Ask ChatGPT about the research

Read More

suggested questions