New community

Subscribe to the gold package and get unlimited access to Shamra Academy

Language-Aligned Waypoint (LAW) Supervision for Vision-and-Language Navigation in Continuous Environments

الإشراف على الطريق المحاذي في اللغة (القانون) للملاحة للرؤية واللغة في البيئات المستمرة

235 0 0 0.0 ( 0 )

Download Cite

Added by Association for Computation Linguistics مقالة

Publication date 2021

fields Artificial Intelligence

and research's language is English

Created by Shamra Editor

continuous environments language-aligned waypoint navigation in continuous البيئات المستمرة نقطة الطريق المحاذاة الملاحة في المستمر صناعة حمض الفوسفور

visit our facebook page

‎Shamra Academia - شمرا أكاديميا‎

Ask ChatGPT about the research

Abstract in Arabic Abstract in English

In the Vision-and-Language Navigation (VLN) task an embodied agent navigates a 3D environment, following natural language instructions. A challenge in this task is how to handle off the path' scenarios where an agent veers from a reference path. Prior work supervises the agent with actions based on the shortest path from the agent's location to the goal, but such goal-oriented supervision is often not in alignment with the instruction. Furthermore, the evaluation metrics employed by prior work do not measure how much of a language instruction the agent is able to follow. In this work, we propose a simple and effective language-aligned supervision scheme, and a new metric that measures the number of sub-instructions the agent has completed during navigation.

References used

https://aclanthology.org/

rate research

Data Efficient Masked Language Modeling for Vision and Language

652 - Association for Computation Linguistics 2021 مقالة

Masked language modeling (MLM) is one of the key sub-tasks in vision-language pretraining. In the cross-modal setting, tokens in the sentence are masked at random, and the model predicts the masked tokens given the image and the text. In this paper, we observe several key disadvantages of MLM in this setting. First, as captions tend to be short, in a third of the sentences no token is sampled. Second, the majority of masked tokens are stop-words and punctuation, leading to under-utilization of the image. We investigate a range of alternative masking strategies specific to the cross-modal setting that address these shortcomings, aiming for better fusion of text and image in the learned representation. When pre-training the LXMERT model, our alternative masking strategies consistently improve over the original masking strategy on three downstream tasks, especially in low resource settings. Further, our pre-training approach substantially outperforms the baseline model on a prompt-based probing task designed to elicit image objects. These results and our analysis indicate that our method allows for better utilization of the training data.

efficient masked language لغة ملثمفة فعالة صناعة حمض الفوسفور

VisualSem: a high-quality knowledge graph for vision and language

662 - Association for Computation Linguistics 2021 مقالة

An exciting frontier in natural language understanding (NLU) and generation (NLG) calls for (vision-and-) language models that can efficiently access external structured knowledge repositories. However, many existing knowledge bases only cover limite d domains, or suffer from noisy data, and most of all are typically hard to integrate into neural language pipelines. To fill this gap, we release VisualSem: a high-quality knowledge graph (KG) which includes nodes with multilingual glosses, multiple illustrative images, and visually relevant relations. We also release a neural multi-modal retrieval model that can use images or sentences as inputs and retrieves entities in the KG. This multi-modal retrieval model can be integrated into any (neural network) model pipeline. We encourage the research community to use VisualSem for data augmentation and/or as a source of grounding, among other possible uses. VisualSem as well as the multi-modal retrieval models are publicly available and can be downloaded in this URL: https://github.com/iacercalixto/visualsem.

high-quality knowledge graph high-quality knowledge multi-modal retrieval الرسم البياني المعرفة عالية الجودة المعرفة عالية الجودة استرجاع متعددة مشروط صناعة حمض الفوسفور المزيد..

FLIN: A Flexible Natural Language Interface for Web Navigation

393 - Association for Computation Linguistics 2021 مقالة

AI assistants can now carry out tasks for users by directly interacting with website UIs. Current semantic parsing and slot-filling techniques cannot flexibly adapt to many different websites without being constantly re-trained. We propose FLIN, a na tural language interface for web navigation that maps user commands to concept-level actions (rather than low-level UI actions), thus being able to flexibly adapt to different websites and handle their transient nature. We frame this as a ranking problem: given a user command and a webpage, FLIN learns to score the most relevant navigation instruction (involving action and parameter values). To train and evaluate FLIN, we collect a dataset using nine popular websites from three domains. Our results show that FLIN was able to adapt to new websites in a given domain.

flexible natural language natural language interface flexible natural لغة طبيعية مرنة واجهة اللغة الطبيعية طبيعي مرن صناعة حمض الفوسفور المزيد..

Improving Pre-trained Vision-and-Language Embeddings for Phrase Grounding

340 - Association for Computation Linguistics 2021 مقالة

Phrase grounding aims to map textual phrases to their associated image regions, which can be a prerequisite for multimodal reasoning and can benefit tasks requiring identifying objects based on language. With pre-trained vision-and-language models ac hieving impressive performance across tasks, it remains unclear if we can directly utilize their learned embeddings for phrase grounding without fine-tuning. To this end, we propose a method to extract matched phrase-region pairs from pre-trained vision-and-language embeddings and propose four fine-tuning objectives to improve the model phrase grounding ability using image-caption data without any supervised grounding signals. Experiments on two representative datasets demonstrate the effectiveness of our objectives, outperforming baseline models in both weakly-supervised and supervised phrase grounding settings. In addition, we evaluate the aligned embeddings on several other downstream tasks and show that we can achieve better phrase grounding without sacrificing representation generality.

phrase grounding grounding improving pre-trained عبارة التأريض التأريض تحسين مدرب مسبقا صناعة حمض الفوسفور المزيد..

Semantic Aligned Multi-modal Transformer for Vision-LanguageUnderstanding: A Preliminary Study on Visual QA

399 - Association for Computation Linguistics 2021 مقالة

Recent vision-language understanding approaches adopt a multi-modal transformer pre-training and finetuning paradigm. Prior work learns representations of text tokens and visual features with cross-attention mechanisms and captures the alignment sole ly based on indirect signals. In this work, we propose to enhance the alignment mechanism by incorporating image scene graph structures as the bridge between the two modalities, and learning with new contrastive objectives. In our preliminary study on the challenging compositional visual question answering task, we show the proposed approach achieves improved results, demonstrating potentials to enhance vision-language understanding.

semantic aligned multi-modal aligned multi-modal transformer semantic aligned المحاذاة الدلالية متعددة مشروط محاذاة محول متعدد مشروط المحاذاة الدلالية صناعة حمض الفوسفور المزيد..

يمكنك البدء بجني المال وتحقيق ربح مادي من أبحاثك العلمية، المزيد

Language-Aligned Waypoint (LAW) Supervision for Vision-and-Language Navigation in Continuous Environments

الإشراف على الطريق المحاذي في اللغة (القانون) للملاحة للرؤية واللغة في البيئات المستمرة

Ask ChatGPT about the research

Read More

suggested questions