New community

Subscribe to the gold package and get unlimited access to Shamra Academy

Natural Language Video Localization with Learnable Moment Proposals

توطين الفيديو باللغة الطبيعية مع مقترحات لحظة معرفة

254 0 0 0.0 ( 0 )

Download Cite

Added by Association for Computation Linguistics مقالة

Publication date 2021

fields Artificial Intelligence

and research's language is English

Created by Shamra Editor

visit our facebook page

‎Shamra Academia - شمرا أكاديميا‎

Ask ChatGPT about the research

Abstract in Arabic Abstract in English

نظرا لفيديو غير جذوع واستعلام لغة طبيعية، يهدف توطين فيديو اللغة الطبيعي (NLVL) إلى تحديد لحظة الفيديو الموصوفة بواسطة الاستعلام. لمعالجة هذه المهمة، يمكن تجميع الأساليب الحالية تقريبا إلى مجموعتين: 1) نماذج اقتراح ورتبة تحدد أولا مجموعة من المرشحين لحظة مصممة باليد، ثم اكتشفوا أفضل واحد مطابقة. 2) النماذج الخالية من الاقتراح تنبئ مباشرة اثنين من الحدود الزمنية لحظة المرجعية من الإطارات. حاليا، تقريبا جميع طرق الاقتراح والرتبة لها أداء أدنى أقل من نظرائها الخالية من الاقتراح. في هذه الورقة، نجادل بأن أداء نماذج الاقتراح والرسوم يتم تقليله بسبب الإصابة المحددة مسبقا: 1) من الصعب ضمان القواعد المصممة باليد التغطية الكاملة للقطاعات المستهدفة. 2) لحظات مرشح العينات كثيفة تسبب حسابا زائدة عن الحاجة ويخفض أداء عملية الترتيب. تحقيقا لهذه الغاية، نقترح نموذجا جديدا نموذج LPNET (شبكة اقتراح مقترح ل NLVL) مع مجموعة ثابتة من مقترحات اللحظات المحددة. يتم تعديل موضع وطول هذه المقترحات ديناميكيا أثناء عملية التدريب. علاوة على ذلك، تم اقتراح خسارة على علم الحدود لاستفادة من المعلومات على مستوى الإطار وأيضا تحسين الأداء. أظهرت الاعتداءات الواسعة على اثنين من معايير NLVL التحدي فعالية LPNET على الطرق الحالية من الأساليب الحالية.

Given an untrimmed video and a natural language query, Natural Language Video Localization (NLVL) aims to identify the video moment described by query. To address this task, existing methods can be roughly grouped into two groups: 1) propose-and-rank models first define a set of hand-designed moment candidates and then find out the best-matching one. 2) proposal-free models directly predict two temporal boundaries of the referential moment from frames. Currently, almost all the propose-and-rank methods have inferior performance than proposal-free counterparts. In this paper, we argue that the performance of propose-and-rank models are underestimated due to the predefined manners: 1) Hand-designed rules are hard to guarantee the complete coverage of targeted segments. 2) Densely sampled candidate moments cause redundant computation and degrade the performance of ranking process. To this end, we propose a novel model termed LPNet (Learnable Proposal Network for NLVL) with a fixed set of learnable moment proposals. The position and length of these proposals are dynamically adjusted during training process. Moreover, a boundary-aware loss has been proposed to leverage frame-level information and further improve performance. Extensive ablations on two challenging NLVL benchmarks have demonstrated the effectiveness of LPNet over existing state-of-the-art methods.

References used

https://aclanthology.org/

rate research

NeuralLog: Natural Language Inference with Joint Neural and Logical Reasoning

506 - Association for Computation Linguistics 2021 مقالة

Deep learning (DL) based language models achieve high performance on various benchmarks for Natural Language Inference (NLI). And at this time, symbolic approaches to NLI are receiving less attention. Both approaches (symbolic and DL) have their adva ntages and weaknesses. However, currently, no method combines them in a system to solve the task of NLI. To merge symbolic and deep learning methods, we propose an inference framework called NeuralLog, which utilizes both a monotonicity-based logical inference engine and a neural network language model for phrase alignment. Our framework models the NLI task as a classic search problem and uses the beam search algorithm to search for optimal inference paths. Experiments show that our joint logic and neural inference system improves accuracy on the NLI task and can achieve state-of-art accuracy on the SICK and MED datasets.

natural language inference logical reasoning الاستدلال باللغة الطبيعية التفكير المنطقي صناعة حمض الفوسفور

Disentangling Generative Factors in Natural Language with Discrete Variational Autoencoders

258 - Association for Computation Linguistics 2021 مقالة

The ability of learning disentangled representations represents a major step for interpretable NLP systems as it allows latent linguistic features to be controlled. Most approaches to disentanglement rely on continuous variables, both for images and text. We argue that despite being suitable for image datasets, continuous variables may not be ideal to model features of textual data, due to the fact that most generative factors in text are discrete. We propose a Variational Autoencoder based method which models language features as discrete variables and encourages independence between variables for learning disentangled representations. The proposed model outperforms continuous and discrete baselines on several qualitative and quantitative benchmarks for disentanglement as well as on a text style transfer downstream application.

disentangling generative factors factors in natural تحرير العوامل المولدة صناعة حمض الفوسفور

Improving Policing with Natural Language Processing

599 - Association for Computation Linguistics 2021 مقالة

This article explores the potential for Natural Language Processing (NLP) to enable a more effective, prevention focused and less confrontational policing model that has hitherto been too resource consuming to implement at scale. Problem-Oriented Pol icing (POP) is a potential replacement, at least in part, for traditional policing which adopts a reactive approach, relying heavily on the criminal justice system. By contrast, POP seeks to prevent crime by manipulating the underlying conditions that allow crimes to be committed. Identifying these underlying conditions requires a detailed understanding of crime events - tacit knowledge that is often held by police officers but which can be challenging to derive from structured police data. One potential source of insight exists in unstructured free text data commonly collected by police for the purposes of investigation or administration. Yet police agencies do not typically have the skills or resources to analyse these data at scale. In this article we argue that NLP offers the potential to unlock these unstructured data and by doing so allow police to implement more POP initiatives. However we caution that using NLP models without adequate knowledge may either allow or perpetuate bias within the data potentially leading to unfavourable outcomes.

قانون جنائي صناعة حمض الفوسفور

Video Question Answering with Phrases via Semantic Roles

388 - Association for Computation Linguistics 2021 مقالة

Video Question Answering (VidQA) evaluation metrics have been limited to a single-word answer or selecting a phrase from a fixed set of phrases. These metrics limit the VidQA models' application scenario. In this work, we leverage semantic roles deri ved from video descriptions to mask out certain phrases, to introduce VidQAP which poses VidQA as a fill-in-the-phrase task. To enable evaluation of answer phrases, we compute the relative improvement of the predicted answer compared to an empty string. To reduce the influence of language bias in VidQA datasets, we retrieve a video having a different answer for the same question. To facilitate research, we construct ActivityNet-SRL-QA and Charades-SRL-QA and benchmark them by extending three vision-language models. We perform extensive analysis and ablative studies to guide future work. Code and data are public.

video question answering video question إجابة سؤال الفيديو سؤال الفيديو صناعة حمض الفوسفور

Building a Video-and-Language Dataset with Human Actions for Multimodal Logical Inference

368 - Association for Computation Linguistics 2021 مقالة

This paper introduces a new video-and-language dataset with human actions for multimodal logical inference, which focuses on intentional and aspectual expressions that describe dynamic human actions. The dataset consists of 200 videos, 5,554 action l abels, and 1,942 action triplets of the form (subject, predicate, object) that can be easily translated into logical semantic representations. The dataset is expected to be useful for evaluating multimodal inference systems between videos and semantically complicated sentences including negation and quantification.

multimodal logical inference human actions dynamic human actions الاستدلال المنطقي متعدد الوسائط الإجراءات البشرية الإجراءات البشرية الديناميكية صناعة حمض الفوسفور المزيد..

يمكنك البدء بجني المال وتحقيق ربح مادي من أبحاثك العلمية، المزيد

Natural Language Video Localization with Learnable Moment Proposals

توطين الفيديو باللغة الطبيعية مع مقترحات لحظة معرفة

Ask ChatGPT about the research

Read More

suggested questions