Advanced search powered by artificial intelligence

New community

Subscribe to the gold package and get unlimited access to Shamra Academy

Integrating Visuospatial, Linguistic, and Commonsense Structure into Story Visualization

دمج هيكل visuospatial واللغوية والعموم في تصور القصة

346 0 0 0.0 ( 0 )

Download Cite

Added by Association for Computation Linguistics مقالة

Publication date 2021

fields Artificial Intelligence

and research's language is English

Created by Shamra Editor

integrating visuospatial story visualization visual story دمج visuospatial تصور القصة القصة البصرية صناعة حمض الفوسفور

visit our facebook page

‎Shamra Academia - شمرا أكاديميا‎

Ask ChatGPT about the research

Abstract in Arabic Abstract in English

في حين أن الكثير من الأبحاث قد تم في توليف الرسائل النصية إلى صورة، فقد تم إجراء القليل من العمل لاستكشاف استخدام الهيكل اللغوي لنص المدخلات. هذه المعلومات أكثر أهمية بالنسبة لتصور القصة لأن مدخلاتها لها هيكل سرد صريح يحتاج إلى ترجمة إلى تسلسل الصورة (أو قصة مرئية). أظهر العمل المسبق في هذا المجال أن هناك مجالا واسعا للتحسين في تسلسل الصور الناتج من حيث الجودة البصرية والاتساق والأهمية. في هذه الورقة، نستكشف أولا استخدام أجهزة تحليل الدائرة باستخدام بنية متكررة قائمة على المحولات لترميز المدخلات المهيكلة. ثانيا، نشجع المدخلات المنظمة مع معلومات المنطقية ودراسة تأثير هذه المعرفة الخارجية على جيل القصة البصرية. ثالثا، نحن أيضا دمج البنية المرئية عبر المربعات المحيطة والتسمية الكثيفة لتوفير ملاحظات حول الأحرف / الكائنات في الصور التي تم إنشاؤها داخل إعداد تعليمي مزدوج. نظهر أن نماذج التسمية الكثيفة غير الرفية التي تم تدريبها على جينوم المرئي يمكن أن تحسن الهيكل المكاني للصور من مجال مستهدف مختلف دون الحاجة إلى ضبط جيد. نحن ندرب طراز النموذج باستخدام فقدان داخل القصة داخل القصة (بين الكلمات والمناطق الفرعية للصور) وإظهار تحسينات كبيرة في الجودة البصرية. أخيرا، نحن نقدم تحليلا للمعلومات اللغوية والمكانية.

While much research has been done in text-to-image synthesis, little work has been done to explore the usage of linguistic structure of the input text. Such information is even more important for story visualization since its inputs have an explicit narrative structure that needs to be translated into an image sequence (or visual story). Prior work in this domain has shown that there is ample room for improvement in the generated image sequence in terms of visual quality, consistency and relevance. In this paper, we first explore the use of constituency parse trees using a Transformer-based recurrent architecture for encoding structured input. Second, we augment the structured input with commonsense information and study the impact of this external knowledge on the generation of visual story. Third, we also incorporate visual structure via bounding boxes and dense captioning to provide feedback about the characters/objects in generated images within a dual learning setup. We show that off-the-shelf dense-captioning models trained on Visual Genome can improve the spatial structure of images from a different target domain without needing fine-tuning. We train the model end-to-end using intra-story contrastive loss (between words and image sub-regions) and show significant improvements in visual quality. Finally, we provide an analysis of the linguistic and visuo-spatial information.

References used

https://aclanthology.org/

rate research

Integrating Higher-Level Semantics into Robust Biomedical Name Representations

331 - Association for Computation Linguistics 2021 مقالة

Neural encoders of biomedical names are typically considered robust if representations can be effectively exploited for various downstream NLP tasks. To achieve this, encoders need to model domain-specific biomedical semantics while rivaling the univ ersal applicability of pretrained self-supervised representations. Previous work on robust representations has focused on learning low-level distinctions between names of fine-grained biomedical concepts. These fine-grained concepts can also be clustered together to reflect higher-level, more general semantic distinctions, such as grouping the names nettle sting and tick-borne fever together under the description puncture wound of skin. It has not yet been empirically confirmed that training biomedical name encoders on fine-grained distinctions automatically leads to bottom-up encoding of such higher-level semantics. In this paper, we show that this bottom-up effect exists, but that it is still relatively limited. As a solution, we propose a scalable multi-task training regime for biomedical name encoders which can also learn robust representations using only higher-level semantic classes. These representations can generalise both bottom-up as well as top-down among various semantic hierarchies. Moreover, we show how they can be used out-of-the-box for improved unsupervised detection of hypernyms, while retaining robust performance on various semantic relatedness benchmarks.

الكشف عن التقارير الذاتية representations downstream nlp tasks التوكيلات مهام الدفيئة NLP صناعة حمض الفوسفور

Integrating Lexical Information into Entity Neighbourhood Representations for Relation Prediction

386 - Association for Computation Linguistics 2021 مقالة

Relation prediction informed from a combination of text corpora and curated knowledge bases, combining knowledge graph completion with relation extraction, is a relatively little studied task. A system that can perform this task has the ability to ex tend an arbitrary set of relational database tables with information extracted from a document corpus. OpenKi[1] addresses this task through extraction of named entities and predicates via OpenIE tools then learning relation embeddings from the resulting entity-relation graph for relation prediction, outperforming previous approaches. We present an extension of OpenKi that incorporates embeddings of text-based representations of the entities and the relations. We demonstrate that this results in a substantial performance increase over a system without this information.

entity neighbourhood representations integrating lexical information entity neighbourhood تمثيل حي الكيان دمج المعلومات المعجمية حي الكيان صناعة حمض الفوسفور المزيد..

Automatic Story Generation: Challenges and Attempts

368 - Association for Computation Linguistics 2021 مقالة

Automated storytelling has long captured the attention of researchers for the ubiquity of narratives in everyday life. The best human-crafted stories exhibit coherent plot, strong characters, and adherence to genres, attributes that current states-of -the-art still struggle to produce, even using transformer architectures. In this paper, we analyze works in story generation that utilize machine learning approaches to (1) address story generation controllability, (2) incorporate commonsense knowledge, (3) infer reasonable character actions, and (4) generate creative language.

challenges and attempts automatic story generation story generation التحديات والمحاولات توليد القصة التلقائي جيل القصة صناعة حمض الفوسفور المزيد..

Integrating Unsupervised Data Generation into Self-Supervised Neural Machine Translation for Low-Resource Languages

406 - Association for Computation Linguistics 2021 مقالة

For most language combinations and parallel data is either scarce or simply unavailable. To address this and unsupervised machine translation (UMT) exploits large amounts of monolingual data by using synthetic data generation techniques such as back- translation and noising and while self-supervised NMT (SSNMT) identifies parallel sentences in smaller comparable data and trains on them. To this date and the inclusion of UMT data generation techniques in SSNMT has not been investigated. We show that including UMT techniques into SSNMT significantly outperforms SSNMT (up to +4.3 BLEU and af2en) as well as statistical (+50.8 BLEU) and hybrid UMT (+51.5 BLEU) baselines on related and distantly-related and unrelated language pairs.

دراسة حالة الهند self-supervised neural machine integrating unsupervised data الجهاز العصبي الخاضع للإشراف دمج البيانات غير المدمرة صناعة حمض الفوسفور

Time in the Quran story Psychological time: Case Study

1889 - Tishreen University 2017 ورقة بحثية

This research deals with time psychology in stories of the Holy Quran. It begins with defining time psychology and shows its kinds and their reasons, from a personal and internal time and the time of ego. Then the research moves on to reveal the ti me psychology in Arabic Literature and it begins with poetry. It displays some lines of verse that illustrate the sense of time, and then it clarifies the meaning of time psychology in the new narrative studies. After that, the research deals with time psychology in the story of the Holy Quran and the wrong estimate of time due to not realizing it because of the loss of loss of life, or unconsciousness. Then it exposes some human moments of different characters from the Holy Quran, like incidents of drowning, giving birth, moments of fear and worry, moments of departure and meeting and the moments of determining and triumph. The research demonstrates that the story in the Holy Quran can transmit human feelings, explain emotions and express psychological senses and the inner depths of the characters, using words with high level of transparency and subtlety. The words are meaningul and carefully chosen. So, all of that what has probably given the miraculous story of the Holy Quran it's superiority in both it's linguistics and semantics. Also, it exceeds high above the limited level of the human capacity.

الزمن Time القصة القرآنية الزمن النفسي The story of the Holy Quran Time psychology

يمكنك البدء بجني المال وتحقيق ربح مادي من أبحاثك العلمية، المزيد

Integrating Visuospatial, Linguistic, and Commonsense Structure into Story Visualization

دمج هيكل visuospatial واللغوية والعموم في تصور القصة

Ask ChatGPT about the research

Read More

suggested questions