{"id":1042,"date":"2024-01-29T10:49:04","date_gmt":"2024-01-29T09:49:04","guid":{"rendered":"https:\/\/adim.web.amu.edu.pl\/?page_id=1042"},"modified":"2024-11-15T12:24:13","modified_gmt":"2024-11-15T11:24:13","slug":"lnnor-corpus","status":"publish","type":"page","link":"https:\/\/adim.web.amu.edu.pl\/en\/lnnor-corpus\/","title":{"rendered":"LnNor Corpus"},"content":{"rendered":"\n<p><strong>The<\/strong>&nbsp;<strong>LnNor Corpus: A spoken multilingual corpus of non-native and native Norwegian, English and Polish<\/strong> <strong>(Part 1) <strong>&amp; <strong>(Part 2)<\/strong><\/strong><\/strong> <\/p>\n\n\n\n<p><strong>Wrembel Magdalena, Hwaszcz Krzysztof, Pludra Agnieszka, Ska\u0142ba Anna, Weckwerth<br>Jaros\u0142aw, Malarski Kamil, Cal Zuzanna, K\u0119dzierska Hanna, Czarnecki-Verner Tristan,<br>Balas Anna, Ka\u017amierski Kamil, \u017bychli\u0144ski Sylwiusz, Gruszecka Justyna<\/strong><\/p>\n\n\n\n<p>The LnNor corpus was created as part of the data collection in two projects: CLIMAD (Cross-linguistic influence in multilingualism across domains: phonology and syntax) and ADIM (Across-domain Investigations in Multilingualism: Modeling L3 Acquisition in Diverse Settings), led by Prof. Magdalena Wrembel at Adam Mickiewicz University in Pozna\u0144, Poland and by Prof. Marit Westergaard at the Arctic University of Norway, from December 2021 to April 2024 with funding from the National Science Centre (NCN) in Poland and Norway Grants.<\/p>\n\n\n\n<p>The CLIMAD and ADIM projects explored cross-linguistic influence (CLI) in the acquisition, processing, and use of a third language (L3\/Ln) across various language domains and focused on different settings and stages of acquisition from a multilingual perspective. A range of sophisticated methodologies, such as perception and production tests, grammaticality judgement tasks and online brain imaging techniques like EEG, were leveraged to unravel the intricacies of multilingual processing. By capturing real-time insights into the interplay of cross-linguistic influences, the projects not only provided valuable contributions to the understanding of L3\/Ln acquisition but also advanced theoretical frameworks in this field.<\/p>\n\n\n\n<p>Corpus data collection covered a broad range of speech elicitation tasks. The recordings consist of word, sentence and text reading, picture story description, video story retelling, spontaneous speech and socio-phonetic interviews in Polish, English and Norwegian. The corpus contains metadata based on the Language History Questionnaire (Li et al. 2020) such as age, gender, native languages, proficiency level, length of language exposure, age of onset.<\/p>\n\n\n\n<p>Data was collected from different <strong>groups of speakers<\/strong>:<\/p>\n\n\n\n<div class=\"wp-block-group is-layout-constrained wp-block-group-is-layout-constrained\">\n<ul class=\"wp-block-list\">\n<li>L1 Polish learners of Norwegian as L3\/Ln, attending Scandinavian studies at Pozna\u0144 College of Modern Languages and the University of Szczecin (instructed learners);<\/li>\n\n\n\n<li>L1 Polish learners of Norwegian as L3\/Ln, living in Norway (naturalistic learners)<\/li>\n\n\n\n<li>L1 English natives as controls<\/li>\n\n\n\n<li>L1 Norwegian natives as controls<\/li>\n\n\n\n<li>speakers of L2\/L3\/Ln English and L2\/L3\/Ln Norwegian with various L1 backgrounds<\/li>\n<\/ul>\n<\/div>\n\n\n\n<p>Seven types of <strong>speech tasks<\/strong> were recorded in Norwegian, English and Polish:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>word reading<\/li>\n\n\n\n<li>sentence reading<\/li>\n\n\n\n<li>text reading (\u201cThe North Wind and the Sun\u201d)<\/li>\n\n\n\n<li>picture description<\/li>\n\n\n\n<li>story telling<\/li>\n\n\n\n<li>video description<\/li>\n\n\n\n<li>translation from Polish\/English into Norwegian<\/li>\n<\/ul>\n\n\n\n<p><strong>Metadata<\/strong> corresponding to the recordings include the following information:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>speaker ID, age, gender, education, current residence, speaker status<br>(instructed\/naturalistic\/native), native language, additional languages spoken<\/li>\n\n\n\n<li>recording ID<\/li>\n\n\n\n<li>language: PL (Polish), EN (English), NO (Norwegian)<\/li>\n\n\n\n<li>status: L1, L2, L3\/Ln<\/li>\n\n\n\n<li>speech task: WR (word reading), SR (sentence reading), TR (text reading), PD (picture description), ST (story telling), VD (video description), translation from Polish (TP) \/ English (TE) into Norwegian<\/li>\n\n\n\n<li>recording date, recording place, iteration, recording environment, recording device, type of microphone, noise level, etc.<\/li>\n<\/ul>\n\n\n\n<p><strong>The labels of the recordings <\/strong>adhere to a structured format: <strong>PROJECT_SPEAKER<br>ID_LANGUAGE STATUS_TASK<\/strong>, wherein:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>PROJECT corresponds to the project within which the data were collected (A for<br>ADIM, C for CLIMAD)<\/li>\n\n\n\n<li>SPEAKER ID corresponds to a unique speaker ID consisting of 8 characters<\/li>\n\n\n\n<li>LANGUAGE STATUS represents the language in which the task was recorded and its<\/li>\n\n\n\n<li>status for the speaker (e.g., L1PL, L2EN, L3NO)<\/li>\n\n\n\n<li>TASK corresponds to the type of speech task recorded (e.g., TR, SR, WR, etc.). If a given task type was done more than once, numbers corresponding to their iterations have been added after TASK.<\/li>\n<\/ul>\n\n\n\n<p>The LnNor corpus has been created to represent multilingual speech with a focus on L3\/Ln<br>Norwegian learners as well as native controls of Norwegian, English and Polish. The corpus is<br>designed to study linguistic variation in learners acquiring Norwegian as a foreign language in instructed and naturalistic settings. Additionally, a subcorpus of native speech patterns is provided to serve as a benchmark, against which the learners&#8217; productions could be compared. Furthermore, parts of the corpus contain word alignment with orthographic transcriptions of speech to facilitate subsequent analyses across various linguistic domains.<\/p>\n\n\n\n<p>All speech samples were recorded with the use of Shure SM-35 unidirectional cardioid<br>head-worn condenser microphones, using portable Marantz PMD620 solid state recorders with signal digitized at 48 kHz, 16-bit. This set-up was selected to minimize ambient noise and provide clear and focused recordings.<\/p>\n\n\n\n<p>The <strong>LnNOR corpus part 1<\/strong> consists of 1073 annotated files from 78 speakers. The speakers included 53 L1 Polish, 16 L1 Norwegian and 9 L1 speakers of other European languages. The total recording time is approximately 35 hours and the full size is 19 GB. The recordings in the released LnNor corpus part 1 cover data collected between 2021-2022.<\/p>\n\n\n\n<p>The <strong>LnNOR corpus part 2<\/strong> consists of 1739 annotated files from 153 speakers. The speakers included 113 L1 Polish, 33 L1 Norwegian and 18 L1 speakers of English. The total recording time is approximately 67 hours and the full size is 31 GB. The recordings in the released LnNor corpus part 2 cover data collected between 2023-2024.<\/p>\n\n\n\n<p>The corpus has been published under an open license in three repositories:<\/p>\n\n\n\n<p>&#8211; AMUReD repository <a href=\"https:\/\/researchportal.amu.edu.pl\/info\/researchdata\/UAM31a0b24bafa34f3688092d72047817c3?r=researchdata&amp;ps=20&amp;tab=&amp;title=Szczeg%25C3%25B3%25C5%2582y%2Brekordu%2B%25E2%2580%2593%2BDane%2Bbadawcze%2B%25E2%2580%2593%2BUniwersytet%2Bim.%2BAdama%2BMickiewicza%2Bw%2BPoznaniu&amp;lang=pl\">Part1<\/a> &amp; <a href=\"https:\/\/researchportal.amu.edu.pl\/info\/researchdata\/UAM22e4ff296ea8497fbe64a214a78c3220\/\">Part 2<\/a><br>&#8211; CLARIN D-Space repository <a href=\"https:\/\/clarin-pl.eu\/dspace\/handle\/11321\/931\">Part1<\/a> &amp; <a href=\"https:\/\/clarin-pl.eu\/dspace\/handle\/11321\/932\">Part2<\/a><br>&#8211; WA server <a href=\"http:\/\/corpora.wa.amu.edu.pl\/LnNor_Corpus_part_1\/\">Part1<\/a> &amp; <a href=\"http:\/\/corpora.wa.amu.edu.pl\/LnNor_Corpus_part_2\/\">Part 2<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>The&nbsp;LnNor Corpus: A spoken multilingual corpus of non-native and native Norwegian, English and Polish (Part 1) &amp; (Part 2) Wrembel [&hellip;]<\/p>\n","protected":false},"author":3,"featured_media":0,"parent":0,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"_monsterinsights_skip_tracking":false,"footnotes":""},"class_list":["post-1042","page","type-page","status-publish","hentry"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v23.8 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>LnNor Corpus - ADIM<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/adim.web.amu.edu.pl\/en\/lnnor-corpus\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"LnNor Corpus - ADIM\" \/>\n<meta property=\"og:description\" content=\"The&nbsp;LnNor Corpus: A spoken multilingual corpus of non-native and native Norwegian, English and Polish (Part 1) &amp; (Part 2) Wrembel [&hellip;]\" \/>\n<meta property=\"og:url\" content=\"https:\/\/adim.web.amu.edu.pl\/en\/lnnor-corpus\/\" \/>\n<meta property=\"og:site_name\" content=\"ADIM\" \/>\n<meta property=\"article:publisher\" content=\"https:\/\/www.facebook.com\/people\/The-ADIM-Project-Insights-on-Multilingual-Minds\/100089017856711\/\" \/>\n<meta property=\"article:modified_time\" content=\"2024-11-15T11:24:13+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/adim.web.amu.edu.pl\/wp-content\/uploads\/2023\/01\/thumbnail_IMG_1406-1024x614-1.jpg\" \/>\n\t<meta property=\"og:image:width\" content=\"1024\" \/>\n\t<meta property=\"og:image:height\" content=\"614\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/jpeg\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:site\" content=\"@ADIMLinguists\" \/>\n<meta name=\"twitter:label1\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data1\" content=\"4 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"WebPage\",\"@id\":\"https:\/\/adim.web.amu.edu.pl\/en\/lnnor-corpus\/\",\"url\":\"https:\/\/adim.web.amu.edu.pl\/en\/lnnor-corpus\/\",\"name\":\"LnNor Corpus - ADIM\",\"isPartOf\":{\"@id\":\"https:\/\/adim.web.amu.edu.pl\/en\/#website\"},\"datePublished\":\"2024-01-29T09:49:04+00:00\",\"dateModified\":\"2024-11-15T11:24:13+00:00\",\"breadcrumb\":{\"@id\":\"https:\/\/adim.web.amu.edu.pl\/en\/lnnor-corpus\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/adim.web.amu.edu.pl\/en\/lnnor-corpus\/\"]}]},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/adim.web.amu.edu.pl\/en\/lnnor-corpus\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/adim.web.amu.edu.pl\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"LnNor Corpus\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/adim.web.amu.edu.pl\/en\/#website\",\"url\":\"https:\/\/adim.web.amu.edu.pl\/en\/\",\"name\":\"ADIM (Across-domain Investigations in Multilingualism)\",\"description\":\"Find out how our team is working on understanding language acquisition in multilinguals!\",\"publisher\":{\"@id\":\"https:\/\/adim.web.amu.edu.pl\/en\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/adim.web.amu.edu.pl\/en\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/adim.web.amu.edu.pl\/en\/#organization\",\"name\":\"ADIM (Across-domain Investigations in Multilingualism)\",\"url\":\"https:\/\/adim.web.amu.edu.pl\/en\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/adim.web.amu.edu.pl\/en\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/adim.web.amu.edu.pl\/wp-content\/uploads\/2022\/11\/logo-with-name.png\",\"contentUrl\":\"https:\/\/adim.web.amu.edu.pl\/wp-content\/uploads\/2022\/11\/logo-with-name.png\",\"width\":458,\"height\":476,\"caption\":\"ADIM (Across-domain Investigations in Multilingualism)\"},\"image\":{\"@id\":\"https:\/\/adim.web.amu.edu.pl\/en\/#\/schema\/logo\/image\/\"},\"sameAs\":[\"https:\/\/www.facebook.com\/people\/The-ADIM-Project-Insights-on-Multilingual-Minds\/100089017856711\/\",\"https:\/\/x.com\/ADIMLinguists\"]}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"LnNor Corpus - ADIM","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/adim.web.amu.edu.pl\/en\/lnnor-corpus\/","og_locale":"en_US","og_type":"article","og_title":"LnNor Corpus - ADIM","og_description":"The&nbsp;LnNor Corpus: A spoken multilingual corpus of non-native and native Norwegian, English and Polish (Part 1) &amp; (Part 2) Wrembel [&hellip;]","og_url":"https:\/\/adim.web.amu.edu.pl\/en\/lnnor-corpus\/","og_site_name":"ADIM","article_publisher":"https:\/\/www.facebook.com\/people\/The-ADIM-Project-Insights-on-Multilingual-Minds\/100089017856711\/","article_modified_time":"2024-11-15T11:24:13+00:00","og_image":[{"width":1024,"height":614,"url":"https:\/\/adim.web.amu.edu.pl\/wp-content\/uploads\/2023\/01\/thumbnail_IMG_1406-1024x614-1.jpg","type":"image\/jpeg"}],"twitter_card":"summary_large_image","twitter_site":"@ADIMLinguists","twitter_misc":{"Est. reading time":"4 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"WebPage","@id":"https:\/\/adim.web.amu.edu.pl\/en\/lnnor-corpus\/","url":"https:\/\/adim.web.amu.edu.pl\/en\/lnnor-corpus\/","name":"LnNor Corpus - ADIM","isPartOf":{"@id":"https:\/\/adim.web.amu.edu.pl\/en\/#website"},"datePublished":"2024-01-29T09:49:04+00:00","dateModified":"2024-11-15T11:24:13+00:00","breadcrumb":{"@id":"https:\/\/adim.web.amu.edu.pl\/en\/lnnor-corpus\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/adim.web.amu.edu.pl\/en\/lnnor-corpus\/"]}]},{"@type":"BreadcrumbList","@id":"https:\/\/adim.web.amu.edu.pl\/en\/lnnor-corpus\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/adim.web.amu.edu.pl\/"},{"@type":"ListItem","position":2,"name":"LnNor Corpus"}]},{"@type":"WebSite","@id":"https:\/\/adim.web.amu.edu.pl\/en\/#website","url":"https:\/\/adim.web.amu.edu.pl\/en\/","name":"ADIM (Across-domain Investigations in Multilingualism)","description":"Find out how our team is working on understanding language acquisition in multilinguals!","publisher":{"@id":"https:\/\/adim.web.amu.edu.pl\/en\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/adim.web.amu.edu.pl\/en\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/adim.web.amu.edu.pl\/en\/#organization","name":"ADIM (Across-domain Investigations in Multilingualism)","url":"https:\/\/adim.web.amu.edu.pl\/en\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/adim.web.amu.edu.pl\/en\/#\/schema\/logo\/image\/","url":"https:\/\/adim.web.amu.edu.pl\/wp-content\/uploads\/2022\/11\/logo-with-name.png","contentUrl":"https:\/\/adim.web.amu.edu.pl\/wp-content\/uploads\/2022\/11\/logo-with-name.png","width":458,"height":476,"caption":"ADIM (Across-domain Investigations in Multilingualism)"},"image":{"@id":"https:\/\/adim.web.amu.edu.pl\/en\/#\/schema\/logo\/image\/"},"sameAs":["https:\/\/www.facebook.com\/people\/The-ADIM-Project-Insights-on-Multilingual-Minds\/100089017856711\/","https:\/\/x.com\/ADIMLinguists"]}]}},"_links":{"self":[{"href":"https:\/\/adim.web.amu.edu.pl\/en\/wp-json\/wp\/v2\/pages\/1042"}],"collection":[{"href":"https:\/\/adim.web.amu.edu.pl\/en\/wp-json\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/adim.web.amu.edu.pl\/en\/wp-json\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/adim.web.amu.edu.pl\/en\/wp-json\/wp\/v2\/users\/3"}],"replies":[{"embeddable":true,"href":"https:\/\/adim.web.amu.edu.pl\/en\/wp-json\/wp\/v2\/comments?post=1042"}],"version-history":[{"count":22,"href":"https:\/\/adim.web.amu.edu.pl\/en\/wp-json\/wp\/v2\/pages\/1042\/revisions"}],"predecessor-version":[{"id":1238,"href":"https:\/\/adim.web.amu.edu.pl\/en\/wp-json\/wp\/v2\/pages\/1042\/revisions\/1238"}],"wp:attachment":[{"href":"https:\/\/adim.web.amu.edu.pl\/en\/wp-json\/wp\/v2\/media?parent=1042"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}