Subtopic annotation and automatic segmentation for news texts in Brazilian Portuguese

Registro completo de metadados
MetadadosDescriçãoIdioma
Autor(es): dc.creatorCardoso, Paula C. F.-
Autor(es): dc.creatorPardo, Thiago A. S.-
Autor(es): dc.creatorTaboada, Maite-
Data de aceite: dc.date.accessioned2026-02-09T12:31:49Z-
Data de disponibilização: dc.date.available2026-02-09T12:31:49Z-
Data de envio: dc.date.issued2019-01-25-
Data de envio: dc.date.issued2019-01-25-
Data de envio: dc.date.issued2017-
Fonte completa do material: dc.identifierhttps://repositorio.ufla.br/handle/1/32557-
Fonte completa do material: dc.identifierhttps://www.euppublishing.com/doi/10.3366/cor.2017.0108-
Fonte: dc.identifier.urihttp://educapes.capes.gov.br/handle/capes/1163137-
Descrição: dc.descriptionSubtopic segmentation aims to break documents into subtopical text passages, which develop a main topic in a text. Being capable of automatically detecting subtopics is very useful for several Natural Language Processing applications. For instance, in automatic summarisation, having the subtopics at hand enables the production of summaries with good subtopic coverage. Given the usefulness of subtopic segmentation, it is common to assemble a reference-annotated corpus that supports the study of the envisioned phenomena and the development and evaluation of systems. In this paper, we describe the subtopic annotation process in a corpus of news texts written in Brazilian Portuguese, following a systematic annotation process and answering the main research questions when performing corpus annotation. Based on this corpus, we propose novel methods for subtopic segmentation following patterns of discourse organisation, specifically using Rhetorical Structure Theory. We show that discourse structures mirror the subtopic changes in news texts. An important outcome of this work is the freely available annotated corpus, which, to the best of our knowledge, is the only one for Portuguese. We demonstrate that some discourse knowledge may significantly help to find boundaries automatically in a text. In particular, the relation type and the level of the tree structure are important features.-
Idioma: dc.languageen-
Publicador: dc.publisherEdinburgh University Press-
Direitos: dc.rightsrestrictAccess-
???dc.source???: dc.sourceCorpora-
Palavras-chave: dc.subjectCorpus annotation-
Palavras-chave: dc.subjectNewspaper discourse-
Palavras-chave: dc.subjectSubtopics-
Palavras-chave: dc.subjectText segmentation-
Palavras-chave: dc.subjectDiscurso de jornal-
Palavras-chave: dc.subjectSubtópicos-
Palavras-chave: dc.subjectSegmentação de texto-
Título: dc.titleSubtopic annotation and automatic segmentation for news texts in Brazilian Portuguese-
Tipo de arquivo: dc.typeArtigo-
Aparece nas coleções:Repositório Institucional da Universidade Federal de Lavras (RIUFLA)

Não existem arquivos associados a este item.