Ayuda
Ir al contenido

Dialnet


English Dataset For Automatic Forum Extraction

  • Autores: Jakub Sido, Miloslav Konopík, Ondřej Pražák
  • Localización: Computación y Sistemas (CyS), ISSN 1405-5546, ISSN-e 2007-9737, Vol. 23, Nº. 3, 2019, págs. 765-771
  • Idioma: inglés
  • Enlaces
  • Resumen
    • Abstract This paper describes the process of collecting, maintaining and exploiting an English dataset of web discussions. The dataset consists of many web discussions with hand-annotated posts in the context of a tree structure of a web page. Each post consists of username, date, text, and citations used by its author. The dataset contains 79 different websites with at least 500 pages from each. Each web page consists of a tree structure of HTML tags with texts taken from selected web pages. In the paper, we also describe algorithms trained on the dataset. The algorithms employ basic architectures (such as a bag of words with an SVM classifier and an LSTM network) to set a baseline for the dataset.

Los metadatos del artículo han sido obtenidos de SciELO México

Fundación Dialnet

Dialnet Plus

  • Más información sobre Dialnet Plus

Opciones de compartir

Opciones de entorno