Hierarchical vs. Flat n-gram-based text categorization: Can we do better?

Graovac, Jelena; Kovačević, Jovana; Pavlović-Lažetić, Gordana

Please use this identifier to cite or link to this item: https://research.matf.bg.ac.rs/handle/123456789/1220

DC Field	Value	Language
dc.contributor.author	Graovac, Jelena	en_US
dc.contributor.author	Kovačević, Jovana	en_US
dc.contributor.author	Pavlović-Lažetić, Gordana	en_US
dc.date.accessioned	2022-09-29T15:57:21Z	-
dc.date.available	2022-09-29T15:57:21Z	-
dc.date.issued	2017-01-01	-
dc.identifier.issn	18200214	en
dc.identifier.uri	https://research.matf.bg.ac.rs/handle/123456789/1220	-
dc.description.abstract	Hierarchical text categorization (HTC) refers to assigning a text document to one or more most suitable categories from a hierarchical category space. In this paper we present two HTC techniques based on kNN and SVM machine learning techniques for categorization process and byte n-gram based document representation. They are fully language independent and do not require any text preprocessing steps, or any prior information about document content or language. The effectiveness of the presented techniques and their language independence are demonstrated in experiments performed on five tree-structured benchmark category hierarchies that differ in many aspects: Reuters-Hier1, Reuters-Hier2, 15NGHier and 20NGHier in English and TanCorpHier in Chinese. The results obtained are compared with the corresponding flat categorization techniques applied to leaf level categories of the considered hierarchies. While kNN-based flat text categorization produced slightly better results than kNN-based HTC on the largest TanCorpHier and 20NGHier datasets, SVM-based HTC results do not considerably differ from the corresponding flat techniques, due to shallow hierarchies; still, they outperform both kNN-based flat and hierarchical categorization on all corpora except the smallest Reuters-Hier1 and Reuters-Hier2 datasets. Formal evaluation confirmed that the proposed techniques obtained state-of-the-art results.	en_US
dc.language.iso	en	en_US
dc.publisher	Novi Sad : ComSIS Consortium	en_US
dc.relation.ispartof	Computer Science and Information Systems	en_US
dc.subject	Hierarchical text categorization	en_US
dc.subject	KNN	en_US
dc.subject	N-grams	en_US
dc.subject	SVM	en_US
dc.title	Hierarchical vs. Flat n-gram-based text categorization: Can we do better?	en_US
dc.type	Article	en_US
dc.identifier.doi	10.2298/CSIS151017030G	-
dc.identifier.scopus	2-s2.0-85011649741	-
dc.identifier.isi	000396389300007	-
dc.identifier.url	https://api.elsevier.com/content/abstract/scopus_id/85011649741	-
dc.contributor.affiliation	Informatics and Computer Science	en_US
dc.contributor.affiliation	Informatics and Computer Science	en_US
dc.relation.issn	1820-0214	en_US
dc.description.rank	M23	en_US
dc.relation.firstpage	103	en_US
dc.relation.lastpage	121	en_US
dc.relation.volume	14	en_US
dc.relation.issue	1	en_US
item.openairecristype	http://purl.org/coar/resource_type/c_18cf	-
item.cerifentitytype	Publications	-
item.openairetype	Article	-
item.grantfulltext	none	-
item.fulltext	No Fulltext	-
item.languageiso639-1	en	-
crisitem.author.dept	Informatics and Computer Science	-
crisitem.author.dept	Informatics and Computer Science	-
crisitem.author.orcid	0000-0002-9323-4695	-
crisitem.author.orcid	0000-0002-0242-2472	-
Appears in Collections:	Research outputs

Show simple item record

SCOPUS^TM
Citations

12

checked on Apr 16, 2026

Page view(s)

12

checked on Jan 19, 2025

Google Scholar^TM

Check

SCOPUS^TM
Citations

Page view(s)

Google Scholar^TM

Altmetric

Altmetric

SCOPUSTM Citations

Page view(s)

Google ScholarTM

Altmetric

Altmetric

SCOPUS^TM
Citations

Google Scholar^TM