Classification of Small Datasets: Why Using Class-Based Weighting Measures?

Flavien Bouillot; Pascal Poncelet; Mathieu Roche

doi:10.1007/978-3-319-08326-1_35

Communication Dans Un Congrès Année : 2014

Classification of Small Datasets: Why Using Class-Based Weighting Measures?

(1, 2) , (1) , (3, 1)

1
2
3

Flavien Bouillot

Fonction : Auteur
PersonId : 927300

ADVanced Analytics for data SciencE

Itesoft R&D

Pascal Poncelet

Fonction : Auteur
PersonId : 6247
IdHAL : pascal-poncelet
ORCID : 0000-0002-8277-3490
IdRef : 069260613

ADVanced Analytics for data SciencE

Mathieu Roche

Fonction : Auteur
PersonId : 4967
IdHAL : mathieu-roche
ORCID : 0000-0003-3272-8568
IdRef : 09042087X

Territoires, Environnement, Télédétection et Information Spatiale

ADVanced Analytics for data SciencE

Résumé

In text classification, providing an efficient classifier even if the number of documents involved in the learning step is small remains an important issue. In this paper we evaluate the performance of traditional classification methods to better evaluate their limitation in the learning phase when dealing with small amount of documents. We thus propose a new way for weighting features which are used for classifying. These features have been integrated in two well known classifiers: Class-Feature-Centroid and Naïve Bayes, and evaluations have been performed on two real datasets. We have also investigated the influence on parameters such as number of classes, documents or words in the classification. Experiments have shown the efficiency of our proposal relatively to state of the art classification methods. Either with a very few amount of data or with a small number of features that can be extracted from poor content documents, we show that our approach performs well.

Domaines

Autre Intelligence artificielle [cs.AI] Recherche d'information [cs.IR] Traitement du texte et du document

Fichier principal

lirmm-01054900.pdf (372.07 Ko)

Origine : Fichiers produits par l'(les) auteur(s)

Mathieu Roche : Connectez-vous pour contacter le contributeur

https://hal-lirmm.ccsd.cnrs.fr/lirmm-01054900

Soumis le : mardi 6 novembre 2018-13:20:55

Dernière modification le : mardi 10 octobre 2023-16:38:10

Archivage à long terme le : jeudi 7 février 2019-15:15:59

Dates et versions

lirmm-01054900 , version 1 (06-11-2018)

Identifiants

HAL Id : lirmm-01054900 , version 1
DOI : 10.1007/978-3-319-08326-1_35

Citer

Flavien Bouillot, Pascal Poncelet, Mathieu Roche. Classification of Small Datasets: Why Using Class-Based Weighting Measures?. ISMIS: International Symposium on Methodologies for Intelligent Systems, Jun 2014, Roskilde, Denmark. pp.345-354, ⟨10.1007/978-3-319-08326-1_35⟩. ⟨lirmm-01054900⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

CIRAD AGROPARISTECH CNRS IRSTEA ADVANSE LIRMM AGROPOLIS TETIS MIPS UNIV-MONTPELLIER INRAE INRAEOCCITANIEMONTPELLIER MATHNUM

230 Consultations

153 Téléchargements

Classification of Small Datasets: Why Using Class-Based Weighting Measures?

Résumé

Domaines

Dates et versions

Identifiants

Citer

Exporter

Collections

Altmetric

Partager