EPISODE · Mar 31, 2026 · 23 MIN
Ep 01 - Aprendizado de máquina e legibilidade em contabilidade: uma abordagem de ensemble learning
from PodAccounting · host PPGCC / UFPE
We employ a language model (FinBERT-PT-BR) trained in Brazilian Portuguese to develop a Informativeness Index, assigning scores to 26.804 yearly financial statement notes from 1152 companies in Brazil over the span of 12 years. Additionally, we calculate the usual readability metrics (Flesch-Kincaid reading ease, Fog index, SMOG index, Loughran-McDonald Index) for all the notes and employ machine learning models to evaluate which readability metric best represents an informativeness index built upon the dimensions of Boilerplateness, Completeness and Density. The evaluation of which readability metric is closets to measuring the informativeness of financial text is based on the feature importance, which indicates the best proxy for financial text readability of Portuguese text should be the Loughran-McDonald Index. These findings are in line with the literature. This research contributes to the literature by employing novel methods (Machinelearning and language models) within a not-so-explored field (Portuguese financial information) with a reasonably large dataset. Further research may be needed to aggregate different Language models or human experiments to increase the validity of the metric concept.
Embed this episode
NOW PLAYING
Ep 01 - Aprendizado de máquina e legibilidade em contabilidade: uma abordagem de ensemble learning
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.