Natural Language Processing Methods for the Analysis of Project Applications within the IPMA Methodology
Purpose: This study explores how natural language processing (NLP) can support the analysis and classification of project applications framed within the IPMA Individual Competence Baseline (ICB 4.0), addressing the labour-intensive nature of manually reviewing large volumes of funding applications. Design/Methodology/Approach: Using a corpus of 6,031 Polish-language project-application titles, we apply the Polish transformer language model HerBERT to represent the titles as embeddings and to classify them into the three IPMA competence areas (Perspective, People and Practice) by means of cosine similarity, validated against expert judgement on a sample ( ). This matching procedure is complemented by unsupervised Latent Dirichlet Allocation (LDA) topic modelling, with the number of topics selected using coherence and perplexity. Findings: HerBERT-based classification reached about agreement with expert assessment; performance varied markedly across the competence areas and between the maximum- and median-similarity aggregation rules. LDA revealed interpretable themes: the two-topic grouping separated a growth-and-recovery orientation from a stability-maintenance orientation, whereas the eight-topic solution disaggregated these into finer strategic and operational themes (e.g., resilience, digitalisation, crisis management and financial stability). Practical Implications: The proposed pipeline can partially automate the triage and comparison of project applications in institutions that process large numbers of submissions and can inform competence-based team composition. Originality/Value: The paper demonstrates the joint use of a Polish-specific transformer model and topic modelling for the competence-oriented analysis of project applications within the IPMA framework – a combination not previously reported for Polish-language project documentation.