MMTM (tri-modal topic modeling)
Extracts themes from long-form video responses with a tri-modal topic modeling workflow.
Open GitHub repositoryMeasure 2 - CIRCLET Documentation
CIRCLET provides the research-driven infrastructure for processing, analyzing, and integrating advanced survey-related data within ENTAILab.
Infrastructure for advanced survey-related data
Built on DUUI, CIRCLET supports scalable and reproducible analysis of heterogeneous data sources, including textual data, multimodal material such as images, audio, and video, as well as behavioral observations.
The infrastructure enables researchers to combine computational methods with domain-specific expertise in a shared analytical environment. It supports the integration of AI-based methods for tasks such as annotation, classification, interpretation, model fine-tuning, and in-context learning. This makes CIRCLET suitable for workflows ranging from individual case studies to large-scale and distributed data processing.
CIRCLET is guided by the principles of interoperability, reusability, extensibility, validity, and reproducibility. Analytical tools and methods are packaged as modular, containerized pipeline components and orchestrated through DUUI. These workflows can be documented, repeated, extended, and reused across projects, making research processes more transparent and sustainable.
CIRCLET follows a project-driven development approach by working closely with individual SPP projects to solve project-specific tasks. At the same time, it aims to generalize and reuse the resulting solutions across the SPP. This allows CIRCLET to remain closely oriented toward the needs of individual projects while also adopting a cross-project perspective.
With CIRCLET's support, projects should be able to adopt and apply state-of-the-art methods, particularly AI-based methods. This supports technology transfer from outside the SPP into individual projects. Conversely, innovations developed within individual projects can be made available for reuse by other projects, enabling technology transfer within and beyond the SPP.
Table 1. Software provided, hosted, and adapted by CIRCLET.
†To be released soon.
Extracts themes from long-form video responses with a tri-modal topic modeling workflow.
Open GitHub repositoryIdentifies themes in donated visual material within the SmartDYN context.
Open GitHub repositorySupports individual-respondent modeling at scale with retrieval-augmented generation and LLM MCQ fine-tuning.
Open GitHub repositoryOpinion measurement in underserved languages.
Open GitHub repositoryDetects inconsistencies between what is said and what is shown across modalities.
Open GitHub repositoryAs shown in Table 1, CIRCLET provides, hosts, and adapts software components for shared use across the SPP. This includes Open WebUI, which provides unified and standardized access to Ollama instances and is suitable for web-based prototyping and API-based NLP tasks. JupyterHub enables the dedicated execution of isolated Jupyter notebooks and can be used as a web-based, interactive development environment for Python-based code and data. Similar to Google Colab, it uses local resources and GPUs, allowing users to prototype and evaluate applications before integrating them into big data infrastructures such as DUUI. Slurm is a high-performance computing solution that enables the distribution and shared, fair use of CPU and GPU resources within a cluster; it can be used both through Jupyter notebooks and via DUUI Zhou, Abrami, and Mehler, 2026.
The project cards summarize pipelines developed through collaborations with individual SPP projects.
CIRCLET has contributed to standardizing the use of tools and systems for natural language processing and multimodal computing, including text Wigbels et al., 2026, image Weiss et al., 2026, and video processing Bundan, Abrami, and Mehler, 2025.
This standardization has been implemented through DUUI, the Docker Unified UIMA Interface Leonhardt et al., 2023. DUUI provides an interoperable environment that makes existing tools and services reusable for distributed big-data processing Abrami et al., 2025.
DUUIgateway extends this infrastructure by providing a web- and API-based access layer for DUUI Borkowski et al., 2026. It facilitates the deployment, management, and monitoring of NLP workflows and makes DUUI-based processing more accessible to non-expert users.
DUUI provides an expanding set of components for text, image, audio, and video processing. These components currently cover tasks ranging from topic modeling, emotion detection, and sentiment analysis to multimodal processing and vision-language models. They are continuously extended in response to the needs of the SPP projects.
Detects and anonymizes personal or sensitive information in text.
Open GitHub repositoryClassifies emotional expression in text using transformer-based models.
Open GitHub repositoryAssesses the factual accuracy of textual claims.
Open GitHub repositoryClassifies documents according to genre.
Open GitHub repositoryGenerates textual descriptions from image content.
Open GitHub repositoryProcesses and integrates text, image, audio, and video data within unified pipelines.
Open GitHub repositoryIdentifies and classifies named entities in text.
Open GitHub repositoryProvides unified, standardized web access to LLMs across text, image, and audio modalities.
Open GitHub repositoryDetermines sentiment polarity of text at a fine-grained level.
Open GitHub repositoryExtracts latent topics from text corpora.
Open GitHub repositoryApplies vision-language models to interpret and classify image content.
Open GitHub repository@unpublished{Zhou:et:al:prep,
title = {DUUISlurm},
author = {Zhou, Jiadong and Abrami, Giuseppe and Mehler, Alexander},
note = {In preparation},
year = {2026}
}@inproceedings{Wigbels:et:al:2026:a,
author = {Wigbels, Christoph and Abusaleh, Ali and Jansen, Markus T. and Mehler, Alexander and Schaaf, Manuel and Hofmann, Markus J.},
title = {Individual Text Corpora Predict User-Specific Knowledge: Benchmarks of Individualized Knowledge Simulation},
booktitle = {KONVENS (Konferenz zur Verarbeitung natürlicher Sprache)},
year = {2026},
keywords = {Individual Text Corpora, Retrieval-Augmented Generation, Personalized Language Models, Probabilistic Calibration, Knowledge Benchmarks, Data Contamination, German NLP},
note = {submitted}
}@inproceedings{weiss:et:al:2026,
series = {WebSci Companion '26},
title = {From Images to Topics: Evaluating Vision-Language Models for Topic Classification of Election Advertising},
url = {http://dx.doi.org/10.1145/3795513.3807426},
doi = {10.1145/3795513.3807426},
booktitle = {Companion Publication of the 2026 18th ACM Web Science Conference},
publisher = {ACM},
author = {Weiss, Julia and Burger, Axel and Roßmann, Joss and Meurer, Jan Eric and Abusaleh, Ali},
year = {2026},
month = {May},
pages = {10--14},
collection = {WebSci Companion '26},
keywords = {Multimodal Large Language Models, Political communication, Privacy-aware AI, new-data-spaces, circlet}
}@inproceedings{Bundan:Abrami:Mehler:2025,
author = {Bundan, Daniel and Abrami, Giuseppe and Mehler, Alexander},
title = {Multimodal Docker Unified {UIMA} Interface: New Horizons for Distributed Microservice-Oriented Processing of Corpora using {UIMA}},
booktitle = {Proceedings of the 21st Conference on Natural Language Processing (KONVENS 2025): Long and Short Papers},
year = {2025},
editor = {Wartena, Christian and Heid, Ulrich},
location = {Hildesheim, Germany},
address = {Hannover, Germany},
publisher = {HsH Applied Academics},
pages = {257--268},
series = {KONVENS '25},
url = {https://aclanthology.org/2025.konvens-1.22/},
pdf = {https://aclanthology.org/2025.konvens-1.22.pdf}
}@inproceedings{Leonhardt:et:al:2023,
title = {Unlocking the Heterogeneous Landscape of Big Data {NLP} with {DUUI}},
author = {Leonhardt, Alexander and Abrami, Giuseppe and Baumartz, Daniel and Mehler, Alexander},
editor = {Bouamor, Houda and Pino, Juan and Bali, Kalika},
booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2023},
year = {2023},
address = {Singapore},
publisher = {Association for Computational Linguistics},
url = {https://aclanthology.org/2023.findings-emnlp.29},
pages = {385--399},
pdf = {https://aclanthology.org/2023.findings-emnlp.29.pdf}
}@article{Abrami:et:al:2025:a,
title = {Docker Unified UIMA Interface: New perspectives for NLP on big data},
journal = {SoftwareX},
volume = {29},
pages = {102033},
year = {2025},
issn = {2352-7110},
doi = {https://doi.org/10.1016/j.softx.2024.102033},
url = {https://www.sciencedirect.com/science/article/pii/S2352711024004047},
author = {Giuseppe Abrami and Markos Genios and Filip Fitzermann and Daniel Baumartz and Alexander Mehler}
}@article{Borkowski:et:al:2026,
title = {{DUUIgateway}: A Web Service for Platform-independent, Ubiquitous Big Data NLP},
journal = {SoftwareX},
volume = {34},
pages = {102549},
year = {2026},
issn = {2352-7110},
doi = {https://doi.org/10.1016/j.softx.2026.102549},
url = {https://www.sciencedirect.com/science/article/pii/S2352711026000439},
author = {Borkowski, Cedric and Abrami, Giuseppe and Terefe, Dawit and Baumartz, Daniel and Mehler, Alexander},
keywords = {duui, neglab, core, core_b05, core_c08, new-data-spaces, circlet}
}