Skip to content

Research-Driven Infrastructure for Advanced Survey-Related Data (CIRCLET)

Measure 2 - CIRCLET Documentation

CIRCLET provides the research-driven infrastructure for processing, analyzing, and integrating advanced survey-related data within ENTAILab.

Infrastructure for advanced survey-related data

Built on DUUI, CIRCLET supports scalable and reproducible analysis of heterogeneous data sources, including textual data, multimodal material such as images, audio, and video, as well as behavioral observations.

Introduction

The infrastructure enables researchers to combine computational methods with domain-specific expertise in a shared analytical environment. It supports the integration of AI-based methods for tasks such as annotation, classification, interpretation, model fine-tuning, and in-context learning. This makes CIRCLET suitable for workflows ranging from individual case studies to large-scale and distributed data processing.

CIRCLET is guided by the principles of interoperability, reusability, extensibility, validity, and reproducibility. Analytical tools and methods are packaged as modular, containerized pipeline components and orchestrated through DUUI. These workflows can be documented, repeated, extended, and reused across projects, making research processes more transparent and sustainable.

Software Provided, Hosted, and Adapted by CIRCLET

CIRCLET follows a project-driven development approach by working closely with individual SPP projects to solve project-specific tasks. At the same time, it aims to generalize and reuse the resulting solutions across the SPP. This allows CIRCLET to remain closely oriented toward the needs of individual projects while also adopting a cross-project perspective.

With CIRCLET's support, projects should be able to adopt and apply state-of-the-art methods, particularly AI-based methods. This supports technology transfer from outside the SPP into individual projects. Conversely, innovations developed within individual projects can be made available for reuse by other projects, enabling technology transfer within and beyond the SPP.

Table 1. Software provided, hosted, and adapted by CIRCLET.

Provided SystemYear
1Open WebUI2025
2ollama-Server2025
3JupyterHub2026
4Slurm2026

To be released soon.

Project-specific pipelines developed with SPP projects

Python 3.12Jupyter NotebookPipeline: MMTMTask: topic modeling

MMTM (tri-modal topic modeling)

Extracts themes from long-form video responses with a tri-modal topic modeling workflow.

Open GitHub repository
Python 3.12Project: SmartDYNModality: visual materialTask: topic modeling

Image-to-topic modeling

Identifies themes in donated visual material within the SmartDYN context.

Open GitHub repository
Python 3.12Project: IndividualTextCorporaPhase1Method: RAGLLM: MCQ fine-tuning

RAG + LLM MCQ fine-tuning

Supports individual-respondent modeling at scale with retrieval-augmented generation and LLM MCQ fine-tuning.

Open GitHub repository
Python 3.12Project: CrossModalNegationModality: multimodalTask: negation analysis

Multimodal negation analysis

Detects inconsistencies between what is said and what is shown across modalities.

Open GitHub repository

As shown in Table 1, CIRCLET provides, hosts, and adapts software components for shared use across the SPP. This includes Open WebUI, which provides unified and standardized access to Ollama instances and is suitable for web-based prototyping and API-based NLP tasks. JupyterHub enables the dedicated execution of isolated Jupyter notebooks and can be used as a web-based, interactive development environment for Python-based code and data. Similar to Google Colab, it uses local resources and GPUs, allowing users to prototype and evaluate applications before integrating them into big data infrastructures such as DUUI. Slurm is a high-performance computing solution that enables the distribution and shared, fair use of CPU and GPU resources within a cluster; it can be used both through Jupyter notebooks and via DUUI Zhou, Abrami, and Mehler, 2026.

The project cards summarize pipelines developed through collaborations with individual SPP projects.

Computational Scalability (DUUI)

CIRCLET has contributed to standardizing the use of tools and systems for natural language processing and multimodal computing, including text Wigbels et al., 2026, image Weiss et al., 2026, and video processing Bundan, Abrami, and Mehler, 2025.

This standardization has been implemented through DUUI, the Docker Unified UIMA Interface Leonhardt et al., 2023. DUUI provides an interoperable environment that makes existing tools and services reusable for distributed big-data processing Abrami et al., 2025.

DUUIgateway extends this infrastructure by providing a web- and API-based access layer for DUUI Borkowski et al., 2026. It facilitates the deployment, management, and monitoring of NLP workflows and makes DUUI-based processing more accessible to non-expert users.

DUUI provides an expanding set of components for text, image, audio, and video processing. These components currently cover tasks ranging from topic modeling, emotion detection, and sentiment analysis to multimodal processing and vision-language models. They are continuously extended in response to the needs of the SPP projects.

DUUI components provided by CIRCLET.

Python 3.10Component: Emotion DetectionModality: textTask: emotion detectionYear: 2024Commits: 19Docker images: 87

Emotion Detection

Classifies emotional expression in text using transformer-based models.

Open GitHub repository
Python 3.10Component: Multimodal ProcessingModality: text, image, audio, videoTask: multimodal processingYear: 2025Commits: 41Docker images: 9

Multimodal Processing

Processes and integrates text, image, audio, and video data within unified pipelines.

Open GitHub repository
Python 3.10Component: Open WebUIModality: text, image, audioTask: LLM accessYear: 2026Commits: 5Docker images: 1

Open WebUI

Provides unified, standardized web access to LLMs across text, image, and audio modalities.

Open GitHub repository
Python 3.10Component: Vision-Language ModelsModality: imageTask: vision-languageYear: 2025Commits: 9Docker images: 1

Vision-Language Models

Applies vision-language models to interpret and classify image content.

Open GitHub repository
BibTeX references

Zhou, Abrami, and Mehler, 2026:

@unpublished{Zhou:et:al:prep,
title  = {DUUISlurm},
author = {Zhou, Jiadong and Abrami, Giuseppe and Mehler, Alexander},
note   = {In preparation},
year   = {2026}
}

Wigbels et al., 2026:

@inproceedings{Wigbels:et:al:2026:a,
author    = {Wigbels, Christoph and Abusaleh, Ali and Jansen, Markus T. and Mehler, Alexander and Schaaf, Manuel and Hofmann, Markus J.},
title     = {Individual Text Corpora Predict User-Specific Knowledge: Benchmarks of Individualized Knowledge Simulation},
booktitle = {KONVENS (Konferenz zur Verarbeitung natürlicher Sprache)},
year      = {2026},
keywords  = {Individual Text Corpora, Retrieval-Augmented Generation, Personalized Language Models, Probabilistic Calibration, Knowledge Benchmarks, Data Contamination, German NLP},
note      = {submitted}
}

Weiss et al., 2026:

@inproceedings{weiss:et:al:2026,
series     = {WebSci Companion '26},
title      = {From Images to Topics: Evaluating Vision-Language Models for Topic Classification of Election Advertising},
url        = {http://dx.doi.org/10.1145/3795513.3807426},
doi        = {10.1145/3795513.3807426},
booktitle  = {Companion Publication of the 2026 18th ACM Web Science Conference},
publisher  = {ACM},
author     = {Weiss, Julia and Burger, Axel and Roßmann, Joss and Meurer, Jan Eric and Abusaleh, Ali},
year       = {2026},
month      = {May},
pages      = {10--14},
collection = {WebSci Companion '26},
keywords   = {Multimodal Large Language Models, Political communication, Privacy-aware AI, new-data-spaces, circlet}
}

Bundan, Abrami, and Mehler, 2025:

@inproceedings{Bundan:Abrami:Mehler:2025,
author    = {Bundan, Daniel and Abrami, Giuseppe and Mehler, Alexander},
title     = {Multimodal Docker Unified {UIMA} Interface: New Horizons for Distributed Microservice-Oriented Processing of Corpora using {UIMA}},
booktitle = {Proceedings of the 21st Conference on Natural Language Processing (KONVENS 2025): Long and Short Papers},
year      = {2025},
editor    = {Wartena, Christian and Heid, Ulrich},
location  = {Hildesheim, Germany},
address   = {Hannover, Germany},
publisher = {HsH Applied Academics},
pages     = {257--268},
series    = {KONVENS '25},
url       = {https://aclanthology.org/2025.konvens-1.22/},
pdf       = {https://aclanthology.org/2025.konvens-1.22.pdf}
}

Leonhardt et al., 2023:

@inproceedings{Leonhardt:et:al:2023,
title     = {Unlocking the Heterogeneous Landscape of Big Data {NLP} with {DUUI}},
author    = {Leonhardt, Alexander and Abrami, Giuseppe and Baumartz, Daniel and Mehler, Alexander},
editor    = {Bouamor, Houda and Pino, Juan and Bali, Kalika},
booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2023},
year      = {2023},
address   = {Singapore},
publisher = {Association for Computational Linguistics},
url       = {https://aclanthology.org/2023.findings-emnlp.29},
pages     = {385--399},
pdf       = {https://aclanthology.org/2023.findings-emnlp.29.pdf}
}

Abrami et al., 2025:

@article{Abrami:et:al:2025:a,
title   = {Docker Unified UIMA Interface: New perspectives for NLP on big data},
journal = {SoftwareX},
volume  = {29},
pages   = {102033},
year    = {2025},
issn    = {2352-7110},
doi     = {https://doi.org/10.1016/j.softx.2024.102033},
url     = {https://www.sciencedirect.com/science/article/pii/S2352711024004047},
author  = {Giuseppe Abrami and Markos Genios and Filip Fitzermann and Daniel Baumartz and Alexander Mehler}
}

Borkowski et al., 2026:

@article{Borkowski:et:al:2026,
title    = {{DUUIgateway}: A Web Service for Platform-independent, Ubiquitous Big Data NLP},
journal  = {SoftwareX},
volume   = {34},
pages    = {102549},
year     = {2026},
issn     = {2352-7110},
doi      = {https://doi.org/10.1016/j.softx.2026.102549},
url      = {https://www.sciencedirect.com/science/article/pii/S2352711026000439},
author   = {Borkowski, Cedric and Abrami, Giuseppe and Terefe, Dawit and Baumartz, Daniel and Mehler, Alexander},
keywords = {duui, neglab, core, core_b05, core_c08, new-data-spaces, circlet}
}