I am a PhD candidate in the Computer Vision Group at the University of Bonn, supervised by Prof. Dr. Juergen Gall, and currently working in the domain of Multimodal Video Understanding. Previously, I completed my master's degree at the University of Hanover, supervised by Prof. Dr. Bodo Rosenhahn. I hold an undergraduate degree in Electrical Engineering, with a minor in computer science, from the South China University of Technology.
My current research focuses on multimodal video understanding, long-term action anticipation and object re-indentification. Previously, I also worked on unsupervised domain adaptation, and applications of computer vision to plant phenotyping.
Introduces the first open-vocabulary evaluation setting for action anticipation, testing models on entirely unseen action vocabularies across egocentric datasets.
An open, locally deployable 8B cybersecurity LLM trained on a large curated corpus and agentic dialogue data, outperforming baselines on both security-specific and general benchmarks.
Combines adaptive deformable attention with hierarchical query learning to handle varying modality importance and spatial misalignment in multi-modal re-identification.
The first encoder-free video-language model, replacing heavy visual encoders with a lightweight alignment block for a large parameter and speed savings.
A parameter-efficient transformer that aggregates spatial, temporal, and keypoint cues with lightweight adapters for video-based person re-identification.
The first method to exploit multiple camera views in both the source and target domains for unsupervised domain adaptation of plant nutrient-deficiency detection.
Replaces the intermediate self-similarity matrix with full-resolution frame embeddings and a reference-based loss for more accurate repetitive action counting.
Benchmarked eight CNN- and transformer-based models on drone imagery of winter wheat and rye to detect specific nutrient-deficiency treatments, reaching 75-81% average accuracy within a season.
A lightweight graph convolutional network that refines noisy skeleton poses before recognition, cutting parameters and computation by roughly 90% with comparable accuracy.