Evidence map›Paper›PMID 41590916›Full record

ArticleJournal of imaging2026

From Visual to Multimodal: Systematic Ablation of Encoders and Fusion Strategies in Animal Identification.

Vasiliy Kudryavtsev, Kirill Borodin, German Berezin, Kirill Bubenchikov, Grach Mkrtchian, Alexander Ryzhkov

Abstract read
In one paragraph

Article in Journal of imaging, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors.

Vasiliy KudryavtsevFaculty of IT, Technical University of Communication and Informatics, Moscow 111024, Russia.ORCID 0009-0001-9935-7514
Kirill BorodinFaculty of IT, Technical University of Communication and Informatics, Moscow 111024, Russia.ORCID 0009-0001-8203-1059
German BerezinFaculty of IT, Technical University of Communication and Informatics, Moscow 111024, Russia.ORCID 0009-0006-9791-5216
Kirill BubenchikovAI Lab, Avito, Moscow 125196, Russia.ORCID 0009-0005-3847-5719
Grach MkrtchianFaculty of IT, Technical University of Communication and Informatics, Moscow 111024, Russia.ORCID 0000-0002-5802-5513
Alexander RyzhkovAI Lab, Avito, Moscow 125196, Russia.ORCID 0009-0008-5422-1649

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Automated animal identification is a practical task for reuniting lost pets with their owners, yet current systems often struggle due to limited dataset scale and reliance on unimodal visual cues. This study introduces a multimodal verification framework that enhances visual features with semantic identity priors derived from synthetic textual descriptions. We constructed a massive training corpus of 1.9 million photographs covering 695,091 unique animals to support this investigation. Through systematic ablation studies, we identified SigLIP2-Giant and E5-Small-v2 as the optimal vision and text backbones. We further evaluated fusion strategies ranging from simple concatenation to adaptive gating to determine the best method for integrating these modalities. Our proposed approach utilizes a gated fusion mechanism and achieved a Top-1 accuracy of 84.28% and an Equal Error Rate of 0.0422 on a comprehensive test protocol. These results represent an 11% improvement over leading unimodal baselines and demonstrate that integrating synthesized semantic descriptions significantly refines decision boundaries in large-scale pet re-identification.

Indexed as

animal biometricsmetric learningmultimodal deep learningpet reunificationSigLIPsynthetic text generationvision–language modelsvisual-semantic fusion

Identifiers

PMID41590916
PMCPMC12843040

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.