A most Frontend Collection and survey of vision-language model papers, and models GitHub repository. Continuous updates.
-
Updated
Jul 29, 2026 - HTML
A most Frontend Collection and survey of vision-language model papers, and models GitHub repository. Continuous updates.
A repository for organizing papers, codes and other resources related to Virtual Try-on Models
[IEEE TII 2025] Official Implementation for "Dual-Detector Reoptimization for Federated Weakly Supervised Video Anomaly Detection via Adaptive Dynamic Recursive Mapping"
Code and data for "Timo: Towards Better Temporal Reasoning for Language Models" (COLM 2024)
[IEEE JSTARS 2026] Official Implementation of Mamba-FCS: Joint Spatio-Frequency Feature Fusion, Change-Guided Attention, and SeK Inspired Loss for Enhanced Semantic Change Detection in Remote Sensing
[DEPRECIATED] Very fast, large music transformer with 8k sequence length, efficient heptabit MIDI notes encoding, true full MIDI instruments range, chords counters and outro tokens
Researchers who published code, models (in some cases), and demo apps (in few cases) along with their SOTA paper
Python SDK for Lighton API, the 🇪🇺 European industrial-grade retrieval infrastructure powered by its frontier retrieval models developed in-house
SOTA pure drums transformer which is capable of drums track generation for any source composition
figsr — a frequency-domain (FFT-based) SISR architecture. Enhances detail reconstruction and inference speed, combining the strengths of CNNs and Transformers while mitigating their core limitations.
[SOTA] MIDI Tempo Detection AI implementation and model (94% accuracy on any MIDI]
[DEPRECIATED] [339M] [88% acc] Fast full-featured drums inpainting transformer with octo-velocity
This repository includes multiple competitions-solutions/tutorials in deep learning and machine learning
SOTA quality fast music transformer with symmetrical quad MIDI notes encoding
B.Sc. Thesis Deep Learning & NLP research on Medical Image Captioning
PyTorch implementation of the current SOTA Transformer. Configurable, efficient, and HuggingFace-compatible, serving as a baseline for research, benchmarking, and architectural experimentation.
Investigation of the capabilities of foundations models in the context of time series forecasting
QTrack: Query Driven Reasoning for Multimodal MOT
LangYuan is a SOTA-level new generation of generalized real-time digital human system.
A multi-agent real-time local discovery system with intent parsing, live place retrieval, review synthesis, transit-aware ranking, explainable recommendations, and SSE progress streaming.
Add a description, image, and links to the sota-model topic page so that developers can more easily learn about it.
To associate your repository with the sota-model topic, visit your repo's landing page and select "manage topics."