Multimodal-AND-Large-Language-Models
Paper list about multimodal and large language models, only used to record papers I read in the daily arxiv for personal needs.
Research Scientist @ NVIDIA, working on LLMs post-training and LLM-based coding agents.
Paper list about multimodal and large language models, only used to record papers I read in the daily arxiv for personal needs.
[TMLR] Public code repo for paper "A Single Transformer for Scalable Vision-Language Modeling"
The released data for paper "Measuring and Improving Chain-of-Thought Reasoning in Vision-Language Models".
Mostly recording papers about models' trustworthy applications. Intending to include topics like model evaluation & analysis, security, calibration, backdoor learning, robustness, et al.
Public code repo for paper "Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training"
Source code for ACL 2023 Findings paper "Making Pre-trained Language Models both Task-solvers and Self-calibrators"
Source code for EMNLP 2023 paper "ViStruct: Visual Structural Knowledge Extraction via Curriculum Guided Code-Vision Representation"
This repo is meant to serve as a guide for Machine Learning/AI technical interviews.
[NeurIPS 2025 D&B] 🚀 SWE-bench Goes Live!
Advanced Melee Artificial Intelligence Mod For Warcraft 3
Public repository for Agent Skills