Computer Science, Computer Engineering, Electerical Engineering, etc.
-3rd Year Students
-4th Year Students- Seniors
214
Fully remote/Remote considered
This project studies large language models (LLMs) that perform reasoning and inter-agent communication directly in latent space rather than through explicit natural language tokens. The core idea is to represent intermediate thoughts, plans, and messages as continuous latent states that are learned end to end and exchanged among agents. The project will develop algorithms and model architectures that allow agents to generate, transform, and interpret latent representations for tasks that require long-horizon reasoning, coordination, and information sharing. Key goals include analyzing how latent-space interaction affects reasoning depth, sample efficiency, and robustness, and comparing these properties with token-level communication. The project will also provide empirical and theoretical analysis of when latent communication improves performance, how it can be aligned with external supervision, and how latent states can be decoded or constrained to support interpretability and control.
✅ Strong in programming (Python/C++, with experience in deep learning frameworks like PyTorch)
✅ Solid in mathematical foundations (e.g., probability, statistics, linear algebra, optimization)
✅ Knowledgeable in AI/ML (machine learning, data mining, or AI fundamentals; has taken CSE 475/476 or equivalent)
✅ Passionate about research (publications in ML or interdisciplinary venues are a plus, but not required)
✅ Self-motivated, curious, and ready to take initiative – we encourage students to lead projects and publish at top venues!
Xiyang Hu
Fully remote/Remote considered