LLM 8
- Anthropic's J-Space: A Global Workspace Inside Language Models
- Policy-Conditioned Policies for Multi-Agent Task Solving
- Talk, Judge, Cooperate: Gossip-Driven Indirect Reciprocity in Self-Interested LLM Agents
- Verbalized Bayesian Persuasion
- MS-Swift GRPO Pipeline Walkthrough
- LLM x RL
- LLM Architecture Speedrun
- Llama Memo