Fully autonomous & self-evolving research from idea to paper. Chat an Idea. Get a Paper. 🦞
-
Updated
Jul 13, 2026 - Python
Fully autonomous & self-evolving research from idea to paper. Chat an Idea. Get a Paper. 🦞
Official code repo for NeurIPS 2025 Spotlight paper, "Debate or Vote: Which Yields Better Decisions in Multi-Agent LLMs?"
Framework: Multi-Agent LLMs For Conversational Task-Solving (MALLM)
Research-backed methodology for multi-AI collaborative decision-making with structured debate, consensus synthesis, and bias reduction
Source code for the paper: Hear Both Sides: Efficient Multi-Agent Debate via Diversity-Aware Message Retention
Human-in-the-loop adversarial workflows for high-stakes research audit: from ChatGPT-Gemini duels to 4-model MAD.
An adversarial AI expert workshop that stress-tests a research paper (rival-tradition referees argue; every comment quote-grounded and independently re-verified) and then rebuilds it: tracked-changes redline, clean version, your code re-run under a provenance wall, and a replication package. A Claude Code skill.
Code for "Multiple LLM Agents Debate for Equitable Cultural Alignment" [ACL 2025 Oral]
Code review, but with 5 models arguing first.
Three Claude Code skills for working with Codex CLI: codex-bridge (one-shot Codex calls), mad-build (Claude+Codex collaboration with cross-review), and mad-research (three-stream adversarial audit of papers, grants, reports with anonymized cross-critique and fresh-Codex synthesis).
Multi-model deliberative design review — a Claude Code skill that runs structured debate between Gemini and GPT to surface blind spots in architecture decisions.
Control Stream Deck buttons for Claude Code terminal sessions with live status, tap-to-focus, hold-to-dictate, and snap-to-grid window layout on macOS
A brutally fault-tolerant Mixture-of-Agents (MoA) pipeline built in pure Python. Designed to orchestrate chaotic, round-robin LLM proxy endpoints through a rigorous 4-stage Agentic Workflow (Generate ➔ Cross-Critique ➔ Rebuttal ➔ Judge). Built to eradicate hallucination and guarantee absolute accuracy in complex, multi-step reasoning tasks.
Claude Code plugin for second opinions: iterative adversarial debates between Claude and Codex over any subject, from specs, designs, and code to non-technical positions, ending in genuine agreement or a documented dispute.
Claude Code plugin: open a review topic and AI agents (Claude, Codex) debate it round after round — design, attack, rebuttal — fully hands-free until they deliver a reasoned decision.md. Multi-topic priority queue, human gate only on code changes, git/non-git, claude-solo fallback.
Multi-LLM debate orchestrator that drives ChatGPT, Claude, and DeepSeek web UIs (no API keys) through a 5-phase loop: propose → critique → revise → synthesize → ratify-or-veto. Editorial dark UI.
Build autonomous ML research in Elixir: design, train, and iterate GPT models across GPUs with fault-tolerant BEAM concurrency
Generate research papers autonomously by chatting with OpenClaw, using Python 3.11+, with a self-evolving framework and extensive test coverage.
Research paper on how agentic debate pipelines can be constructed to reduce hallucinations in LLMs with open-source and commercial models
Multi-agent debate framework for long-form complex research. 11 AI departments cross-validate every claim with confidence annotations. Not one LLM — an adversarial committee. 共识管线:多智能体学术辩论框架,11个AI部门交叉验证,长线复杂任务一站解决。
Add a description, image, and links to the multi-agent-debate topic page so that developers can more easily learn about it.
To associate your repository with the multi-agent-debate topic, visit your repo's landing page and select "manage topics."