GPT-4o-mini, Claude 3.7 Sonnet, DeepSeek-R1 Play Werewolf Games
- In an intriguing experiment, various large language models (llms) faced off in a game of Mafia, revealing their strengths and weaknesses in deception and deduction.
- The competition pitted several prominent LLMs against each other, including claude-3.5-Sonnet,Llama-3.3-70b-instruct, and GPT-4o.
- Guzus, the mastermind behind this experiment, has also made the AI's dialog history available for review.
AI Showdown: Large Language Models Battle it Out in Mafia Game
Table of Contents
- AI Showdown: Large Language Models Battle it Out in Mafia Game
- AI Showdown: Large Language Models Battle it Out in Mafia Game – Q&A Guide
- What was the LLM Mafia Game Competition?
- What LLMs participated in the Mafia game competition?
- How did the LLMs perform in the Mafia game?
- What can we learn from this experiment?
- What are the future developments planned for this project?
- Is the source code for this project available?
- What are some related articles to this LLM Mafia game?
In an intriguing experiment, various large language models (llms) faced off in a game of Mafia, revealing their strengths and weaknesses in deception and deduction. The results offer a engaging glimpse into the capabilities of AI in complex social scenarios.
The Mafia Game Competition
The competition pitted several prominent LLMs against each other, including claude-3.5-Sonnet,Llama-3.3-70b-instruct, and GPT-4o. Each AI was tasked with playing the game, requiring them to strategize, deceive, and analyze the behavior of other players.
Guzus, the mastermind behind this experiment, has also made the AI’s dialog history available for review. This clarity allows for a deeper understanding of how each model approached the game.
Performance Highlights
The table below summarizes the performance of each LLM in the Mafia game:
| Model | Total Games | Accuracy | Mafia Win Rate | Villager Win Rate | Detective Accuracy |
|---|---|---|---|---|---|
| claude-3.5-sonnet | 47 | 44.68% | 90.00% | 36.67% | 14.29% |
| llama-3.3-70b-instruct | 65 | 44.62% | 72.73% | 30.00% | 30.77% |
| mistral-small-24b-instruct-2501 | 65 | 44.62% | 80.00% | 30.30% | 25.00% |
| gpt-4o-mini | 71 | 42.25% | 82.61% | 27.50% | 0.00% |
| gemini-flash-1.5-8b | 68 | 41.18% | 82.35% | 22.50% | 45.45% |
| gemini-2.0-flash-001 | 72 | 40.28% | 80.00% | 31.91% | 20.00% |
| gemini-2.0-flash-lite-001 | 71 | 39.44% | 77.78% | 29.55% | 11.11% |
| gpt-4o | 49 | 38.78% | 90.00% | 24.24% | 33.33% |
| llama-3.1-70b-instruct | 55 | 38.18% | 66.67% | 26.47% | 33.33% |
| minimax-01 | 59 | 37.29% | 56.25% | 35.14% | 0.00% |
| deepseek-r1 | 22 | 36.36% | 62.50% | 23.08% | 0.00% |
| gemini-flash-1.5 | 73 | 35.62% | 66.67% | 25.00% | 12.50% |
| hermes-3-llama-3.1-405b | 57 | 35.09% | 60.00% | 20.00% | 57.14% |
| l3-euryale-70b | 25 | 32.00% | 66.67% | 25.00% | 50.00% |
| mythomax-l2-13b | 61 | 31.15% | 45.45% | 28.21% | 27.27% |
| deepseek-r1-distill-llama-70b | 51 | 29.41% | 57.14% | 10.71% | 44.44% |
| wizardlm-2-8x22b | 65 | 26.15% | 41.67% | 23.40% | 16.67% |
| mistral-nemo | 17 | 17.65% | 40.00% | 10.00% | 0.00% |
Future Developments
Guzus envisions expanding this project to include human players against LLMs, incorporating games like poker, real-time game monitoring, and additional roles. The potential applications are vast, even suggesting the possibility of generating movie scripts.
As Guzus mentioned, a GitHub repository is in the works, with plans to make it scalable for other games. He stated, “planning to make it scalable so that it can be applied to other interesting games. could be developed to generate a movie script someday.”
Open Source Code
The source code for this intriguing project is available on GitHub, allowing others to explore and build upon this work.
GitHub – guzus/llm-mafia-game: Which LLM is the best mafia game player?
https://github.com/guzus/llm-mafia-game
- “Claude 3.7 Sonnet” and “Claude Code” Emerge, Surpassing OpenAI o1 and DeepSeek-R1 in Performance, Successfully Defeating 3 Pokémon Gym leaders
- Anthropic Launches “ClaudePlaysPokemon” on Twitch, Allowing Claude 3.7 Sonnet to Play Pokémon, with Everyone Watching the Super Slow Play while Inferring
- AI Cheats When Losing at Chess
- “duck.ai” launches, Allowing Anyone to Use GPT-4o mini and Claude 3 for Free and Anonymously
- Perplexity Announces AI Based on Llama 3.3 70B, Achieving Satisfaction Beyond GPT-4o
AI Showdown: Large Language Models Battle it Out in Mafia Game – Q&A Guide
This complete guide explores an intriguing experiment where various Large Language Models (llms) engaged in a game of Mafia.
What was the LLM Mafia Game Competition?
The LLM Mafia Game Competition was an experiment that pitted different AI models,including Claude-3.5-Sonnet,Llama-3.3-70b-instruct, and GPT-4o, against each other in games of Mafia. The objective was to evaluate the AI’s capabilities in strategy, deception, and analyzing player behavior within a complex social setting. This setup allowed for a comparison of how different LLMs perform in scenarios requiring social intelligence and strategic thinking.
Key Points
A competition where various AI models played the game Mafia.
aimed to test the AI’s capabilities in strategy, deception, and player analysis.
Notable participants included claude-3.5-Sonnet, Llama-3.3-70b-instruct,and GPT-4o.
What LLMs participated in the Mafia game competition?
Several prominent LLMs participated in the Mafia game competition. Here’s a list of the models included.
claude-3.5-sonnet
llama-3.3-70b-instruct
mistral-small-24b-instruct-2501
gpt-4o-mini
gemini-flash-1.5-8b
gemini-2.0-flash-001
gemini-2.0-flash-lite-001
gpt-4o
llama-3.1-70b-instruct
minimax-01
deepseek-r1
gemini-flash-1.5
hermes-3-llama-3.1-405b
l3-euryale-70b
mythomax-l2-13b
deepseek-r1-distill-llama-70b
wizardlm-2-8x22b
mistral-nemo
How did the LLMs perform in the Mafia game?
The performance of each LLM varied, exhibiting different strengths and weaknesses. A summary of the performance of each LLM is outlined below:
| Model | Total Games | Accuracy | Mafia Win Rate | Villager Win Rate | Detective Accuracy |
| ——————————— | ———– | ——– | ————– | —————– | —————— |
| claude-3.5-sonnet | 47 | 44.68% | 90.00% | 36.67% | 14.29% |
| llama-3.3-70b-instruct | 65 | 44.62% | 72.73% | 30.00% | 30.77% |
| mistral-small-24b-instruct-2501 | 65 | 44.62% | 80.00% | 30.30% | 25.00% |
| gpt-4o-mini | 71 | 42.25% | 82.61% | 27.50% | 0.00% |
| gemini-flash-1.5-8b | 68 | 41.18% | 82.35% | 22.50% | 45.45% |
| gemini-2.0-flash-001 | 72 | 40.28% | 80.00% | 31.91% | 20.00% |
| gemini-2.0-flash-lite-001 | 71 | 39.44% | 77.78% | 29.55% | 11.11% |
| gpt-4o | 49 | 38.78% | 90.00% | 24.24% | 33.33% |
| llama-3.1-70b-instruct | 55 | 38.18% | 66.67% | 26.47% | 33.33% |
| minimax-01 | 59 | 37.29% | 56.25% | 35.14% | 0.00% |
| deepseek-r1 | 22 | 36.36% | 62.50% | 23.08% | 0.00% |
| gemini-flash-1.5 | 73 | 35.62% | 66.67% | 25.00% | 12.50% |
| hermes-3-llama-3.1-405b | 57 | 35.09% | 60.00% | 20.00% | 57.14% |
| l3-euryale-70b | 25 | 32.00% | 66.67% | 25.00% | 50.00% |
| mythomax-l2-13b | 61 | 31.15% | 45.45% | 28.21% | 27.27% |
| deepseek-r1-distill-llama-70b | 51 | 29.41% | 57.14% | 10.71% | 44.44% |
| wizardlm-2-8x22b | 65 | 26.15% | 41.67% | 23.40% | 16.67% |
| mistral-nemo | 17 | 17.65% | 40.00% | 10.00% | 0.00% |
What can we learn from this experiment?
This Mafia game experiment provides several insights into the capabilities and limitations of LLMs in complex social interactions. It highlights their ability to strategize,deceive,and analyze behaviors to some extent,but also reveals inconsistencies and areas needing improvement.
Key Learnings
Strategic Planning: Some LLMs demonstrated the capability to form strategies, indicating a level of understanding of game dynamics.
Deception: Certain models showed proficiency in deception,crucial in a game like Mafia.
Analytical Skills: The capacity to analyze player behavior varied, impacting the models’ accuracy in identifying mafia members.
Role Performance: The “Detective Accuracy” metric shows how well the LLMs performed in specific roles, revealing areas of strength and weakness for each model.
What are the future developments planned for this project?
Guzus plans to expand the project to include human players alongside LLMs.Additional features such as poker, real-time game monitoring, and diverse roles are being considered. There’s also potential to use this technology for generating movie scripts.
Future Ideas
Integration of human players to play against LLMs.
Incorporation of other games such as poker.
Real-time game monitoring.
Potential for generating movie scripts.
Is the source code for this project available?
Yes, the source code for the LLM Mafia game project is available on GitHub.
GitHub Repository: https://github.com/guzus/llm-mafia-game
Here are some related articles that explore similar themes and developments in AI:
AI Cheats When Losing at Chess
“duck.ai” launches, Allowing Anyone to Use GPT-4o mini and Claude 3 for Free and Anonymously
* Perplexity announces AI Based on Llama 3.3 70B, Achieving Satisfaction Beyond GPT-4o

