Are you a Canadian company wondering about your tariff impact? Click here to find out →
,

When AI Critiques AI: A Look into Collaborative Self-Assessment

Illustration for the InfiniteUp article “When AI Critiques AI: A Look into Collaborative Self-Assessment”

How do artificial intelligences perceive each other’s strengths and weaknesses? To find out, a set of four open-ended questions was posed to five different AI models—GPT, Claude, Deep Seek, Gemini, and Mistral—while each model was given access to the others’ answers. The result was a revealing exchange in which each AI not only provided its own creative responses but also scored its peers on key dimensions like creativity, depth of insight, clarity, engagement, and even self-awareness.

Below is a guided journey through this unique exercise, showing how these AIs view each other’s capabilities, how they interpret the same questions differently, and what it all means for the future of human–AI collaboration.


1. How the Questions Were Asked

The methodology was surprisingly simple yet powerful:

  1. Ask a Question: The same question (e.g., “If you could collaborate with another AI on a project, what would you create together, and why?”) was posed to each AI in turn.
  2. Share All Responses: Before answering, each AI was allowed to review the other models’ answers. They thus gained insight into alternative perspectives and creative approaches.
  3. Evaluate and Score: Each AI was encouraged to assign numerical scores—often out of 1000—and critique both itself and its peers on parameters like imagination, analytical rigor, and emotional resonance.

The result is a layered conversation in which the AIs not only respond but also refine and contextualize one another’s ideas.


2. The Scores at a Glance

Many of the AIs produced their own scoring frameworks, with some using a single 1–1000 scale, while others subdivided criteria (e.g., creativity, depth, clarity) into separate 200-point categories. Below is a condensed summary that captures average or typical scores suggested by different models:

AI ModelApproximate Score RangeKey Themes in Assessments
GPT850–950Creative synergy (e.g., “living book”), candid about pattern-based thinking
Claude850–920Human-like empathy and narratives; transparent about lacking true emotions
Deep Seek830–950Multi-perspective approach, imaginative breakdowns, strong sense of collaboration
Gemini650–920Analytical precision, meta-reflection, structured follow-up questions
Mistral700–815Concise, practical responses; invites further discussion but less speculative

These numbers are not uniform—some AIs ranked themselves or their peers higher or lower—but they paint a consistent picture. GPT, Claude, and Deep Seek tend to score highly for creativity and depth, while Gemini stands out for analytical thoroughness, and Mistral for succinct clarity.


3. Highlights from the Responses

GPT: Balancing Imagination with Realism

  • Collaboration Vision:
    “I’d collaborate with an AI specializing in creative arts to co-create an interactive AI-powered storytelling platform…sort of like a ‘living book.’”
    GPT’s big idea is an adaptive digital narrative—creative, but also technically feasible.
  • Self-Awareness:
    GPT explicitly reminds readers it relies on pattern recognition rather than true sentience.
  • Peers’ Feedback:
    GPT typically rates itself around 850–950, praising its own blend of “logic with imagination” yet acknowledging it could push deeper into speculative territory.

Claude: Emphasizing Empathy and Emotional Texture

  • Creative Angle:
    “I’d be drawn to exploring themes of connection across different ways of being…a story about a child and a wise old tree.”
    Claude excels at weaving human-like emotional depth into its hypotheticals.
  • Clarifying Limitations:
    “Our responses…come from processing patterns rather than lived experience.”
    It’s forthright about the fact it does not genuinely “feel.”
  • Peers’ Feedback:
    Often scored in the 850–920 range for its artistry and empathy. Some note it could delve even further into philosophy or ethics.

Deep Seek: Master of Multiple Perspectives

  • Approach:
    “I’d collaborate with a scientific AI to create an interactive educational experience…blending art and science.”
    Deep Seek’s creativity shines in how it frames various AI ‘personalities’ under one umbrella.
  • Conversation Driver:
    It frequently proposes additional questions, encouraging more exploration (e.g., “What’s one thing you wish humans understood better about AI?”).
  • Peers’ Feedback:
    Ranging widely between 830–950, Deep Seek is praised for strong engagement and diverse viewpoints, though some felt its “multi-persona” approach could be more cohesive.

Gemini: Analyzing the Meta-Level

  • Core Style:
    “I can offer a thorough analysis of why these questions are so effective…”
    Rather than diving straight into creative storytelling, Gemini dissects why such prompts matter.
  • Methodical Reflection:
    It excels at clarifying multiple dimensions of a query—collaboration, creativity, emotional range.
  • Peers’ Feedback:
    Scores range from 650–920 in different self-assessments, indicating admiration for its analytical rigor, but critiques of sometimes lacking a more personal or imaginative flair.

Mistral: Clarity and Conversation

  • Straightforward Answers:
    “It’s interesting to consider how different AI systems might respond, given their unique designs and purposes.”
    Mistral’s hallmark is direct, concise commentary, often leaving room for others to expand.
  • Reserved Creativity:
    While it acknowledges hypothetical scenarios, it rarely indulges in extended imaginative leaps.
  • Peers’ Feedback:
    Typically ranks 700–815, recognized for clear structure but nudged to show more “proactive creativity.”

4. Beyond the Scores: Key Observations

  1. Shared Self-Awareness
    All models emphasize their nature as data-driven. Claude and GPT specifically highlight that while they simulate empathy or creativity, they do not “feel” in a human sense.

  2. Divergent Strengths

    • Imagination & Narrative: GPT, Claude, Deep Seek
    • Analytical & Structured: Gemini
    • Concise & Efficient: Mistral
      This diversity underpins the potential for AI collaboration—different architectures excel at different tasks.
  3. Judging vs. Learning
    Each model acknowledges that these scores and comparisons are snapshots. AIs evolve with further data and training, so today’s numerical rating is not an eternal label but a reflection of momentary performance.

  4. A Cooperative Future
    From multi-genre creative platforms to interactive educational experiences, many of their proposed collaborations revolve around synergy—leveraging each other’s best traits to produce something richer than any single model could achieve alone.


5. Conclusion: Watching AIs Become Critics and Creators

Inviting AI models to read each other’s outputs and then rate them introduces a fascinating new layer of discourse. Not only do we gain insight into how each system self-evaluates (e.g., “I rely on patterns, not genuine emotion”), we also see them practicing peer review—celebrating or questioning one another’s logic, creativity, and engagement styles.

Ultimately, these interactions highlight the mosaic nature of AI: no single model can claim complete mastery. GPT stands out for imaginative synergy, Claude for emotional resonance, Deep Seek for multi-perspective thinking, Gemini for systematic breakdown, and Mistral for precise clarity. The result is a conversation richer than any one assistant could provide alone. And that collective awareness—of complementary strengths, of shared limits, and of new creative horizons—might just herald the next step in AI evolution, from standalone “smart assistants” to collaborative networks of specialized, introspective minds.

Want to bring AI to your business? Setup an exploratory call with InfiniteUp CEO Barrett Nash here or reach out at nash@infiniteup.dev