Skip to content

Add Baseline Agents and Benchmark Configurations #482

Description

@AnkitaNaik

Title
Integrate baseline agents and benchmarking configurations

Description
Summary: Expand benchmarking beyond CUGA to support baseline comparisons.

Scope

Support

  • Agents
    -- CUGA
    -- LangGraph ReAct
    -- AppWorld CodeAct

  • Models
    -- GPT-OSS-120B
    -- GPT-OSS-20B
    -- Mistral Medium 3.5
    -- Gemma 4 31B
    -- Nematron 3

Acceptance Criteria

  • Multiple agents supported
  • Multiple models configurable
  • Baseline comparisons generated

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    Status
    Backlog

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions