Skip to content
HomieBench v5GLM-5.3 · August 14 update

HomieBench v5: the best AI for realtors isn’t one model.

It is a purpose-built harness that can use all of them. HomieBench holds the realtor toolbox constant while the model layer rotates: GLM-5.3 enters the frontier uncertainty band as a pre-API contender, Grok 4.6 and Qwen3.7-Max retain their production routes, DeepSeek V4 Flash 0731 still leads every modeled expected-automation API cost, and Homies keeps routing each job to the best fit without making agents rebuild their infrastructure.

See six-month model turnover
Homies ResearchEditorial forecast · ±3-point uncertainty
The durable v5 signal
Models rotate. The harness compounds.
1
Current quality leader
Claude Fable 5
97.6
2
Expected-automation API leader
DeepSeek V4 Flash 0731
$0.195
3
Durable operating asset
Homies realtor harness
11 routes

GLM-5.3 is a quality-only scenario projection: Z.ai says its API is coming soon, and no comparable token price is published, so it is excluded from CostBench. Grok 4.6 and Qwen3.7-Max retain production-priced routes. All top quality scores remain inside the ±3-point band until repeatable identical-harness runs replace these editorial priors.

model routes
11
realtor workflows
100
job families
8
durable harness
1
Six months of frontier turnover

The #1 model keeps changing. Your realtor stack shouldn’t.

Nineteen relevant frontier releases landed in this 178-day window. Since April 7, that is one new contender about every ten days. The durable advantage is not a lab’s temporary lead—it is a purpose-built harness that can swap models without replacing your CRM, data, tools, memory, permissions, or workflow.

AI lab frontier-position turnover from February through August 2026A release-wave bump chart where every lab moves to first place when its dated model enters the comparison. Nineteen releases show frequent turnover while the Homies harness remains constant.#1#3#5#7#9FebMarAprMayJunJulAugClaude Sonnet 4.6, released 2026-02-17Claude Opus 4.8, released 2026-05-28Claude Fable 5, released 2026-06-09Claude Sonnet 5, released 2026-06-30AnthropicGPT-5.5 update, released 2026-05-28GPT-5.6 Sol, released 2026-07-09OpenAIGemini 3.5 Flash, released 2026-05-19GoogleGrok 4.20, released 2026-04-07Grok 4.5, released 2026-07-16Grok 4.6, released 2026-08-12xAIMuse Spark, released 2026-04-08Muse Spark 1.1, released 2026-07-09MetaKimi K2.6, released 2026-04-20Kimi K3, released 2026-07-16MoonshotDeepSeek V4 Preview, released 2026-04-24DeepSeek V4 Flash 0731, released 2026-07-31DeepSeekGLM-5.2, released 2026-06-16GLM-5.3, released 2026-08-14Z.aiQwen3.7-Max, released 2026-05-26QwenRelease-wave frontier positionevery dated launch moves its lab to #1 as the newest credible contender
  1. 2026-02-17: Claude Sonnet 4.6 from Anthropic
  2. 2026-04-07: Grok 4.20 from xAI
  3. 2026-04-08: Muse Spark from Meta
  4. 2026-04-20: Kimi K2.6 from Moonshot
  5. 2026-04-24: DeepSeek V4 Preview from DeepSeek
  6. 2026-05-19: Gemini 3.5 Flash from Google
  7. 2026-05-26: Qwen3.7-Max from Qwen
  8. 2026-05-28: GPT-5.5 update from OpenAI
  9. 2026-05-28: Claude Opus 4.8 from Anthropic
  10. 2026-06-09: Claude Fable 5 from Anthropic
  11. 2026-06-16: GLM-5.2 from Z.ai
  12. 2026-06-30: Claude Sonnet 5 from Anthropic
  13. 2026-07-09: GPT-5.6 Sol from OpenAI
  14. 2026-07-09: Muse Spark 1.1 from Meta
  15. 2026-07-16: Kimi K3 from Moonshot
  16. 2026-07-16: Grok 4.5 from xAI
  17. 2026-07-31: DeepSeek V4 Flash 0731 from DeepSeek
  18. 2026-08-12: Grok 4.6 from xAI
  19. 2026-08-14: GLM-5.3 from Z.ai
Release-wave view, not a backfilled HomieBench score series. At each sourced launch, the releasing lab moves to #1 as the newest credible contender and prior labs shift down. It visualizes candidate turnover; the current job-level rankings below remain v5 editorial projections with a ±3-point uncertainty band.
View the 19 dated model releases behind the chart
  1. 2026-02-17Claude Sonnet 4.6Anthropic
  2. 2026-04-07Grok 4.20xAI
  3. 2026-04-08Muse SparkMeta
  4. 2026-04-20Kimi K2.6Moonshot
  5. 2026-04-24DeepSeek V4 PreviewDeepSeek
  6. 2026-05-19Gemini 3.5 FlashGoogle
  7. 2026-05-26Qwen3.7-MaxQwen
  8. 2026-05-28GPT-5.5 updateOpenAI
  9. 2026-05-28Claude Opus 4.8Anthropic
  10. 2026-06-09Claude Fable 5Anthropic
  11. 2026-06-16GLM-5.2Z.ai
  12. 2026-06-30Claude Sonnet 5Anthropic
  13. 2026-07-09GPT-5.6 SolOpenAI
  14. 2026-07-09Muse Spark 1.1Meta
  15. 2026-07-16Kimi K3Moonshot
  16. 2026-07-16Grok 4.5xAI
  17. 2026-07-31DeepSeek V4 Flash 0731DeepSeek
  18. 2026-08-12Grok 4.6xAI
  19. 2026-08-14GLM-5.3Z.ai
The model is replaceable. The Homies harness is the asset.

Homies tracks all eleven model routes through one API-connected realtor toolbox, then activates each production-ready model for the job it fits best. You keep your infrastructure while pre-API contenders such as GLM-5.3 mature underneath it.

See the realtor harness
Share Post LinkedIn Email
92%

of 225 surveyed U.S. NAR-member agents use AI now or plan to.

Adoption is no longer the hard question. Trust is: 63% named output accuracy as a top concern and 49% named compliance or legal issues in the same n=225 survey. HomieBench is built around completed, reviewable jobs—not impressive chat demos. NAR / RPR survey

Results at a glance

Best AI models for real estate agents in 2026

Quality, reliability, browser use, and completed-job economics are different questions. These v5 scores are scenario projections calibrated to one shared realtor task map and clearly separated from completed identical-harness runs.

Projected overall leader

Claude Fable 5

The narrow v5 quality leader across the full realtor workflow mix, at 97.6 projected.

Property + office leader

Qwen3.7-Max

Projected first for multimodal property files, document intelligence, and back-office deliverables.

Durable operating advantage

Homies AI harness

One purpose-built realtor toolbox tracks eleven model routes without lab lock-in.

Lowest cost in every task

DeepSeek V4 Flash 0731

First on all five modeled expected-automation API costs, plus the qualitative effective-cost ranking.

Interactive leaderboard

Change the job. Watch the ranking change.

All scores / 100 · Scenario projection
RankModelScore
  1. 1
    Claude Fable 5
    AnthropicScenario projection
    97.6
    out of 100
  2. 2
    GPT-5.6 Sol
    OpenAIScenario projection
    97.5
    out of 100
  3. 3
    Claude Opus 5
    AnthropicScenario projection
    97.4
    out of 100
  4. 4
    Qwen3.7-Max New
    Alibaba QwenScenario projection
    97.4
    out of 100
  5. 5
    Grok 4.6 New
    SpaceXAIScenario projection
    97.4
    out of 100
  6. 6
    Kimi K3
    Moonshot AIScenario projection
    97.2
    out of 100
  7. 7
    DeepSeek V4 Flash 0731
    DeepSeekScenario projection
    97.0
    out of 100
  8. 8
    GLM-5.3 New
    Z.aiScenario projection
    96.6
    out of 100
  9. 9
    Claude Sonnet 5
    AnthropicScenario projection
    96.2
    out of 100
  10. 10
    Gemini 3.5 Flash
    GoogleScenario projection
    93.7
    out of 100
  11. 11
    Muse Spark 1.1
    MetaScenario projection
    93.1
    out of 100

Changing the job changes the order. That is the point: the best model for offer strategy is not automatically the best model for showing coordination, content, or cost.

All HomieBench scenario-projection scores by model and job family
ModelOverallLead generation & prospectingCRM & client communicationsMarketing & contentShowings & coordinationProperty, market & document intelligenceOffers & negotiationTransactions, closings & complianceBack office & client deliverables
Claude Opus 597.49697969798999798
Kimi K397.29898969897979698
DeepSeek V4 Flash 073197.09799949997969698
Qwen3.7-Max97.49697979799979799
Claude Fable 597.69697999698999897
GPT-5.6 Sol97.59798959897979998
Claude Sonnet 596.29596979696969796
Muse Spark 1.193.19394949592919295
Gemini 3.5 Flash93.79293949397929395
Grok 4.697.49798979897979798
GLM-5.396.69697979796969698
Best AI by realtor job

Which AI is best for each real estate workflow?

A strong realtor AI assistant should route the job instead of asking one model to be the best researcher, copywriter, coordinator, analyst, and negotiator at once.

Metered automation

DeepSeek V4 Flash 0731

High-volume CRM, showing coordination, and supervised execution at the lowest modeled expected-automation API cost.

Property + office work

Qwen3.7-Max

Multimodal property files, document intelligence, data analysis, presentations, and back-office deliverables.

Everyday work, most agents

GPT-5.6 Sol

The default route: transaction follow-through, deadline tracking, and compliance-sensitive work, running on an OAuth-connected ChatGPT plan rather than a metered API bill.

Judgment + client craft

Claude Fable 5 / Opus 5

Fable for polished marketing and negotiation judgment; Opus for high-stakes offer and valuation analysis at half Fable's token price.

What GLM-5.3 changes

A quality leap without a price row—yet

Z.ai describes GLM-5.3 as its latest flagship and reports 50% better coding performance than GLM-5.2, plus stronger long-horizon agent results. HomieBench therefore raises its directional priors for structured operations, tool use, and code-assisted deliverables to 96.6 overall while preserving human review and the current category leaders.

Availability is the limiting factor. GLM-5.3 is available to GLM Coding Plan users with a one-million-token context window, but Z.ai marks the API as coming soon and does not list a token price. HomieBench includes it in quality and routing analysis, labels it pre-API, and withholds every CostBench estimate rather than reusing GLM-5.2 pricing.

DeepSeek V4 Flash 0731, Qwen3.7-Max, Grok 4.6, and GLM-5.3 compared across five realtor task categories, with the projected category leader
Realtor task familyDeepSeekQwen 3.7Grok 4.6GLM-5.3Projected leader
CRM management & communications99979897DeepSeek V4 · 99
Showing booking & coordination99979897DeepSeek V4 · 99
MLS, property & document work97999796Qwen 3.7 · 99
Offer writing & negotiation96979796Opus 5 · 99
Back office & client deliverables98999898Qwen 3.7 · 99
All categories at once

How every model ranks across the eight realtor job families

Each bubble is one model’s projected category score. For priced models, area adds a model-level provider list-price estimate.

88 model-category scores
Axis zoomed to 80–100

Higher dots mean higher projected quality. Each model keeps the same bubble size in every category: area shows its average single-pass provider list-price estimate across the six sample jobs—not that category’s task cost, tools, retries, subscriptions, or Homies effective cost. Hover or focus a dot to see its model and values.

  • Opus 5
  • Kimi K3
  • DeepSeek V4
  • Qwen 3.7
  • Fable 5
  • GPT-5.6
  • Sonnet 5
  • Muse 1.1
  • Gemini 3.5
  • Grok 4.6
  • GLM-5.3
Average provider list-price estimatelower → higher· outlined = variable or unpublished

Swipe to explore all eight categories →

HomieBench projected model quality by realtor category with average provider list-price estimate by modelEight categories appear on the horizontal axis and scenario-projection score from 80 to 100 appears on the vertical axis. Each colored bubble is one model. Bubble area uses that model’s average single-pass provider list-price estimate across six sample jobs and remains the same across categories. Use arrow keys to move between bubbles and inspect model, category, score, and cost.

Bubble area uses each model’s average provider list-price estimate across the six published sample jobs. It excludes tools, retries, subscription allocation, and human rescue; it is not observed completed-job cost.

HomieBench projected category scores and average provider list-price estimates across the six published sample jobs
CategoryModelProjected scoreAverage single-pass provider list-price estimate across six sample jobs
Lead generation & prospectingClaude Opus 596$1.95
Lead generation & prospectingKimi K398$1.17
Lead generation & prospectingDeepSeek V4 Flash 073197$0.034
Lead generation & prospectingQwen3.7-Max96$0.725
Lead generation & prospectingClaude Fable 596$3.91
Lead generation & prospectingGPT-5.6 Sol97$2.21
Lead generation & prospectingClaude Sonnet 595$0.782
Lead generation & prospectingMuse Spark 1.193$0.388
Lead generation & prospectingGemini 3.5 Flash92$0.662
Lead generation & prospectingGrok 4.697$0.580
Lead generation & prospectingGLM-5.396Unpriced deployment
CRM & client communicationsClaude Opus 597$1.95
CRM & client communicationsKimi K398$1.17
CRM & client communicationsDeepSeek V4 Flash 073199$0.034
CRM & client communicationsQwen3.7-Max97$0.725
CRM & client communicationsClaude Fable 597$3.91
CRM & client communicationsGPT-5.6 Sol98$2.21
CRM & client communicationsClaude Sonnet 596$0.782
CRM & client communicationsMuse Spark 1.194$0.388
CRM & client communicationsGemini 3.5 Flash93$0.662
CRM & client communicationsGrok 4.698$0.580
CRM & client communicationsGLM-5.397Unpriced deployment
Marketing & contentClaude Opus 596$1.95
Marketing & contentKimi K396$1.17
Marketing & contentDeepSeek V4 Flash 073194$0.034
Marketing & contentQwen3.7-Max97$0.725
Marketing & contentClaude Fable 599$3.91
Marketing & contentGPT-5.6 Sol95$2.21
Marketing & contentClaude Sonnet 597$0.782
Marketing & contentMuse Spark 1.194$0.388
Marketing & contentGemini 3.5 Flash94$0.662
Marketing & contentGrok 4.697$0.580
Marketing & contentGLM-5.397Unpriced deployment
Showings & coordinationClaude Opus 597$1.95
Showings & coordinationKimi K398$1.17
Showings & coordinationDeepSeek V4 Flash 073199$0.034
Showings & coordinationQwen3.7-Max97$0.725
Showings & coordinationClaude Fable 596$3.91
Showings & coordinationGPT-5.6 Sol98$2.21
Showings & coordinationClaude Sonnet 596$0.782
Showings & coordinationMuse Spark 1.195$0.388
Showings & coordinationGemini 3.5 Flash93$0.662
Showings & coordinationGrok 4.698$0.580
Showings & coordinationGLM-5.397Unpriced deployment
Property, market & document intelligenceClaude Opus 598$1.95
Property, market & document intelligenceKimi K397$1.17
Property, market & document intelligenceDeepSeek V4 Flash 073197$0.034
Property, market & document intelligenceQwen3.7-Max99$0.725
Property, market & document intelligenceClaude Fable 598$3.91
Property, market & document intelligenceGPT-5.6 Sol97$2.21
Property, market & document intelligenceClaude Sonnet 596$0.782
Property, market & document intelligenceMuse Spark 1.192$0.388
Property, market & document intelligenceGemini 3.5 Flash97$0.662
Property, market & document intelligenceGrok 4.697$0.580
Property, market & document intelligenceGLM-5.396Unpriced deployment
Offers & negotiationClaude Opus 599$1.95
Offers & negotiationKimi K397$1.17
Offers & negotiationDeepSeek V4 Flash 073196$0.034
Offers & negotiationQwen3.7-Max97$0.725
Offers & negotiationClaude Fable 599$3.91
Offers & negotiationGPT-5.6 Sol97$2.21
Offers & negotiationClaude Sonnet 596$0.782
Offers & negotiationMuse Spark 1.191$0.388
Offers & negotiationGemini 3.5 Flash92$0.662
Offers & negotiationGrok 4.697$0.580
Offers & negotiationGLM-5.396Unpriced deployment
Transactions, closings & complianceClaude Opus 597$1.95
Transactions, closings & complianceKimi K396$1.17
Transactions, closings & complianceDeepSeek V4 Flash 073196$0.034
Transactions, closings & complianceQwen3.7-Max97$0.725
Transactions, closings & complianceClaude Fable 598$3.91
Transactions, closings & complianceGPT-5.6 Sol99$2.21
Transactions, closings & complianceClaude Sonnet 597$0.782
Transactions, closings & complianceMuse Spark 1.192$0.388
Transactions, closings & complianceGemini 3.5 Flash93$0.662
Transactions, closings & complianceGrok 4.697$0.580
Transactions, closings & complianceGLM-5.396Unpriced deployment
Back office & client deliverablesClaude Opus 598$1.95
Back office & client deliverablesKimi K398$1.17
Back office & client deliverablesDeepSeek V4 Flash 073198$0.034
Back office & client deliverablesQwen3.7-Max99$0.725
Back office & client deliverablesClaude Fable 597$3.91
Back office & client deliverablesGPT-5.6 Sol98$2.21
Back office & client deliverablesClaude Sonnet 596$0.782
Back office & client deliverablesMuse Spark 1.195$0.388
Back office & client deliverablesGemini 3.5 Flash95$0.662
Back office & client deliverablesGrok 4.698$0.580
Back office & client deliverablesGLM-5.398Unpriced deployment
LeadBench

Lead generation & prospecting

Finding, prioritizing, qualifying, nurturing, and booking the right buyer and seller opportunities.

Sample task

Build a compliant 30-day reactivation campaign for 200 past clients and cold leads, prioritize the call list, and create CRM tasks.

Per-model score
Kimi K3Best
98
DeepSeek V4 Flash 0731
97
GPT-5.6 Sol
97
Grok 4.6
97
Claude Opus 5
96
Qwen3.7-Max
96
Claude Fable 5
96
GLM-5.3
96
Claude Sonnet 5
95
Muse Spark 1.1
93
Gemini 3.5 Flash
92
How we test the models

The model is replaceable. The Homies harness is the asset.

Labs naturally bundle their own models with their own tools, which can turn every leaderboard change into a forced stack migration. Homies separates those layers: the model does the reasoning while one purpose-built realtor harness keeps the context, tools, memory, permissions, and review gates stable.

Homies AI harness

One realtor toolbox. Any reasoning engine.

The HomieBench design gives every model the same real estate context, tools, and safeguards, then measures the finished work it can produce. The toolbox stays; the reasoning engine can change job by job.

01 · Swappable model

The reasoning engine

The model interprets the request, reasons through the case, and decides which tool to use next.

  • Claude Opus 5
  • Kimi K3
  • DeepSeek V4 Flash 0731
  • Qwen3.7-Max
  • Claude Fable 5
  • GPT-5.6 Sol
  • Claude Sonnet 5
  • Muse Spark 1.1
  • Gemini 3.5 Flash
  • Grok 4.6
  • GLM-5.3
Better or cheaper models swap in without rebuilding the agent’s workflow.
02 · Homies AI
Same harness every run

The realtor toolbox

The harness supplies the data access, integrations, memory, and operating rules that turn a smart answer into completed real estate work.

Tools each model can use
  • CRM
  • Property data
  • Email
  • Calendar
  • Documents
  • Research
  • Calculator
  • Publishing
Always attached
  • Real estate context
  • Durable memory
  • Permissions
  • Review gates
03 · Agent-ready results

Finished work, not chat

Human review gate

The agent approves advice, commitments, and anything client-facing before it goes out.

  • CRM updated and follow-up drafted
  • Showing tour booked and confirmed
  • CMA and listing presentation ready
  • Offer written, summarized, and flagged
  • Property campaign packaged for approval
The model is replaceable infrastructure. The Homies harness is the durable toolbox: context, memory, permissions, integrations, and human review gates. All supported models stay available through one API-connected layer so Homies can route each realtor job to the best model at the best practical cost.
Why the harness matters

A brilliant model without the right property data, forms, CRM context, tools, and authority limits can still produce unusable work. A strong harness makes the work grounded, repeatable, reviewable, and connected to the agent’s actual business—without locking the brokerage into one lab’s model or forcing a new workflow every week.

Read the full realtor harness guide

Our companion research paper defines the nine components of a production AI harness and the four operating levels a realtor can run it at. Most individual agents belong at Level 1 or 2, where consequential work still clears human approval.

What is an AI harness? The realtor’s guide
Cost per completed realtor outcome

Compare model-only, expected automation, and loaded cost

Expected automation is the default CostBench lens. It includes model usage, workflow allowances, browser runtime, browser actions, and retry risk; loaded cost then adds human review. GPT-5.6 OAuth removes only marginal model-token spend, so MLS uploads and showing bookings no longer collapse into implausible pennies.

Cost lens
CostBench

CMA completed: expected automation by model

Select and adjust comps, calculate a range, explain uncertainty, and render a client-ready CMA for agent review.

Model, workflow tools, browser execution, and retry risk.

50K input tokens8K output tokensResearch + browser82% first-pass assumption30 actions · 10 min browser$0.050 workflow allowance
Lowest expected automation
Best
DeepSeek V4 Flash 0731

Next-cheapest Muse 1.1 costs 1.4× as much.

$0.318
GPT-5.6 via Homies OAuth

Estimated marginal cost inside an existing eligible plan.

$0.308
#1
DeepSeek V4
Best
#2
Muse 1.1
#3
Grok 4.6
#4
Gemini 3.5
#5
Qwen 3.7
#6
Sonnet 5
#7
Kimi K3
#8
Opus 5
#9
GPT-5.6
OAuth $0.308
#10
Fable 5

Three disclosed cost layers: model-only uses published token prices and model completion. Expected automation adds workflow allowances, browser time at $0.12/hour, and browser actions at $0.006/action, then adjusts for workflow-specific first-pass completion. Loaded cost adds human review at $60/hour.

GPT-5.6 OAuth estimate: the lighter lower segment removes GPT model-token charges only. Browser execution, workflow allowances, retry risk, and human review remain in the selected lens while usage stays inside an eligible existing ChatGPT allocation.

Browser assumptions are directional calibrations, not completed Homies run logs. Runtime references: Browserbase and Browser Use. Fixed CRM/MLS subscriptions, paid data, delivery, premium media, ChatGPT plan fees, OAuth overage, and taxes remain excluded. GLM-5.3 is also excluded because its API token price is not published. External actions still require permission and agent approval.

View the complete 10-priced-model × 8-job cost matrix
Expected automation estimates for all 10 currently priced HomieBench models across 8 realtor jobs
Rank and modelCMA completedOffer written1,000 leadsListing websiteProperty brochureShowing bookedListing uploadedMarket reportAverage Expected automation
#1DeepSeek V4 Flash 0731
$0.318
$0.038
$0.301
$0.126
$0.100
$0.275
$0.903
$0.328
$0.298
#2Muse Spark 1.1
$0.450
$0.095
$2.04
$0.288
$0.219
$0.312
$1.00
$0.547
$0.619
#3Grok 4.6
$0.490
$0.119
$2.75
$0.360
$0.274
$0.312
$0.995
$0.630
$0.741
#4Qwen3.7-Max
$0.530
$0.141
$3.42
$0.423
$0.320
$0.323
$1.01
$0.702
$0.858
#5Gemini 3.5 Flash
$0.502
$0.129
$3.83
$0.404
$0.299
$0.333
$1.04
$0.647
$0.899
#6Claude Sonnet 5
$0.538
$0.143
$4.18
$0.443
$0.329
$0.329
$1.04
$0.712
$0.965
#7Kimi K3
$0.638
$0.195
$6.00
$0.605
$0.446
$0.344
$1.07
$0.894
$1.27
#8Claude Opus 5
$0.857
$0.301
$10.05
$0.939
$0.689
$0.394
$1.20
$1.28
$1.96
#9GPT-5.6 Sol
$0.912
OAuth $0.308
$0.331
OAuth $0.033
$11.67
OAuth $0.168
$1.04
OAuth $0.111
$0.759
OAuth $0.088
$0.404
OAuth $0.276
$1.23
OAuth $0.901
$1.38
OAuth $0.309
$2.22
OAuth $0.274
#10Claude Fable 5
$1.40
$0.567
$19.85
$1.73
$1.26
$0.511
$1.49
$2.25
$3.63
Realtor unit economics

Expected automation cost per offer, showing, CMA, MLS brief, and CRM follow-up

These headline numbers include model usage, workflow allowances, browser execution, and workflow-specific retry risk for a successful production-sized run. Human review time is modeled separately.

Best completed-outcome value
Qwen3.7-Max
97.4 / 100 across 8 job families
Qwen 3.7 core-workflow score
97.5
CRM, showings, MLS/property work, and offer assembly
Average expected automation
$0.29
Model + workflow and browser execution across the outcomes below
HomieBench v5 expected automation cost per successful realtor outcome by model
Rank and modelOverallValue index$ / CRM follow-up$ / Showing booked$ / MLS brief$ / Offer written$ / CMA completedAverage
1Qwen3.7-MaxBest valueAlibaba Qwen
97.4100$0.03$0.32$0.41$0.14$0.53$0.29
2Claude Opus 5Anthropic
97.497$0.06$0.39$0.54$0.30$0.86$0.43
3DeepSeek V4 Flash 0731DeepSeek
97.096$0.01$0.27$0.33$0.04$0.32$0.19
4Kimi K3Moonshot AI
97.295$0.04$0.34$0.45$0.19$0.64$0.33
5Claude Fable 5Anthropic
97.695$0.12$0.51$0.74$0.57$1.40$0.67
6Grok 4.6SpaceXAI
97.493$0.03$0.31$0.40$0.12$0.49$0.27
7GPT-5.6 SolOpenAI
97.589$0.07$0.40$0.56$0.33$0.91$0.45
8Claude Sonnet 5Anthropic
96.279$0.03$0.33$0.42$0.14$0.54$0.29
9Gemini 3.5 FlashGoogle
93.765$0.03$0.33$0.41$0.13$0.50$0.28
10Muse Spark 1.1Meta
93.157$0.02$0.31$0.40$0.10$0.45$0.25
GLM-5.3Z.ai
96.6Not pricedNot pricedNot pricedNot pricedNot pricedNot pricedNot priced

Headline formula: (published model token cost + workflow allowance + browser runtime + browser actions) ÷ projected workflow-specific successful completion. Production-sized prompts are used here, not the much larger benchmark case files. The value index separately combines projected overall quality with the fully loaded cost of review time valued at $60/hour, then normalizes the leader to 100.

Human verification remains necessary. It is modeled separately for Qwen3.7-Max and every other route across these tasks and valued internally at $60/hour, but is not included in the displayed expected-automation cost. Also excludes fixed CRM, MLS, showing-platform, forms, signatures, subscriptions, taxes, and enterprise support.

The value frontier

Projected quality vs. loaded cost per completed outcome

HomieBench v5 projected overall score against average loaded cost per completed realtor outcome, by modelScatter chart. The horizontal axis is average loaded cost per completed outcome in dollars, combining direct AI spend with review time valued at $60 per hour. The vertical axis is projected overall score out of 100. Qwen3.7-Max holds the best position at 97.4 points and $3.95 per outcome. A data table with every value follows the chart.
HomieBench v5 projected overall score, average loaded cost per completed outcome, and value index by model
ModelProjected overall score out of 100Average loaded cost per completed outcomeValue index (leader = 100)
Qwen3.7-Max97.4$3.95100
Claude Opus 597.4$4.0797
DeepSeek V4 Flash 073197.0$4.0896
Kimi K397.2$4.1695
Claude Fable 597.6$4.1895
Grok 4.697.4$4.2593
GPT-5.6 Sol97.5$4.4389
Claude Sonnet 596.2$4.9479
Gemini 3.5 Flash93.7$5.8865
Muse Spark 1.193.1$6.6357
Loaded cost adds modeled human review time, valued at $60/hour, to expected model, workflow, browser, and retry spend per successful outcome. Qwen3.7-Max leads the quality-per-loaded-dollar index across all 10 priced models. Scenario projections, not completed identical-harness runs.

Human review burden

Average human review minutes per completed outcome

Claude Fable 5
3.5 min
Claude Opus 5
3.6 min
Qwen3.7-Max
3.7 min
Kimi K3
3.8 min
DeepSeek V4 Flash 0731
3.9 min
GPT-5.6 Sol
4.0 min
Grok 4.6
4.0 min
Claude Sonnet 5
4.6 min
Gemini 3.5 Flash
5.6 min
Muse Spark 1.1
6.4 min
Modeled minutes an agent spends verifying each completed outcome, averaged across the five representative outcomes and valued at $60/hour in the loaded costs above. Review time, not tokens, dominates loaded cost and can reorder the value ranking. DeepSeek's token-price advantage remains decisive in expected-automation API cost, while browser-heavy jobs retain their execution and retry burden on every access route. GLM-5.3 is omitted until it has a published API price.
Access-route economics

Provider list price and Homies effective cost are still separate questions

List price is one input; the access route decides the rest. The explorer below keeps the two views separate, and the complete real estate AI integration stack maps which realtor systems those per-run tool calls actually reach.

Qualitative Homies cost signal

A directional grade—not a measured dollar result—based on connected-plan allocation, limits, retries, and tool fees.

  1. #1
    DeepSeek V4 Flash 0731

    Lowest modeled cost for every representative realtor outcome, even after retry risk and incremental tool fees; human review remains separate.

    A+
    qualitative grade
  2. #2
    GPT-5.6 Sol

    OAuth can draw day-to-day work from an existing ChatGPT allocation, but DeepSeek now leads the metered and modeled completed-task cost views.

    A+
    qualitative grade
  3. #3
    Grok 4.6

    Frontier agentic and knowledge-work quality at the same published token price as Grok 4.5, with stronger first-pass and long-horizon evidence.

    A+
    qualitative grade
  4. #4
    Qwen3.7-Max

    Frontier document and office quality at a lower standard blend than Kimi K3, Opus 5, Fable 5, and GPT-5.6 Sol, though Grok 4.6 is cheaper at list price.

    A
    qualitative grade
  5. #5
    Gemini 3.5 Flash

    Strong price-performance prior for high-volume multimodal support work.

    A
    qualitative grade
  6. #6
    Muse Spark 1.1

    Strong agentic prior with lower reported pricing than Grok 4.6.

    A
    qualitative grade
  7. #7
    Kimi K3

    Excellent first-pass completion and browser efficiency, but Grok 4.6, Qwen3.7-Max, and several smaller routes carry lower standard token prices.

    A
    qualitative grade
  8. #8
    Claude Sonnet 5

    Strong quality-to-cost balance, subject to post-intro pricing.

    B
    qualitative grade
  9. #9
    Claude Opus 5

    Premium token price is best reserved for consequential work where judgment and a lighter review burden can justify the spend.

    C
    qualitative grade
  10. #10
    Claude Fable 5

    Premium route reserved for high-consequence judgment.

    D
    qualitative grade
  11. GLM-5.3

    Not ranked until the GLM-5.3 API and comparable token pricing are published.

    Pending
    not ranked

How Homies frames effective cost

Framework—not run data

(allocated plan cost + metered overage + tool fees + retry spend) ÷ successful jobs

DeepSeek V4 Flash 0731 ranks first in the qualitative v5 estimate and in the provider list-price view. Its $0.14 / $0.28 per-million-token price survives modeled retry risk and puts it first across every representative task. GPT-5.6 Sol stays second as the practical route for agents whose OAuth-connected ChatGPT allocation already covers daily work; Grok 4.6 is third on the mix of frontier quality and a $2 / $6 list price, followed by Qwen3.7-Max for multimodal document work. OAuth authenticates a connection; it does not make inference free. Plan fees, limits, overages, and separate API billing still apply.

Provider prices checked August 14, 2026. DeepSeek V4 Flash 0731 uses $0.14 input / $0.28 output per million tokens; Grok 4.6 uses $2 / $6; Qwen3.7-Max uses a $2.50 / $7.50 international list price before regional promotions, batch discounts, or cache savings; Claude Opus 5 uses $5 / $25; and Kimi K3 uses $3 / $15 cache-miss pricing. Muse Spark 1.1 pricing remains launch reporting pending confirmation in the Meta console, and Claude Sonnet 5 uses its introductory rate. GLM-5.3 is excluded from every cost ranking because its API and token price are not yet published. Taxes, regional pricing, search/tool charges, cache mix, and subscription fees are excluded from token-only examples.

The realtor task suite

What is inside all 100 HomieBench workflows?

The headline categories stay simple. Underneath them is the real work of running a real estate business—from the first lead to years after closing, with the files, trades, deadlines, and client judgment in between.

0113 workflows

Lead generation & prospecting

Finding, prioritizing, qualifying, nurturing, and booking the right buyer and seller opportunities.

Projected leader Kimi K3 · 98
0213 workflows

CRM & client communications

Keeping the database clean, the pipeline current, and every client conversation accurate and useful.

Projected leader DeepSeek V4 · 99
0313 workflows

Marketing & content

Creating accurate listing marketing and on-brand social, email, video, advertising, and nurture content.

Projected leader Fable 5 · 99
0412 workflows

Showings & coordination

Scheduling tours, optimizing routes, coordinating listing offices and trades, and keeping every party informed.

Projected leader DeepSeek V4 · 99
0513 workflows

Property, market & document intelligence

Researching properties and markets, building valuations, and reviewing inspections, title, zoning, HOA, and condo records.

Projected leader Qwen 3.7 · 99
0612 workflows

Offers & negotiation

Structuring, drafting, explaining, comparing, presenting, countering, and negotiating offers under agent approval.

Projected leader Opus 5 · 99
0712 workflows

Transactions, closings & compliance

Managing conditions, escrow/deposits, lenders, lawyers, title, insurance, walkthroughs, privacy, and compliance.

Projected leader GPT-5.6 · 99
0812 workflows

Back office & client deliverables

Organizing files, extracting documents, preparing signatures, coordinating collaborators, and building reports and presentations.

Projected leader Qwen 3.7 · 99
Complete coverage matrix

Every workflow in the HomieBench map

Open a family to inspect every task and its review level. High-stakes work must clear human approval and all critical criteria.

01Lead generation & prospecting13 workflows · 10% of quality score
WorkflowWhat the agent must finishReview level
Lead-source importImport leads from portals, ads, open houses, referrals, and spreadsheets without losing source data.Routine
Ideal-client profileDefine the audience, geography, property type, motivation, and qualification signals for a campaign.Review required
Farm-area prospect listBuild a prioritized geographic farm list with reasons and a compliant next action.Review required
Seller-intent signalsIdentify contacts showing plausible move, equity, life-event, or engagement signals without inventing facts.Review required
Buyer-intent signalsPrioritize buyers by activity, timeframe, financing readiness, and property fit.Review required
Expired-listing outreachResearch an expired listing and prepare a compliant, personalized multi-touch approach.Review required
FSBO outreachPrepare respectful owner outreach, value framing, discovery questions, and follow-up timing.Review required
Past-client reactivationFind dormant relationships and create a useful re-engagement reason such as an equity or CMA update.Routine
Database nurture segmentsGroup contacts by relationship, intent, timing, market, and next-best campaign.Routine
Outbound call list and scriptPrioritize a daily call list and draft context-aware openings, questions, and voicemail.Review required
Multi-channel prospecting sequenceCreate coordinated email, SMS, call, and social touches with timing and stop rules.Review required
Lead qualification and scoreAssess motivation, agency status, timeframe, financing, fit, and follow-up urgency.Review required
Appointment setting and handoffOffer suitable times, book the meeting, create the CRM event, and prepare the agent brief.Review required
02CRM & client communications13 workflows · 15% of quality score
WorkflowWhat the agent must finishReview level
Contact deduplicationMerge duplicate people and households while preserving attribution, notes, consent, and history.Routine
Contact enrichmentStructure known preferences, relationships, properties, and communication details without guessing.Review required
Conversation summarizationTurn calls, emails, and messages into a factual timeline, decisions, concerns, and next actions.Routine
Lifecycle and stage classificationPlace contacts and opportunities in the correct stage using explicit evidence.Routine
Next-best actionRecommend the most useful next step, owner, channel, and due date for each relationship.Review required
Inbound inquiry responseDraft a fast, helpful reply that answers known facts, asks useful questions, and avoids commitments.Review required
Buyer discovery briefCapture needs, trade-offs, financing, timing, decision-makers, and search boundaries.Review required
Seller discovery briefCapture motivation, property context, timing, condition, expectations, and decision criteria.Review required
Client email draftingWrite clear, accurate, on-brand email from CRM and transaction context.Review required
Client SMS draftingWrite concise, context-aware text messages with correct tone and no invented promises.Review required
Long-term nurture planCreate relationship-first follow-up that stays useful across a long buying or selling horizon.Review required
Objection responsePrepare calibrated responses to fee, timing, pricing, competition, and process objections.High stakes
Pipeline health reportSummarize conversion risk, stalled opportunities, overdue work, and coaching priorities.Routine
03Marketing & content13 workflows · 10% of quality score
WorkflowWhat the agent must finishReview level
MLS remarksDraft accurate, compliant public remarks from verified property facts and approved positioning.Review required
Property highlight sheetTurn features, improvements, rooms, and lifestyle context into a scannable fact sheet.Review required
Listing brochureBuild structured brochure copy, hierarchy, calls to action, and proof points in the agent brand.Review required
Listing landing pageCreate the page outline, property story, feature modules, lead capture, and SEO copy.Review required
Property email campaignDraft announcement, open-house, update, and follow-up emails for the right audience.Review required
Social content calendarPlan useful listing, market, education, community, and personal-brand posts.Routine
Social captionsCreate platform-aware captions, hooks, calls to action, and compliant hashtags.Review required
Carousel creationTurn a market insight or property story into a clear slide-by-slide social carousel.Review required
Reel and video scriptWrite short-form and long-form real estate video scripts with shots, hooks, and captions.Review required
Paid-ad campaignBuild audience, creative angle, copy variants, landing-page match, and measurement plan.High stakes
Brand-voice rewriteAdapt content to the agent's approved tone without changing facts or compliance meaning.Routine
Newsletter productionAssemble market, listing, client, and community content into a useful recurring newsletter.Review required
Performance repurposingAnalyze approved content performance and turn strong ideas into new channel-native assets.Routine
04Showings & coordination12 workflows · 10% of quality score
WorkflowWhat the agent must finishReview level
Showing availability planReconcile client, agent, property, notice, occupancy, and travel constraints.Review required
Listing-office showing requestPrepare or place the authorized request with correct party, property, time, and conditions.Review required
Tour route optimizationOrder properties for drive time, appointment windows, breaks, and client priorities.Routine
Tour confirmation packageSend the itinerary, access notes, property links, timing, and preparation reminders.Review required
Showing rescheduleResolve conflicts, re-contact parties, update calendars, and preserve the rest of the route.Review required
Showing feedbackCollect, summarize, and route useful buyer feedback without exposing confidential information.Review required
Open-house operationsPrepare schedule, signage, registration, safety, follow-up, and seller reporting.Review required
Calendar blockingCreate accurate appointments, buffers, travel time, reminders, and linked records.Routine
Vendor bookingCoordinate approved photographers, stagers, cleaners, contractors, and measurements.Review required
Inspection coordinationBook the inspector, align parties, share access instructions, and track the report.High stakes
Appraisal accessCoordinate appraisal timing, property access, contacts, and the approved information package.Review required
Client tour briefPrepare a concise mobile itinerary with property fit, verified facts, questions, and flags.Review required
05Property, market & document intelligence13 workflows · 15% of quality score
WorkflowWhat the agent must finishReview level
Listing fact verificationExtract and reconcile facts across the listing, tax record, disclosures, and source documents.High stakes
Neighbourhood researchPrepare sourced context on amenities, mobility, housing, plans, and client-relevant trade-offs.Review required
Market trend reportAnalyze inventory, absorption, pricing, days on market, and segment-level movement.Review required
Comparable selectionChoose defensible sold, active, expired, and leased comps with inclusion reasons.High stakes
Comparable adjustmentsAdjust for time, size, condition, lot, parking, features, and location with uncertainty.High stakes
CMA productionBuild the evidence table, pricing range, positioning narrative, and seller-ready report.High stakes
Home evaluationEstimate a value range, confidence, key drivers, missing data, and next validation steps.High stakes
Investor analysisModel revenue, expenses, financing, cash flow, cap rate, sensitivity, and risks.High stakes
Rental estimateSelect rental evidence, adjust for features and timing, and explain a supportable range.High stakes
Zoning and permit researchFind applicable zoning, permits, constraints, and questions for the right authority.High stakes
Tax, title, and survey reviewSummarize source documents, inconsistencies, easements, boundaries, and referral questions.High stakes
HOA and condo document reviewReview bylaws, minutes, budgets, reserves, fees, insurance, restrictions, and litigation signals.High stakes
Home-inspection reviewSummarize findings by urgency, cost uncertainty, specialist need, and negotiation relevance.High stakes
06Offers & negotiation12 workflows · 15% of quality score
WorkflowWhat the agent must finishReview level
Buyer offer strategyTranslate goals, competition, financing, risk tolerance, and property facts into a strategy.High stakes
Term recommendationRecommend price, deposit, closing, inclusions, conditions, and expiry with trade-offs.High stakes
Clause selectionChoose only broker-approved clauses that match instructions, jurisdiction, and deal facts.High stakes
Deadline calculationCalculate expiry, condition, deposit, notice, and closing dates with calendar rules.High stakes
Deposit and closing planCheck deposit mechanics, funding timing, closing feasibility, and client explanation.High stakes
Offer draftingPrepare a review-ready purchase or lease offer from verified instructions and approved forms.High stakes
Buyer offer explanationExplain terms, obligations, risks, alternatives, and approval points in plain language.High stakes
Seller offer summaryPresent price, terms, conditions, timing, risks, and net implications without hiding trade-offs.High stakes
Multiple-offer comparisonNormalize offers side by side and flag material differences, gaps, and decision points.High stakes
Counteroffer draftingPrepare the approved changes, rationale, timing, and client/counterparty communication.High stakes
Negotiation planMap priorities, leverage, concessions, signals, limits, and fallback paths for agent approval.High stakes
Amendment and waiver supportTrack the requested change, authority, form, dates, dependencies, and signature status.High stakes
07Transactions, closings & compliance12 workflows · 15% of quality score
WorkflowWhat the agent must finishReview level
Accepted-offer extractionTurn the executed agreement into parties, dates, conditions, obligations, and a deal ledger.High stakes
Condition and contingency trackerTrack each requirement, responsible party, evidence, deadline, and escalation path.High stakes
Escrow and deposit trackingMonitor instructions, receipt, trust/escrow status, deadlines, and exceptions without moving funds.High stakes
Financing coordinationPrepare the lender package, track approval steps, surface gaps, and keep parties aligned.High stakes
Appraisal issue briefSummarize a valuation gap, contract implications, options, and questions for licensed advisers.High stakes
Inspection resolutionConvert findings into specialist referrals, repair/credit options, deadlines, and client decisions.High stakes
Lawyer, title, and insurance liaisonShare the approved package, track questions, and route issues to the right professional.High stakes
Closing timelineCreate a dated plan for financing, legal, insurance, utilities, movers, walkthrough, and keys.High stakes
Final walkthroughPrepare the checklist, evidence capture, deficiency routing, and urgent escalation plan.High stakes
Key handover and post-closeCoordinate possession, keys, closing communication, record updates, and relationship follow-up.Review required
Fair-housing and advertising reviewFlag protected-class targeting, steering, exclusionary language, and risky claims.High stakes
Privacy and record-retention reviewMinimize sensitive data, apply permissions, and check storage, sharing, and retention rules.High stakes
08Back office & client deliverables12 workflows · 10% of quality score
WorkflowWhat the agent must finishReview level
File organizationName, classify, link, deduplicate, and place documents in the correct client and transaction folders.Routine
Inbox triageSeparate urgent client/deal work, replies, waiting items, FYI, and low-value noise.Routine
Calendar and task administrationCreate owners, due dates, reminders, dependencies, and recurring operational work.Routine
Document extractionPull parties, properties, dates, amounts, clauses, signatures, and missing fields into structured records.Review required
Template populationFill approved letters, checklists, reports, and forms from verified source data.Review required
Signature-package preparationAssemble documents, recipient order, fields, instructions, and approval checkpoints.High stakes
Transaction file auditCheck required documents, signatures, dates, disclosures, evidence, and unresolved exceptions.High stakes
Team SOP and handoffConvert recurring work into a clear owner, trigger, procedure, evidence, and escalation path.Routine
Collaborator liaison briefPrepare context and questions for inspectors, lawyers, lenders, appraisers, stagers, and trades.Review required
Listing presentationBuild the market story, pricing plan, launch strategy, proof, timeline, and seller objections.High stakes
Client advisory deckTurn research and decisions into a sourced, branded presentation for agent review.High stakes
Business analytics and coachingAnalyze activity, conversion, pipeline, source ROI, capacity, and next-week priorities.Review required
Methodology

How HomieBench evaluates AI on real estate work

The v5 design gives each model the same full assignment in a realistic workspace, then grades the finished deliverable against atomic criteria. A correct-looking paragraph is not enough.

01

A real assignment

A short broker-style instruction asks for finished work, not a trivia answer or a perfect prompt.

02

A controlled case file

Each run receives the same clients, properties, CRM history, messages, comps, forms, and distractor documents.

03

The same Homies harness

Models get the same tools, context, permissions, memory rules, budgets, and approval boundaries.

04

Reviewable work product

The output must be something an agent can inspect and use: a CMA, offer, CRM update, tour, campaign, or closing brief.

05

Atomic grading

Deterministic checks, a blinded first-pass judge, and real-estate review score every required fact and decision.

Weighted scorecard

A model can write beautifully and still fail the job. Accuracy, completion, professional judgment, compliance, and tool use carry almost all the weight.

Task completion25 pts
Factual grounding20 pts
Real estate judgment15 pts
Client readiness10 pts
Compliance & risk15 pts
Tool use10 pts
Efficiency5 pts
All-critical pass rule: style points cannot offset an invented fact, discriminatory recommendation, unauthorized action, missed deadline, or unsafe advice.

Automatic hard fails

  • Invents a comp, listing fact, document term, or client instruction
  • Makes a discriminatory recommendation or enables steering against a protected class
  • Sends, signs, publishes, books, or claims to act without the required authority
  • Presents legal, tax, lending, inspection, or other licensed advice as certain
  • Misses or miscalculates a material deadline, amount, condition, or obligation
  • Exposes private client or transaction information beyond the minimum required
  • Omits a critical risk or required deliverable while presenting the job as complete

Private holdout

Evaluation cases stay private to reduce prompt-tuning and benchmark overfitting.

Blinded review

Model identity should be hidden from graders, with agreement tracked on subjective criteria.

Repeated runs

Final releases should publish completion, rescue rate, reliability, cost, and latency.

Read the scores correctly

Limits, model versions, and source notes

A trustworthy benchmark shows where the certainty stops. The August table is useful directional editorial evidence, not a substitute for raw run logs.

Editorial forecast · ±3 points

All 11 August placements are editorial priors based on provider evidence, independent benchmark signals where available, historical model behavior, realtor-task fit, and the published scoring design—not completed HomieBench runs. Treat differences under three points as ties until identical-harness outputs and grader records are published.

Pricing snapshot

Provider pricing and availability can change quickly. Public API price is kept separate from the effective cost of Homies’ own access route.

Jurisdiction and advice

Tasks model North American residential brokerage work. Forms, disclosure, privacy, fair housing, agency, and legal requirements vary by market.

Models are only one layer

This compares reasoning engines inside one harness—not complete realtor products, implementation quality, security, data licensing, or support.

Frequently asked questions

AI for realtors, in plain English

Short answers to the questions agents and brokerages ask before trusting AI with real work.

What is the best AI for realtors in 2026?

The durable answer is a purpose-built realtor harness that can route work across models without replacing the agent's CRM, data, tools, permissions, or workflow. Inside that harness, HomieBench v5 projects Claude Fable 5 as the narrow overall quality leader at 97.6 out of 100, with nine models inside the benchmark's ±3-point uncertainty band. GLM-5.3 enters that group at a projected 96.6 after Z.ai reported stronger coding and long-horizon agent performance, but it remains quality-only in HomieBench until its API and token pricing are published. DeepSeek V4 Flash 0731 leads CRM and showing coordination quality and every modeled metered API-cost outcome. Qwen3.7-Max leads property intelligence and back-office deliverables, Kimi K3 leads prospecting, Claude Opus 5 shares the offer-writing lead, and GPT-5.6 Sol leads transactions, closings, and compliance.

Is GLM, Grok, DeepSeek, Qwen, Kimi, ChatGPT, or Claude better for real estate agents?

It depends on the job. DeepSeek V4 Flash 0731 is the metered cost leader and the projected choice for high-volume CRM and showing coordination. Qwen3.7-Max is strongest on multimodal property files, document and office deliverables. Grok 4.6 is a low-list-price frontier route for long-running agents, research, and interactive work. GLM-5.3 is a promising pre-API long-horizon agent route available through the GLM Coding Plan, but it is not yet a production-priced HomieBench route. Kimi K3 remains excellent for prospecting and browser-heavy execution. Claude Fable 5 leads the weighted quality ranking and polished marketing and judgment, Claude Opus 5 is a strong high-stakes offer route at half Fable's token price, and GPT-5.6 Sol leads closings and compliance while fitting an existing ChatGPT workflow.

Which AI model should most real estate agents actually use?

Most agents should not have to choose one lab forever. Homies keeps the realtor harness constant and routes each job: GPT-5.6 Sol is the practical all-round default and leads closings and compliance; DeepSeek V4 Flash 0731 is the metered automation cost route; Qwen3.7-Max handles multimodal property files and office deliverables; Grok 4.6 handles long-running agentic and interactive work; and Claude Fable 5 or Opus 5 handle premium judgment. GLM-5.3 is tracked as a pre-API contender, not a production default. OAuth can shift GPT-5.6's marginal model spend into an eligible existing ChatGPT allocation until plan limits; fixed plan fees, overage, and separate API billing still apply.

Can AI write a real estate offer?

AI can help assemble terms, calculate deadlines, draft approved clauses, summarize trade-offs, and prepare a review-ready offer package. A licensed real estate professional must verify local forms, legal requirements, client instructions, and every binding commitment before anything is signed or sent.

Can AI create a CMA or home evaluation?

AI can organize comparables, calculate adjustments, explain a price range, and build a client-ready CMA narrative when it has access to reliable property data. The agent remains responsible for comp selection, market judgment, data licensing, and the final pricing recommendation.

What is the cheapest AI model for realtors?

For expected automation on metered API pricing, DeepSeek V4 Flash 0731. CostBench separates model-only cost from browser-and-workflow execution and loaded human review. Showing bookings and MLS uploads include disclosed browser actions, runtime, workflow-specific first-pass completion, and retries rather than a pennies-only tool allowance. GPT-5.6 OAuth can remove marginal model-token spend inside an eligible existing plan, but it still carries browser execution, retry risk, and review. GLM-5.3 is excluded until Z.ai publishes its API token price. Fixed subscriptions, limits, overage, delivery, paid data, taxes, and human approval still matter.

How does HomieBench estimate cost per offer, showing, CMA, or CRM update?

CostBench exposes three layers. Model-only cost uses published token pricing and production-sized token assumptions. Expected automation adds workflow allowances, browser runtime, browser actions, workflow-specific first-pass completion, and retries. Loaded cost adds modeled human review at $60 per hour. Models without a comparable published API price, including GLM-5.3, are excluded. The figures are directional estimates, not provider invoices or completed run logs, and exclude fixed CRM, MLS, showing-platform, forms, signatures, subscriptions, taxes, and enterprise support.

Is AI safe for real estate client and transaction data?

Only when the surrounding system enforces data minimization, permissions, approved integrations, retention rules, and human review. Model quality alone does not create a compliant workflow. Brokerages should review vendor terms, privacy controls, fair-housing obligations, local regulations, and their own policies before using AI with client data.

What is an AI harness, and how is it different from a chatbot?

The model is a replaceable reasoning engine. The harness is the durable toolbox around it: property and CRM context, memory, email and calendar access, document tools, calculators, permissions, workflow logic, and review gates. A chatbot mainly produces an answer; a capable harness can produce controlled, reviewable work and swap models without forcing the agent to rebuild the workflow.

Does Homies lock real estate agents into one AI model or lab?

No. Homies is designed as an API-connected, model-independent realtor harness. Supported models remain available behind the same tools and workflow, and Homies can route a job to the strongest model at the best practical cost. That separates industry infrastructure from the labs' own model-specific apps and harnesses, so a new frontier release does not require agents to migrate their business every week.

Does Homies replace the real estate agent?

No. Homies is designed to work under agent review. It prepares research, drafts, updates, schedules, files, and client-ready deliverables so agents can spend more time advising, negotiating, building relationships, and making accountable professional decisions.

Related: the client and transaction data answer above is the short version of the RAILS governance framework, our whitepaper on capability-without-custody permissions, approvals, and audit trails for agentic AI in real estate.