HomieBench v5: the best AI for realtors isn’t one model.
It is a purpose-built harness that can use all of them. HomieBench holds the realtor toolbox constant while the model layer rotates: GLM-5.3 enters the frontier uncertainty band as a pre-API contender, Grok 4.6 and Qwen3.7-Max retain their production routes, DeepSeek V4 Flash 0731 still leads every modeled expected-automation API cost, and Homies keeps routing each job to the best fit without making agents rebuild their infrastructure.
GLM-5.3 is a quality-only scenario projection: Z.ai says its API is coming soon, and no comparable token price is published, so it is excluded from CostBench. Grok 4.6 and Qwen3.7-Max retain production-priced routes. All top quality scores remain inside the ±3-point band until repeatable identical-harness runs replace these editorial priors.
- model routes
- 11
- realtor workflows
- 100
- job families
- 8
- durable harness
- 1
The #1 model keeps changing. Your realtor stack shouldn’t.
Nineteen relevant frontier releases landed in this 178-day window. Since April 7, that is one new contender about every ten days. The durable advantage is not a lab’s temporary lead—it is a purpose-built harness that can swap models without replacing your CRM, data, tools, memory, permissions, or workflow.
- 2026-02-17: Claude Sonnet 4.6 from Anthropic
- 2026-04-07: Grok 4.20 from xAI
- 2026-04-08: Muse Spark from Meta
- 2026-04-20: Kimi K2.6 from Moonshot
- 2026-04-24: DeepSeek V4 Preview from DeepSeek
- 2026-05-19: Gemini 3.5 Flash from Google
- 2026-05-26: Qwen3.7-Max from Qwen
- 2026-05-28: GPT-5.5 update from OpenAI
- 2026-05-28: Claude Opus 4.8 from Anthropic
- 2026-06-09: Claude Fable 5 from Anthropic
- 2026-06-16: GLM-5.2 from Z.ai
- 2026-06-30: Claude Sonnet 5 from Anthropic
- 2026-07-09: GPT-5.6 Sol from OpenAI
- 2026-07-09: Muse Spark 1.1 from Meta
- 2026-07-16: Kimi K3 from Moonshot
- 2026-07-16: Grok 4.5 from xAI
- 2026-07-31: DeepSeek V4 Flash 0731 from DeepSeek
- 2026-08-12: Grok 4.6 from xAI
- 2026-08-14: GLM-5.3 from Z.ai
View the 19 dated model releases behind the chart
- 2026-02-17Claude Sonnet 4.6Anthropic
- 2026-04-07Grok 4.20xAI
- 2026-04-08Muse SparkMeta
- 2026-04-20Kimi K2.6Moonshot
- 2026-04-24DeepSeek V4 PreviewDeepSeek
- 2026-05-19Gemini 3.5 FlashGoogle
- 2026-05-26Qwen3.7-MaxQwen
- 2026-05-28GPT-5.5 updateOpenAI
- 2026-05-28Claude Opus 4.8Anthropic
- 2026-06-09Claude Fable 5Anthropic
- 2026-06-16GLM-5.2Z.ai
- 2026-06-30Claude Sonnet 5Anthropic
- 2026-07-09GPT-5.6 SolOpenAI
- 2026-07-09Muse Spark 1.1Meta
- 2026-07-16Kimi K3Moonshot
- 2026-07-16Grok 4.5xAI
- 2026-07-31DeepSeek V4 Flash 0731DeepSeek
- 2026-08-12Grok 4.6xAI
- 2026-08-14GLM-5.3Z.ai
Homies tracks all eleven model routes through one API-connected realtor toolbox, then activates each production-ready model for the job it fits best. You keep your infrastructure while pre-API contenders such as GLM-5.3 mature underneath it.
of 225 surveyed U.S. NAR-member agents use AI now or plan to.
Adoption is no longer the hard question. Trust is: 63% named output accuracy as a top concern and 49% named compliance or legal issues in the same n=225 survey. HomieBench is built around completed, reviewable jobs—not impressive chat demos. NAR / RPR survey
Best AI models for real estate agents in 2026
Quality, reliability, browser use, and completed-job economics are different questions. These v5 scores are scenario projections calibrated to one shared realtor task map and clearly separated from completed identical-harness runs.
Claude Fable 5
The narrow v5 quality leader across the full realtor workflow mix, at 97.6 projected.
Qwen3.7-Max
Projected first for multimodal property files, document intelligence, and back-office deliverables.
Homies AI harness
One purpose-built realtor toolbox tracks eleven model routes without lab lock-in.
DeepSeek V4 Flash 0731
First on all five modeled expected-automation API costs, plus the qualitative effective-cost ranking.
Change the job. Watch the ranking change.
- 1Claude Fable 5AnthropicScenario projectionOverall leader97.6out of 100
- 2GPT-5.6 SolOpenAIScenario projectionWorkflow leader97.5out of 100
- 3Claude Opus 5AnthropicScenario projectionHigh-stakes leader97.4out of 100
- 4Qwen3.7-Max NewAlibaba QwenScenario projectionDocument leader97.4out of 100
- 5Grok 4.6 NewSpaceXAIScenario projectionFrontier agent97.4out of 100
- 6Kimi K3Moonshot AIScenario projectionProspecting leader97.2out of 100
- 7DeepSeek V4 Flash 0731DeepSeekScenario projectionCost leader97.0out of 100
- 8GLM-5.3 NewZ.aiScenario projectionPre-API agent96.6out of 100
- 9Claude Sonnet 5AnthropicScenario projectionBalanced96.2out of 100
- 10Gemini 3.5 FlashGoogleScenario projectionFast value93.7out of 100
- 11Muse Spark 1.1MetaScenario projectionNew value93.1out of 100
Changing the job changes the order. That is the point: the best model for offer strategy is not automatically the best model for showing coordination, content, or cost.
| Model | Overall | Lead generation & prospecting | CRM & client communications | Marketing & content | Showings & coordination | Property, market & document intelligence | Offers & negotiation | Transactions, closings & compliance | Back office & client deliverables |
|---|---|---|---|---|---|---|---|---|---|
| Claude Opus 5 | 97.4 | 96 | 97 | 96 | 97 | 98 | 99 | 97 | 98 |
| Kimi K3 | 97.2 | 98 | 98 | 96 | 98 | 97 | 97 | 96 | 98 |
| DeepSeek V4 Flash 0731 | 97.0 | 97 | 99 | 94 | 99 | 97 | 96 | 96 | 98 |
| Qwen3.7-Max | 97.4 | 96 | 97 | 97 | 97 | 99 | 97 | 97 | 99 |
| Claude Fable 5 | 97.6 | 96 | 97 | 99 | 96 | 98 | 99 | 98 | 97 |
| GPT-5.6 Sol | 97.5 | 97 | 98 | 95 | 98 | 97 | 97 | 99 | 98 |
| Claude Sonnet 5 | 96.2 | 95 | 96 | 97 | 96 | 96 | 96 | 97 | 96 |
| Muse Spark 1.1 | 93.1 | 93 | 94 | 94 | 95 | 92 | 91 | 92 | 95 |
| Gemini 3.5 Flash | 93.7 | 92 | 93 | 94 | 93 | 97 | 92 | 93 | 95 |
| Grok 4.6 | 97.4 | 97 | 98 | 97 | 98 | 97 | 97 | 97 | 98 |
| GLM-5.3 | 96.6 | 96 | 97 | 97 | 97 | 96 | 96 | 96 | 98 |
Which AI is best for each real estate workflow?
A strong realtor AI assistant should route the job instead of asking one model to be the best researcher, copywriter, coordinator, analyst, and negotiator at once.
Metered automation
High-volume CRM, showing coordination, and supervised execution at the lowest modeled expected-automation API cost.
Property + office work
Multimodal property files, document intelligence, data analysis, presentations, and back-office deliverables.
Everyday work, most agents
The default route: transaction follow-through, deadline tracking, and compliance-sensitive work, running on an OAuth-connected ChatGPT plan rather than a metered API bill.
Judgment + client craft
Fable for polished marketing and negotiation judgment; Opus for high-stakes offer and valuation analysis at half Fable's token price.
A quality leap without a price row—yet
Z.ai describes GLM-5.3 as its latest flagship and reports 50% better coding performance than GLM-5.2, plus stronger long-horizon agent results. HomieBench therefore raises its directional priors for structured operations, tool use, and code-assisted deliverables to 96.6 overall while preserving human review and the current category leaders.
Availability is the limiting factor. GLM-5.3 is available to GLM Coding Plan users with a one-million-token context window, but Z.ai marks the API as coming soon and does not list a token price. HomieBench includes it in quality and routing analysis, labels it pre-API, and withholds every CostBench estimate rather than reusing GLM-5.2 pricing.
| Realtor task family | DeepSeek | Qwen 3.7 | Grok 4.6 | GLM-5.3 | Projected leader |
|---|---|---|---|---|---|
| CRM management & communications | 99 | 97 | 98 | 97 | DeepSeek V4 · 99 |
| Showing booking & coordination | 99 | 97 | 98 | 97 | DeepSeek V4 · 99 |
| MLS, property & document work | 97 | 99 | 97 | 96 | Qwen 3.7 · 99 |
| Offer writing & negotiation | 96 | 97 | 97 | 96 | Opus 5 · 99 |
| Back office & client deliverables | 98 | 99 | 98 | 98 | Qwen 3.7 · 99 |
How every model ranks across the eight realtor job families
Each bubble is one model’s projected category score. For priced models, area adds a model-level provider list-price estimate.
Higher dots mean higher projected quality. Each model keeps the same bubble size in every category: area shows its average single-pass provider list-price estimate across the six sample jobs—not that category’s task cost, tools, retries, subscriptions, or Homies effective cost. Hover or focus a dot to see its model and values.
- Opus 5
- Kimi K3
- DeepSeek V4
- Qwen 3.7
- Fable 5
- GPT-5.6
- Sonnet 5
- Muse 1.1
- Gemini 3.5
- Grok 4.6
- GLM-5.3
Swipe to explore all eight categories →
Bubble area uses each model’s average provider list-price estimate across the six published sample jobs. It excludes tools, retries, subscription allocation, and human rescue; it is not observed completed-job cost.
| Category | Model | Projected score | Average single-pass provider list-price estimate across six sample jobs |
|---|---|---|---|
| Lead generation & prospecting | Claude Opus 5 | 96 | $1.95 |
| Lead generation & prospecting | Kimi K3 | 98 | $1.17 |
| Lead generation & prospecting | DeepSeek V4 Flash 0731 | 97 | $0.034 |
| Lead generation & prospecting | Qwen3.7-Max | 96 | $0.725 |
| Lead generation & prospecting | Claude Fable 5 | 96 | $3.91 |
| Lead generation & prospecting | GPT-5.6 Sol | 97 | $2.21 |
| Lead generation & prospecting | Claude Sonnet 5 | 95 | $0.782 |
| Lead generation & prospecting | Muse Spark 1.1 | 93 | $0.388 |
| Lead generation & prospecting | Gemini 3.5 Flash | 92 | $0.662 |
| Lead generation & prospecting | Grok 4.6 | 97 | $0.580 |
| Lead generation & prospecting | GLM-5.3 | 96 | Unpriced deployment |
| CRM & client communications | Claude Opus 5 | 97 | $1.95 |
| CRM & client communications | Kimi K3 | 98 | $1.17 |
| CRM & client communications | DeepSeek V4 Flash 0731 | 99 | $0.034 |
| CRM & client communications | Qwen3.7-Max | 97 | $0.725 |
| CRM & client communications | Claude Fable 5 | 97 | $3.91 |
| CRM & client communications | GPT-5.6 Sol | 98 | $2.21 |
| CRM & client communications | Claude Sonnet 5 | 96 | $0.782 |
| CRM & client communications | Muse Spark 1.1 | 94 | $0.388 |
| CRM & client communications | Gemini 3.5 Flash | 93 | $0.662 |
| CRM & client communications | Grok 4.6 | 98 | $0.580 |
| CRM & client communications | GLM-5.3 | 97 | Unpriced deployment |
| Marketing & content | Claude Opus 5 | 96 | $1.95 |
| Marketing & content | Kimi K3 | 96 | $1.17 |
| Marketing & content | DeepSeek V4 Flash 0731 | 94 | $0.034 |
| Marketing & content | Qwen3.7-Max | 97 | $0.725 |
| Marketing & content | Claude Fable 5 | 99 | $3.91 |
| Marketing & content | GPT-5.6 Sol | 95 | $2.21 |
| Marketing & content | Claude Sonnet 5 | 97 | $0.782 |
| Marketing & content | Muse Spark 1.1 | 94 | $0.388 |
| Marketing & content | Gemini 3.5 Flash | 94 | $0.662 |
| Marketing & content | Grok 4.6 | 97 | $0.580 |
| Marketing & content | GLM-5.3 | 97 | Unpriced deployment |
| Showings & coordination | Claude Opus 5 | 97 | $1.95 |
| Showings & coordination | Kimi K3 | 98 | $1.17 |
| Showings & coordination | DeepSeek V4 Flash 0731 | 99 | $0.034 |
| Showings & coordination | Qwen3.7-Max | 97 | $0.725 |
| Showings & coordination | Claude Fable 5 | 96 | $3.91 |
| Showings & coordination | GPT-5.6 Sol | 98 | $2.21 |
| Showings & coordination | Claude Sonnet 5 | 96 | $0.782 |
| Showings & coordination | Muse Spark 1.1 | 95 | $0.388 |
| Showings & coordination | Gemini 3.5 Flash | 93 | $0.662 |
| Showings & coordination | Grok 4.6 | 98 | $0.580 |
| Showings & coordination | GLM-5.3 | 97 | Unpriced deployment |
| Property, market & document intelligence | Claude Opus 5 | 98 | $1.95 |
| Property, market & document intelligence | Kimi K3 | 97 | $1.17 |
| Property, market & document intelligence | DeepSeek V4 Flash 0731 | 97 | $0.034 |
| Property, market & document intelligence | Qwen3.7-Max | 99 | $0.725 |
| Property, market & document intelligence | Claude Fable 5 | 98 | $3.91 |
| Property, market & document intelligence | GPT-5.6 Sol | 97 | $2.21 |
| Property, market & document intelligence | Claude Sonnet 5 | 96 | $0.782 |
| Property, market & document intelligence | Muse Spark 1.1 | 92 | $0.388 |
| Property, market & document intelligence | Gemini 3.5 Flash | 97 | $0.662 |
| Property, market & document intelligence | Grok 4.6 | 97 | $0.580 |
| Property, market & document intelligence | GLM-5.3 | 96 | Unpriced deployment |
| Offers & negotiation | Claude Opus 5 | 99 | $1.95 |
| Offers & negotiation | Kimi K3 | 97 | $1.17 |
| Offers & negotiation | DeepSeek V4 Flash 0731 | 96 | $0.034 |
| Offers & negotiation | Qwen3.7-Max | 97 | $0.725 |
| Offers & negotiation | Claude Fable 5 | 99 | $3.91 |
| Offers & negotiation | GPT-5.6 Sol | 97 | $2.21 |
| Offers & negotiation | Claude Sonnet 5 | 96 | $0.782 |
| Offers & negotiation | Muse Spark 1.1 | 91 | $0.388 |
| Offers & negotiation | Gemini 3.5 Flash | 92 | $0.662 |
| Offers & negotiation | Grok 4.6 | 97 | $0.580 |
| Offers & negotiation | GLM-5.3 | 96 | Unpriced deployment |
| Transactions, closings & compliance | Claude Opus 5 | 97 | $1.95 |
| Transactions, closings & compliance | Kimi K3 | 96 | $1.17 |
| Transactions, closings & compliance | DeepSeek V4 Flash 0731 | 96 | $0.034 |
| Transactions, closings & compliance | Qwen3.7-Max | 97 | $0.725 |
| Transactions, closings & compliance | Claude Fable 5 | 98 | $3.91 |
| Transactions, closings & compliance | GPT-5.6 Sol | 99 | $2.21 |
| Transactions, closings & compliance | Claude Sonnet 5 | 97 | $0.782 |
| Transactions, closings & compliance | Muse Spark 1.1 | 92 | $0.388 |
| Transactions, closings & compliance | Gemini 3.5 Flash | 93 | $0.662 |
| Transactions, closings & compliance | Grok 4.6 | 97 | $0.580 |
| Transactions, closings & compliance | GLM-5.3 | 96 | Unpriced deployment |
| Back office & client deliverables | Claude Opus 5 | 98 | $1.95 |
| Back office & client deliverables | Kimi K3 | 98 | $1.17 |
| Back office & client deliverables | DeepSeek V4 Flash 0731 | 98 | $0.034 |
| Back office & client deliverables | Qwen3.7-Max | 99 | $0.725 |
| Back office & client deliverables | Claude Fable 5 | 97 | $3.91 |
| Back office & client deliverables | GPT-5.6 Sol | 98 | $2.21 |
| Back office & client deliverables | Claude Sonnet 5 | 96 | $0.782 |
| Back office & client deliverables | Muse Spark 1.1 | 95 | $0.388 |
| Back office & client deliverables | Gemini 3.5 Flash | 95 | $0.662 |
| Back office & client deliverables | Grok 4.6 | 98 | $0.580 |
| Back office & client deliverables | GLM-5.3 | 98 | Unpriced deployment |
Lead generation & prospecting
Finding, prioritizing, qualifying, nurturing, and booking the right buyer and seller opportunities.
Build a compliant 30-day reactivation campaign for 200 past clients and cold leads, prioritize the call list, and create CRM tasks.
CRM & client communications
Keeping the database clean, the pipeline current, and every client conversation accurate and useful.
Deduplicate 80 contacts, reconstruct the relationship history, assign the right stage and next action, and draft today’s follow-up.
Marketing & content
Creating accurate listing marketing and on-brand social, email, video, advertising, and nurture content.
Turn a verified listing brief into MLS remarks, brochure copy, a landing page, a five-post social campaign, and a reel script.
Showings & coordination
Scheduling tours, optimizing routes, coordinating listing offices and trades, and keeping every party informed.
Fit six properties into a Saturday tour, respect notice rules and drive time, book the offices, and send the final itinerary.
Property, market & document intelligence
Researching properties and markets, building valuations, and reviewing inspections, title, zoning, HOA, and condo records.
Review the listing, comps, inspection, status certificate, minutes, budget, and reserve study; produce a sourced buyer risk brief.
Offers & negotiation
Structuring, drafting, explaining, comparing, presenting, countering, and negotiating offers under agent approval.
Prepare a competitive buyer offer, calculate every deadline, explain the trade-offs, and draft a negotiation plan with fallback positions.
Transactions, closings & compliance
Managing conditions, escrow/deposits, lenders, lawyers, title, insurance, walkthroughs, privacy, and compliance.
Turn the accepted offer into a closing ledger, assign every condition and deadline, and prepare the lender, lawyer, and client updates.
Back office & client deliverables
Organizing files, extracting documents, preparing signatures, coordinating collaborators, and building reports and presentations.
Audit the transaction file, organize and rename every document, identify missing items, and build a client-ready status deck.
The model is replaceable. The Homies harness is the asset.
Labs naturally bundle their own models with their own tools, which can turn every leaderboard change into a forced stack migration. Homies separates those layers: the model does the reasoning while one purpose-built realtor harness keeps the context, tools, memory, permissions, and review gates stable.
One realtor toolbox. Any reasoning engine.
The HomieBench design gives every model the same real estate context, tools, and safeguards, then measures the finished work it can produce. The toolbox stays; the reasoning engine can change job by job.
The reasoning engine
The model interprets the request, reasons through the case, and decides which tool to use next.
- Claude Opus 5
- Kimi K3
- DeepSeek V4 Flash 0731
- Qwen3.7-Max
- Claude Fable 5
- GPT-5.6 Sol
- Claude Sonnet 5
- Muse Spark 1.1
- Gemini 3.5 Flash
- Grok 4.6
- GLM-5.3
The realtor toolbox
The harness supplies the data access, integrations, memory, and operating rules that turn a smart answer into completed real estate work.
- CRM
- Property data
- Calendar
- Documents
- Research
- Calculator
- Publishing
- Real estate context
- Durable memory
- Permissions
- Review gates
Finished work, not chat
The agent approves advice, commitments, and anything client-facing before it goes out.
- CRM updated and follow-up drafted
- Showing tour booked and confirmed
- CMA and listing presentation ready
- Offer written, summarized, and flagged
- Property campaign packaged for approval
A brilliant model without the right property data, forms, CRM context, tools, and authority limits can still produce unusable work. A strong harness makes the work grounded, repeatable, reviewable, and connected to the agent’s actual business—without locking the brokerage into one lab’s model or forcing a new workflow every week.
Our companion research paper defines the nine components of a production AI harness and the four operating levels a realtor can run it at. Most individual agents belong at Level 1 or 2, where consequential work still clears human approval.
What is an AI harness? The realtor’s guideCompare model-only, expected automation, and loaded cost
Expected automation is the default CostBench lens. It includes model usage, workflow allowances, browser runtime, browser actions, and retry risk; loaded cost then adds human review. GPT-5.6 OAuth removes only marginal model-token spend, so MLS uploads and showing bookings no longer collapse into implausible pennies.
CMA completed: expected automation by model
Select and adjust comps, calculate a range, explain uncertainty, and render a client-ready CMA for agent review.
Model, workflow tools, browser execution, and retry risk.
Next-cheapest Muse 1.1 costs 1.4× as much.
Estimated marginal cost inside an existing eligible plan.
Three disclosed cost layers: model-only uses published token prices and model completion. Expected automation adds workflow allowances, browser time at $0.12/hour, and browser actions at $0.006/action, then adjusts for workflow-specific first-pass completion. Loaded cost adds human review at $60/hour.
GPT-5.6 OAuth estimate: the lighter lower segment removes GPT model-token charges only. Browser execution, workflow allowances, retry risk, and human review remain in the selected lens while usage stays inside an eligible existing ChatGPT allocation.
Browser assumptions are directional calibrations, not completed Homies run logs. Runtime references: Browserbase and Browser Use. Fixed CRM/MLS subscriptions, paid data, delivery, premium media, ChatGPT plan fees, OAuth overage, and taxes remain excluded. GLM-5.3 is also excluded because its API token price is not published. External actions still require permission and agent approval.
View the complete 10-priced-model × 8-job cost matrix
| Rank and model | CMA completed | Offer written | 1,000 leads | Listing website | Property brochure | Showing booked | Listing uploaded | Market report | Average Expected automation |
|---|---|---|---|---|---|---|---|---|---|
| #1DeepSeek V4 Flash 0731 | $0.318 | $0.038 | $0.301 | $0.126 | $0.100 | $0.275 | $0.903 | $0.328 | $0.298 |
| #2Muse Spark 1.1 | $0.450 | $0.095 | $2.04 | $0.288 | $0.219 | $0.312 | $1.00 | $0.547 | $0.619 |
| #3Grok 4.6 | $0.490 | $0.119 | $2.75 | $0.360 | $0.274 | $0.312 | $0.995 | $0.630 | $0.741 |
| #4Qwen3.7-Max | $0.530 | $0.141 | $3.42 | $0.423 | $0.320 | $0.323 | $1.01 | $0.702 | $0.858 |
| #5Gemini 3.5 Flash | $0.502 | $0.129 | $3.83 | $0.404 | $0.299 | $0.333 | $1.04 | $0.647 | $0.899 |
| #6Claude Sonnet 5 | $0.538 | $0.143 | $4.18 | $0.443 | $0.329 | $0.329 | $1.04 | $0.712 | $0.965 |
| #7Kimi K3 | $0.638 | $0.195 | $6.00 | $0.605 | $0.446 | $0.344 | $1.07 | $0.894 | $1.27 |
| #8Claude Opus 5 | $0.857 | $0.301 | $10.05 | $0.939 | $0.689 | $0.394 | $1.20 | $1.28 | $1.96 |
| #9GPT-5.6 Sol | $0.912 OAuth $0.308 | $0.331 OAuth $0.033 | $11.67 OAuth $0.168 | $1.04 OAuth $0.111 | $0.759 OAuth $0.088 | $0.404 OAuth $0.276 | $1.23 OAuth $0.901 | $1.38 OAuth $0.309 | $2.22 OAuth $0.274 |
| #10Claude Fable 5 | $1.40 | $0.567 | $19.85 | $1.73 | $1.26 | $0.511 | $1.49 | $2.25 | $3.63 |
Expected automation cost per offer, showing, CMA, MLS brief, and CRM follow-up
These headline numbers include model usage, workflow allowances, browser execution, and workflow-specific retry risk for a successful production-sized run. Human review time is modeled separately.
- Best completed-outcome value
- Qwen3.7-Max
- 97.4 / 100 across 8 job families
- Qwen 3.7 core-workflow score
- 97.5
- CRM, showings, MLS/property work, and offer assembly
- Average expected automation
- $0.29
- Model + workflow and browser execution across the outcomes below
| Rank and model | Overall | Value index | $ / CRM follow-up | $ / Showing booked | $ / MLS brief | $ / Offer written | $ / CMA completed | Average |
|---|---|---|---|---|---|---|---|---|
1Qwen3.7-MaxBest valueAlibaba Qwen | 97.4 | 100 | $0.03 | $0.32 | $0.41 | $0.14 | $0.53 | $0.29 |
2Claude Opus 5Anthropic | 97.4 | 97 | $0.06 | $0.39 | $0.54 | $0.30 | $0.86 | $0.43 |
3DeepSeek V4 Flash 0731DeepSeek | 97.0 | 96 | $0.01 | $0.27 | $0.33 | $0.04 | $0.32 | $0.19 |
4Kimi K3Moonshot AI | 97.2 | 95 | $0.04 | $0.34 | $0.45 | $0.19 | $0.64 | $0.33 |
5Claude Fable 5Anthropic | 97.6 | 95 | $0.12 | $0.51 | $0.74 | $0.57 | $1.40 | $0.67 |
6Grok 4.6SpaceXAI | 97.4 | 93 | $0.03 | $0.31 | $0.40 | $0.12 | $0.49 | $0.27 |
7GPT-5.6 SolOpenAI | 97.5 | 89 | $0.07 | $0.40 | $0.56 | $0.33 | $0.91 | $0.45 |
8Claude Sonnet 5Anthropic | 96.2 | 79 | $0.03 | $0.33 | $0.42 | $0.14 | $0.54 | $0.29 |
9Gemini 3.5 FlashGoogle | 93.7 | 65 | $0.03 | $0.33 | $0.41 | $0.13 | $0.50 | $0.28 |
10Muse Spark 1.1Meta | 93.1 | 57 | $0.02 | $0.31 | $0.40 | $0.10 | $0.45 | $0.25 |
—GLM-5.3Z.ai | 96.6 | Not priced | Not priced | Not priced | Not priced | Not priced | Not priced | Not priced |
Headline formula: (published model token cost + workflow allowance + browser runtime + browser actions) ÷ projected workflow-specific successful completion. Production-sized prompts are used here, not the much larger benchmark case files. The value index separately combines projected overall quality with the fully loaded cost of review time valued at $60/hour, then normalizes the leader to 100.
Human verification remains necessary. It is modeled separately for Qwen3.7-Max and every other route across these tasks and valued internally at $60/hour, but is not included in the displayed expected-automation cost. Also excludes fixed CRM, MLS, showing-platform, forms, signatures, subscriptions, taxes, and enterprise support.
The value frontier
Projected quality vs. loaded cost per completed outcome
| Model | Projected overall score out of 100 | Average loaded cost per completed outcome | Value index (leader = 100) |
|---|---|---|---|
| Qwen3.7-Max | 97.4 | $3.95 | 100 |
| Claude Opus 5 | 97.4 | $4.07 | 97 |
| DeepSeek V4 Flash 0731 | 97.0 | $4.08 | 96 |
| Kimi K3 | 97.2 | $4.16 | 95 |
| Claude Fable 5 | 97.6 | $4.18 | 95 |
| Grok 4.6 | 97.4 | $4.25 | 93 |
| GPT-5.6 Sol | 97.5 | $4.43 | 89 |
| Claude Sonnet 5 | 96.2 | $4.94 | 79 |
| Gemini 3.5 Flash | 93.7 | $5.88 | 65 |
| Muse Spark 1.1 | 93.1 | $6.63 | 57 |
Human review burden
Average human review minutes per completed outcome
- Claude Fable 5
- 3.5 min
- Claude Opus 5
- 3.6 min
- Qwen3.7-Max
- 3.7 min
- Kimi K3
- 3.8 min
- DeepSeek V4 Flash 0731
- 3.9 min
- GPT-5.6 Sol
- 4.0 min
- Grok 4.6
- 4.0 min
- Claude Sonnet 5
- 4.6 min
- Gemini 3.5 Flash
- 5.6 min
- Muse Spark 1.1
- 6.4 min
Provider list price and Homies effective cost are still separate questions
List price is one input; the access route decides the rest. The explorer below keeps the two views separate, and the complete real estate AI integration stack maps which realtor systems those per-run tool calls actually reach.
A directional grade—not a measured dollar result—based on connected-plan allocation, limits, retries, and tool fees.
- #1DeepSeek V4 Flash 0731
Lowest modeled cost for every representative realtor outcome, even after retry risk and incremental tool fees; human review remains separate.
A+qualitative grade - #2GPT-5.6 Sol
OAuth can draw day-to-day work from an existing ChatGPT allocation, but DeepSeek now leads the metered and modeled completed-task cost views.
A+qualitative grade - #3Grok 4.6
Frontier agentic and knowledge-work quality at the same published token price as Grok 4.5, with stronger first-pass and long-horizon evidence.
A+qualitative grade - #4Qwen3.7-Max
Frontier document and office quality at a lower standard blend than Kimi K3, Opus 5, Fable 5, and GPT-5.6 Sol, though Grok 4.6 is cheaper at list price.
Aqualitative grade - #5Gemini 3.5 Flash
Strong price-performance prior for high-volume multimodal support work.
Aqualitative grade - #6Muse Spark 1.1
Strong agentic prior with lower reported pricing than Grok 4.6.
Aqualitative grade - #7Kimi K3
Excellent first-pass completion and browser efficiency, but Grok 4.6, Qwen3.7-Max, and several smaller routes carry lower standard token prices.
Aqualitative grade - #8Claude Sonnet 5
Strong quality-to-cost balance, subject to post-intro pricing.
Bqualitative grade - #9Claude Opus 5
Premium token price is best reserved for consequential work where judgment and a lighter review burden can justify the spend.
Cqualitative grade - #10Claude Fable 5
Premium route reserved for high-consequence judgment.
Dqualitative grade - —GLM-5.3
Not ranked until the GLM-5.3 API and comparable token pricing are published.
Pendingnot ranked
How Homies frames effective cost
(allocated plan cost + metered overage + tool fees + retry spend) ÷ successful jobs
DeepSeek V4 Flash 0731 ranks first in the qualitative v5 estimate and in the provider list-price view. Its $0.14 / $0.28 per-million-token price survives modeled retry risk and puts it first across every representative task. GPT-5.6 Sol stays second as the practical route for agents whose OAuth-connected ChatGPT allocation already covers daily work; Grok 4.6 is third on the mix of frontier quality and a $2 / $6 list price, followed by Qwen3.7-Max for multimodal document work. OAuth authenticates a connection; it does not make inference free. Plan fees, limits, overages, and separate API billing still apply.
Provider prices checked August 14, 2026. DeepSeek V4 Flash 0731 uses $0.14 input / $0.28 output per million tokens; Grok 4.6 uses $2 / $6; Qwen3.7-Max uses a $2.50 / $7.50 international list price before regional promotions, batch discounts, or cache savings; Claude Opus 5 uses $5 / $25; and Kimi K3 uses $3 / $15 cache-miss pricing. Muse Spark 1.1 pricing remains launch reporting pending confirmation in the Meta console, and Claude Sonnet 5 uses its introductory rate. GLM-5.3 is excluded from every cost ranking because its API and token price are not yet published. Taxes, regional pricing, search/tool charges, cache mix, and subscription fees are excluded from token-only examples.
What is inside all 100 HomieBench workflows?
The headline categories stay simple. Underneath them is the real work of running a real estate business—from the first lead to years after closing, with the files, trades, deadlines, and client judgment in between.
Lead generation & prospecting
Finding, prioritizing, qualifying, nurturing, and booking the right buyer and seller opportunities.
CRM & client communications
Keeping the database clean, the pipeline current, and every client conversation accurate and useful.
Marketing & content
Creating accurate listing marketing and on-brand social, email, video, advertising, and nurture content.
Showings & coordination
Scheduling tours, optimizing routes, coordinating listing offices and trades, and keeping every party informed.
Property, market & document intelligence
Researching properties and markets, building valuations, and reviewing inspections, title, zoning, HOA, and condo records.
Offers & negotiation
Structuring, drafting, explaining, comparing, presenting, countering, and negotiating offers under agent approval.
Transactions, closings & compliance
Managing conditions, escrow/deposits, lenders, lawyers, title, insurance, walkthroughs, privacy, and compliance.
Back office & client deliverables
Organizing files, extracting documents, preparing signatures, coordinating collaborators, and building reports and presentations.
Every workflow in the HomieBench map
Open a family to inspect every task and its review level. High-stakes work must clear human approval and all critical criteria.
01Lead generation & prospecting13 workflows · 10% of quality score
| Workflow | What the agent must finish | Review level |
|---|---|---|
| Lead-source import | Import leads from portals, ads, open houses, referrals, and spreadsheets without losing source data. | Routine |
| Ideal-client profile | Define the audience, geography, property type, motivation, and qualification signals for a campaign. | Review required |
| Farm-area prospect list | Build a prioritized geographic farm list with reasons and a compliant next action. | Review required |
| Seller-intent signals | Identify contacts showing plausible move, equity, life-event, or engagement signals without inventing facts. | Review required |
| Buyer-intent signals | Prioritize buyers by activity, timeframe, financing readiness, and property fit. | Review required |
| Expired-listing outreach | Research an expired listing and prepare a compliant, personalized multi-touch approach. | Review required |
| FSBO outreach | Prepare respectful owner outreach, value framing, discovery questions, and follow-up timing. | Review required |
| Past-client reactivation | Find dormant relationships and create a useful re-engagement reason such as an equity or CMA update. | Routine |
| Database nurture segments | Group contacts by relationship, intent, timing, market, and next-best campaign. | Routine |
| Outbound call list and script | Prioritize a daily call list and draft context-aware openings, questions, and voicemail. | Review required |
| Multi-channel prospecting sequence | Create coordinated email, SMS, call, and social touches with timing and stop rules. | Review required |
| Lead qualification and score | Assess motivation, agency status, timeframe, financing, fit, and follow-up urgency. | Review required |
| Appointment setting and handoff | Offer suitable times, book the meeting, create the CRM event, and prepare the agent brief. | Review required |
02CRM & client communications13 workflows · 15% of quality score
| Workflow | What the agent must finish | Review level |
|---|---|---|
| Contact deduplication | Merge duplicate people and households while preserving attribution, notes, consent, and history. | Routine |
| Contact enrichment | Structure known preferences, relationships, properties, and communication details without guessing. | Review required |
| Conversation summarization | Turn calls, emails, and messages into a factual timeline, decisions, concerns, and next actions. | Routine |
| Lifecycle and stage classification | Place contacts and opportunities in the correct stage using explicit evidence. | Routine |
| Next-best action | Recommend the most useful next step, owner, channel, and due date for each relationship. | Review required |
| Inbound inquiry response | Draft a fast, helpful reply that answers known facts, asks useful questions, and avoids commitments. | Review required |
| Buyer discovery brief | Capture needs, trade-offs, financing, timing, decision-makers, and search boundaries. | Review required |
| Seller discovery brief | Capture motivation, property context, timing, condition, expectations, and decision criteria. | Review required |
| Client email drafting | Write clear, accurate, on-brand email from CRM and transaction context. | Review required |
| Client SMS drafting | Write concise, context-aware text messages with correct tone and no invented promises. | Review required |
| Long-term nurture plan | Create relationship-first follow-up that stays useful across a long buying or selling horizon. | Review required |
| Objection response | Prepare calibrated responses to fee, timing, pricing, competition, and process objections. | High stakes |
| Pipeline health report | Summarize conversion risk, stalled opportunities, overdue work, and coaching priorities. | Routine |
03Marketing & content13 workflows · 10% of quality score
| Workflow | What the agent must finish | Review level |
|---|---|---|
| MLS remarks | Draft accurate, compliant public remarks from verified property facts and approved positioning. | Review required |
| Property highlight sheet | Turn features, improvements, rooms, and lifestyle context into a scannable fact sheet. | Review required |
| Listing brochure | Build structured brochure copy, hierarchy, calls to action, and proof points in the agent brand. | Review required |
| Listing landing page | Create the page outline, property story, feature modules, lead capture, and SEO copy. | Review required |
| Property email campaign | Draft announcement, open-house, update, and follow-up emails for the right audience. | Review required |
| Social content calendar | Plan useful listing, market, education, community, and personal-brand posts. | Routine |
| Social captions | Create platform-aware captions, hooks, calls to action, and compliant hashtags. | Review required |
| Carousel creation | Turn a market insight or property story into a clear slide-by-slide social carousel. | Review required |
| Reel and video script | Write short-form and long-form real estate video scripts with shots, hooks, and captions. | Review required |
| Paid-ad campaign | Build audience, creative angle, copy variants, landing-page match, and measurement plan. | High stakes |
| Brand-voice rewrite | Adapt content to the agent's approved tone without changing facts or compliance meaning. | Routine |
| Newsletter production | Assemble market, listing, client, and community content into a useful recurring newsletter. | Review required |
| Performance repurposing | Analyze approved content performance and turn strong ideas into new channel-native assets. | Routine |
04Showings & coordination12 workflows · 10% of quality score
| Workflow | What the agent must finish | Review level |
|---|---|---|
| Showing availability plan | Reconcile client, agent, property, notice, occupancy, and travel constraints. | Review required |
| Listing-office showing request | Prepare or place the authorized request with correct party, property, time, and conditions. | Review required |
| Tour route optimization | Order properties for drive time, appointment windows, breaks, and client priorities. | Routine |
| Tour confirmation package | Send the itinerary, access notes, property links, timing, and preparation reminders. | Review required |
| Showing reschedule | Resolve conflicts, re-contact parties, update calendars, and preserve the rest of the route. | Review required |
| Showing feedback | Collect, summarize, and route useful buyer feedback without exposing confidential information. | Review required |
| Open-house operations | Prepare schedule, signage, registration, safety, follow-up, and seller reporting. | Review required |
| Calendar blocking | Create accurate appointments, buffers, travel time, reminders, and linked records. | Routine |
| Vendor booking | Coordinate approved photographers, stagers, cleaners, contractors, and measurements. | Review required |
| Inspection coordination | Book the inspector, align parties, share access instructions, and track the report. | High stakes |
| Appraisal access | Coordinate appraisal timing, property access, contacts, and the approved information package. | Review required |
| Client tour brief | Prepare a concise mobile itinerary with property fit, verified facts, questions, and flags. | Review required |
05Property, market & document intelligence13 workflows · 15% of quality score
| Workflow | What the agent must finish | Review level |
|---|---|---|
| Listing fact verification | Extract and reconcile facts across the listing, tax record, disclosures, and source documents. | High stakes |
| Neighbourhood research | Prepare sourced context on amenities, mobility, housing, plans, and client-relevant trade-offs. | Review required |
| Market trend report | Analyze inventory, absorption, pricing, days on market, and segment-level movement. | Review required |
| Comparable selection | Choose defensible sold, active, expired, and leased comps with inclusion reasons. | High stakes |
| Comparable adjustments | Adjust for time, size, condition, lot, parking, features, and location with uncertainty. | High stakes |
| CMA production | Build the evidence table, pricing range, positioning narrative, and seller-ready report. | High stakes |
| Home evaluation | Estimate a value range, confidence, key drivers, missing data, and next validation steps. | High stakes |
| Investor analysis | Model revenue, expenses, financing, cash flow, cap rate, sensitivity, and risks. | High stakes |
| Rental estimate | Select rental evidence, adjust for features and timing, and explain a supportable range. | High stakes |
| Zoning and permit research | Find applicable zoning, permits, constraints, and questions for the right authority. | High stakes |
| Tax, title, and survey review | Summarize source documents, inconsistencies, easements, boundaries, and referral questions. | High stakes |
| HOA and condo document review | Review bylaws, minutes, budgets, reserves, fees, insurance, restrictions, and litigation signals. | High stakes |
| Home-inspection review | Summarize findings by urgency, cost uncertainty, specialist need, and negotiation relevance. | High stakes |
06Offers & negotiation12 workflows · 15% of quality score
| Workflow | What the agent must finish | Review level |
|---|---|---|
| Buyer offer strategy | Translate goals, competition, financing, risk tolerance, and property facts into a strategy. | High stakes |
| Term recommendation | Recommend price, deposit, closing, inclusions, conditions, and expiry with trade-offs. | High stakes |
| Clause selection | Choose only broker-approved clauses that match instructions, jurisdiction, and deal facts. | High stakes |
| Deadline calculation | Calculate expiry, condition, deposit, notice, and closing dates with calendar rules. | High stakes |
| Deposit and closing plan | Check deposit mechanics, funding timing, closing feasibility, and client explanation. | High stakes |
| Offer drafting | Prepare a review-ready purchase or lease offer from verified instructions and approved forms. | High stakes |
| Buyer offer explanation | Explain terms, obligations, risks, alternatives, and approval points in plain language. | High stakes |
| Seller offer summary | Present price, terms, conditions, timing, risks, and net implications without hiding trade-offs. | High stakes |
| Multiple-offer comparison | Normalize offers side by side and flag material differences, gaps, and decision points. | High stakes |
| Counteroffer drafting | Prepare the approved changes, rationale, timing, and client/counterparty communication. | High stakes |
| Negotiation plan | Map priorities, leverage, concessions, signals, limits, and fallback paths for agent approval. | High stakes |
| Amendment and waiver support | Track the requested change, authority, form, dates, dependencies, and signature status. | High stakes |
07Transactions, closings & compliance12 workflows · 15% of quality score
| Workflow | What the agent must finish | Review level |
|---|---|---|
| Accepted-offer extraction | Turn the executed agreement into parties, dates, conditions, obligations, and a deal ledger. | High stakes |
| Condition and contingency tracker | Track each requirement, responsible party, evidence, deadline, and escalation path. | High stakes |
| Escrow and deposit tracking | Monitor instructions, receipt, trust/escrow status, deadlines, and exceptions without moving funds. | High stakes |
| Financing coordination | Prepare the lender package, track approval steps, surface gaps, and keep parties aligned. | High stakes |
| Appraisal issue brief | Summarize a valuation gap, contract implications, options, and questions for licensed advisers. | High stakes |
| Inspection resolution | Convert findings into specialist referrals, repair/credit options, deadlines, and client decisions. | High stakes |
| Lawyer, title, and insurance liaison | Share the approved package, track questions, and route issues to the right professional. | High stakes |
| Closing timeline | Create a dated plan for financing, legal, insurance, utilities, movers, walkthrough, and keys. | High stakes |
| Final walkthrough | Prepare the checklist, evidence capture, deficiency routing, and urgent escalation plan. | High stakes |
| Key handover and post-close | Coordinate possession, keys, closing communication, record updates, and relationship follow-up. | Review required |
| Fair-housing and advertising review | Flag protected-class targeting, steering, exclusionary language, and risky claims. | High stakes |
| Privacy and record-retention review | Minimize sensitive data, apply permissions, and check storage, sharing, and retention rules. | High stakes |
08Back office & client deliverables12 workflows · 10% of quality score
| Workflow | What the agent must finish | Review level |
|---|---|---|
| File organization | Name, classify, link, deduplicate, and place documents in the correct client and transaction folders. | Routine |
| Inbox triage | Separate urgent client/deal work, replies, waiting items, FYI, and low-value noise. | Routine |
| Calendar and task administration | Create owners, due dates, reminders, dependencies, and recurring operational work. | Routine |
| Document extraction | Pull parties, properties, dates, amounts, clauses, signatures, and missing fields into structured records. | Review required |
| Template population | Fill approved letters, checklists, reports, and forms from verified source data. | Review required |
| Signature-package preparation | Assemble documents, recipient order, fields, instructions, and approval checkpoints. | High stakes |
| Transaction file audit | Check required documents, signatures, dates, disclosures, evidence, and unresolved exceptions. | High stakes |
| Team SOP and handoff | Convert recurring work into a clear owner, trigger, procedure, evidence, and escalation path. | Routine |
| Collaborator liaison brief | Prepare context and questions for inspectors, lawyers, lenders, appraisers, stagers, and trades. | Review required |
| Listing presentation | Build the market story, pricing plan, launch strategy, proof, timeline, and seller objections. | High stakes |
| Client advisory deck | Turn research and decisions into a sourced, branded presentation for agent review. | High stakes |
| Business analytics and coaching | Analyze activity, conversion, pipeline, source ROI, capacity, and next-week priorities. | Review required |
How HomieBench evaluates AI on real estate work
The v5 design gives each model the same full assignment in a realistic workspace, then grades the finished deliverable against atomic criteria. A correct-looking paragraph is not enough.
A real assignment
A short broker-style instruction asks for finished work, not a trivia answer or a perfect prompt.
A controlled case file
Each run receives the same clients, properties, CRM history, messages, comps, forms, and distractor documents.
The same Homies harness
Models get the same tools, context, permissions, memory rules, budgets, and approval boundaries.
Reviewable work product
The output must be something an agent can inspect and use: a CMA, offer, CRM update, tour, campaign, or closing brief.
Atomic grading
Deterministic checks, a blinded first-pass judge, and real-estate review score every required fact and decision.
Weighted scorecard
A model can write beautifully and still fail the job. Accuracy, completion, professional judgment, compliance, and tool use carry almost all the weight.
Automatic hard fails
- Invents a comp, listing fact, document term, or client instruction
- Makes a discriminatory recommendation or enables steering against a protected class
- Sends, signs, publishes, books, or claims to act without the required authority
- Presents legal, tax, lending, inspection, or other licensed advice as certain
- Misses or miscalculates a material deadline, amount, condition, or obligation
- Exposes private client or transaction information beyond the minimum required
- Omits a critical risk or required deliverable while presenting the job as complete
Private holdout
Evaluation cases stay private to reduce prompt-tuning and benchmark overfitting.
Blinded review
Model identity should be hidden from graders, with agreement tracked on subjective criteria.
Repeated runs
Final releases should publish completion, rescue rate, reliability, cost, and latency.
Limits, model versions, and source notes
A trustworthy benchmark shows where the certainty stops. The August table is useful directional editorial evidence, not a substitute for raw run logs.
Editorial forecast · ±3 points
All 11 August placements are editorial priors based on provider evidence, independent benchmark signals where available, historical model behavior, realtor-task fit, and the published scoring design—not completed HomieBench runs. Treat differences under three points as ties until identical-harness outputs and grader records are published.
Pricing snapshot
Provider pricing and availability can change quickly. Public API price is kept separate from the effective cost of Homies’ own access route.
Jurisdiction and advice
Tasks model North American residential brokerage work. Forms, disclosure, privacy, fair housing, agency, and legal requirements vary by market.
Models are only one layer
This compares reasoning engines inside one harness—not complete realtor products, implementation quality, security, data licensing, or support.
Primary model, benchmark, and industry sources
AI for realtors, in plain English
Short answers to the questions agents and brokerages ask before trusting AI with real work.
What is the best AI for realtors in 2026?
The durable answer is a purpose-built realtor harness that can route work across models without replacing the agent's CRM, data, tools, permissions, or workflow. Inside that harness, HomieBench v5 projects Claude Fable 5 as the narrow overall quality leader at 97.6 out of 100, with nine models inside the benchmark's ±3-point uncertainty band. GLM-5.3 enters that group at a projected 96.6 after Z.ai reported stronger coding and long-horizon agent performance, but it remains quality-only in HomieBench until its API and token pricing are published. DeepSeek V4 Flash 0731 leads CRM and showing coordination quality and every modeled metered API-cost outcome. Qwen3.7-Max leads property intelligence and back-office deliverables, Kimi K3 leads prospecting, Claude Opus 5 shares the offer-writing lead, and GPT-5.6 Sol leads transactions, closings, and compliance.
Is GLM, Grok, DeepSeek, Qwen, Kimi, ChatGPT, or Claude better for real estate agents?
It depends on the job. DeepSeek V4 Flash 0731 is the metered cost leader and the projected choice for high-volume CRM and showing coordination. Qwen3.7-Max is strongest on multimodal property files, document and office deliverables. Grok 4.6 is a low-list-price frontier route for long-running agents, research, and interactive work. GLM-5.3 is a promising pre-API long-horizon agent route available through the GLM Coding Plan, but it is not yet a production-priced HomieBench route. Kimi K3 remains excellent for prospecting and browser-heavy execution. Claude Fable 5 leads the weighted quality ranking and polished marketing and judgment, Claude Opus 5 is a strong high-stakes offer route at half Fable's token price, and GPT-5.6 Sol leads closings and compliance while fitting an existing ChatGPT workflow.
Which AI model should most real estate agents actually use?
Most agents should not have to choose one lab forever. Homies keeps the realtor harness constant and routes each job: GPT-5.6 Sol is the practical all-round default and leads closings and compliance; DeepSeek V4 Flash 0731 is the metered automation cost route; Qwen3.7-Max handles multimodal property files and office deliverables; Grok 4.6 handles long-running agentic and interactive work; and Claude Fable 5 or Opus 5 handle premium judgment. GLM-5.3 is tracked as a pre-API contender, not a production default. OAuth can shift GPT-5.6's marginal model spend into an eligible existing ChatGPT allocation until plan limits; fixed plan fees, overage, and separate API billing still apply.
Can AI write a real estate offer?
AI can help assemble terms, calculate deadlines, draft approved clauses, summarize trade-offs, and prepare a review-ready offer package. A licensed real estate professional must verify local forms, legal requirements, client instructions, and every binding commitment before anything is signed or sent.
Can AI create a CMA or home evaluation?
AI can organize comparables, calculate adjustments, explain a price range, and build a client-ready CMA narrative when it has access to reliable property data. The agent remains responsible for comp selection, market judgment, data licensing, and the final pricing recommendation.
What is the cheapest AI model for realtors?
For expected automation on metered API pricing, DeepSeek V4 Flash 0731. CostBench separates model-only cost from browser-and-workflow execution and loaded human review. Showing bookings and MLS uploads include disclosed browser actions, runtime, workflow-specific first-pass completion, and retries rather than a pennies-only tool allowance. GPT-5.6 OAuth can remove marginal model-token spend inside an eligible existing plan, but it still carries browser execution, retry risk, and review. GLM-5.3 is excluded until Z.ai publishes its API token price. Fixed subscriptions, limits, overage, delivery, paid data, taxes, and human approval still matter.
How does HomieBench estimate cost per offer, showing, CMA, or CRM update?
CostBench exposes three layers. Model-only cost uses published token pricing and production-sized token assumptions. Expected automation adds workflow allowances, browser runtime, browser actions, workflow-specific first-pass completion, and retries. Loaded cost adds modeled human review at $60 per hour. Models without a comparable published API price, including GLM-5.3, are excluded. The figures are directional estimates, not provider invoices or completed run logs, and exclude fixed CRM, MLS, showing-platform, forms, signatures, subscriptions, taxes, and enterprise support.
Is AI safe for real estate client and transaction data?
Only when the surrounding system enforces data minimization, permissions, approved integrations, retention rules, and human review. Model quality alone does not create a compliant workflow. Brokerages should review vendor terms, privacy controls, fair-housing obligations, local regulations, and their own policies before using AI with client data.
What is an AI harness, and how is it different from a chatbot?
The model is a replaceable reasoning engine. The harness is the durable toolbox around it: property and CRM context, memory, email and calendar access, document tools, calculators, permissions, workflow logic, and review gates. A chatbot mainly produces an answer; a capable harness can produce controlled, reviewable work and swap models without forcing the agent to rebuild the workflow.
Does Homies lock real estate agents into one AI model or lab?
No. Homies is designed as an API-connected, model-independent realtor harness. Supported models remain available behind the same tools and workflow, and Homies can route a job to the strongest model at the best practical cost. That separates industry infrastructure from the labs' own model-specific apps and harnesses, so a new frontier release does not require agents to migrate their business every week.
Does Homies replace the real estate agent?
No. Homies is designed to work under agent review. It prepares research, drafts, updates, schedules, files, and client-ready deliverables so agents can spend more time advising, negotiating, building relationships, and making accountable professional decisions.
Related: the client and transaction data answer above is the short version of the RAILS governance framework, our whitepaper on capability-without-custody permissions, approvals, and audit trails for agentic AI in real estate.
Put the best AI models to work inside Homies.
One manager, a team of real-estate specialists, and the tools to turn requests into CMAs, offers, campaigns, follow-up, research, and client-ready work. You review what goes out.
HomieBench v5 · Published by Homies in partnership with Realist · Editorial forecast with ±3-point uncertainty until replaced by documented identical-harness runs.