DOI: 10.3390/realestate3030010 ISSN: 2813-8090

Multimodal LLM-Based Property ConditionAssessment: A Per-Room Analysis Framework with Investor-Perspective Calibration

Ragul Shanmugam

Property condition assessment is a critical step in residential real estate investment underwriting, motivating after-repair value (ARV) estimates and rehabilitation cost projections. Traditional approaches rely on in-person inspections or manual photo review by experienced investors—processes that are time-consuming, subjective, and do not scale. Prior computer vision work on building analysis has focused on structural defect detection using convolutional neural networks but has not addressed the holistic, room-level condition assessment needed for residential investment decision-making. This paper presents a per-room analysis framework that leverages multimodal large language models (MLLMs) to assess the condition of residential properties from photographs. The framework analyzes each photo independently at the room level—detecting the room type, condition category, condition score, material features, and visible issues. Condition output is intended to feed a separate downstream rehabilitation cost and ARV estimation model that is outside the scope of this paper; the present empirical evaluation is restricted to per-photo condition assessment and inter-rater agreement with human experts. I evaluate the framework on two complementary datasets: (i) a primary per-image condition evaluation on 57 photographs from 14 real off-market properties in the Memphis, TN MSA, spanning three condition tiers (Fixer, Outdated, Standard), with independent labels from two experienced real estate investors; (ii) a secondary room classification evaluation on the public REI Dataset (51 attempted, 39 successful, 12 HTTP-503 failures). The room classification accuracy was 76.5% intention-to-analyze on REI (100% per-protocol on the 39 successful calls; 23.5% API failure rate) and 82.5% on the concierge dataset. The inter-rater agreement on the concierge dataset, with 95% bootstrap CIs (5000 resamples) and Spearman’s ρ as primary score statistic, was as follows: Cohen’s κ=0.773 (95% CI [0.64,0.90]) between Labeler A and the MLLM (weighted κ=0.853 [0.76,0.94]; ρ=0.906); and κ=0.502 [0.35,0.66] between Labeler B and the MLLM (ρ=0.858); both bracket the human–human reliability of κ=0.590 [0.42,0.74] (ρ=0.807). The MLLM’s κ asymmetry across the two labelers is statistically significant (Δκ=0.271, 95% bootstrap CI [0.115,0.429], p=0.0004), which I attribute to plausible training distribution and labeling style differences. A blind re-labeling sensitivity analysis on a stratified 15-image subsample yields anchoring-corrected κ estimates of approximately 0.65 (Labeler A) and 0.35 (Labeler B); the headline anchored values therefore sit at the upper bound of plausible blind-equivalent agreement. Failure modes concentrate at the Outdated tier and at the Outdated–Standard boundary, where humans themselves disagree most, indicating intrinsic taxonomy ambiguity rather than a model artifact. I make no claim to multi-market generalization and present multi-market extension as ongoing work.

More from our Archive