
We present ArchiClip, a method for learning multi‑modal joint representations of text and freeform architectural surfaces. Building on the widely adopted multi‑modal contrastive learning paradigm, our approach focuses specifically on distilling knowledge from the architectural domain and enabling robust understanding of freeform 3D geometry. To support this, we introduce a strategy for constructing ArchiShape, a curated 3D dataset enriched with domain‑specific, rich textual descriptions. ArchiShape is generated automatically by combining procedural modeling and geometric shape descriptors with pre‑trained large language models, eliminating the need for manual annotation. We then design a bi‑modal pre‑training framework that learns aligned embeddings for both textual descriptions and architectural freeform surfaces. The framework incorporates a 3D backbone network tailored to architectural geometry and a custom batch‑sampling scheme to ensure efficient training. We evaluate ArchiClip on cross‑modal 3D retrieval of architectural freeforms, demonstrating its ability to encode rich geometric and domain‑specific concepts (e.g., curvature, spatial organization), and highlighting the benefits of domain‑aware multi‑modal representation learning.