docs: add talk-to-data case study
Add anonymized architecture, governance examples, diagrams, and interview materials.\n\nCo-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
This commit is contained in:
299
README.md
Normal file
299
README.md
Normal file
@@ -0,0 +1,299 @@
|
||||
# Talk-to-Data: Semantic Analytics for Enterprise Quotations
|
||||
|
||||
**A POC case study in governance-first semantic layer design: SAP → Snowflake → Cortex → Chatbot.**
|
||||
|
||||
---
|
||||
|
||||
## Problem Statement
|
||||
|
||||
Sales and operations teams across multiple regions handle thousands of quote requests monthly. Each inquiry requires:
|
||||
- **Manual research**: Discount policy lookups, product classification, customer contract data
|
||||
- **Cross-system queries**: SAP for pricing, internal databases for agreements, email chains with ops
|
||||
- **Coordination overhead**: Sales emails ops, ops researches and replies (5-30 minutes per quote)
|
||||
- **Scale challenge**: 30+ salespersons generating inquiries daily across the sales organization
|
||||
|
||||
**Result**: Quote turnaround times of 5-30 minutes, even for routine requests. Ops team context-switches between email, SAP, and spreadsheets. Sales can't get answers fast enough to close deals.
|
||||
|
||||
---
|
||||
|
||||
## Solution: Semantic Analytics Layer + Cortex Guardrails
|
||||
|
||||
**Core insight**: Build a semantic layer that serves as the single source of truth for quote logic, then consume it via two interfaces: BI dashboards and an LLM-powered chatbot. Governance lives in the data layer, not in prompts.
|
||||
|
||||
### Architecture Overview
|
||||
|
||||
```
|
||||
┌─────────────────────────────────────────────────────────────┐
|
||||
│ SAP ERP (Sales, Pricing, Discounts) │
|
||||
│ Updated daily via batch extract (05:00 UTC) │
|
||||
└────────────────────────┬────────────────────────────────────┘
|
||||
│
|
||||
Daily ELT Job
|
||||
│
|
||||
┌────────────────────────▼────────────────────────────────────┐
|
||||
│ Snowflake Raw Layer │
|
||||
│ - raw.quotations (document, line item, amount, ...) │
|
||||
│ - raw.discount_conditions (discount tiers, policies) │
|
||||
│ - raw.customer_master (customer codes, agreements) │
|
||||
└────────────────────────┬────────────────────────────────────┘
|
||||
│
|
||||
dbt Staging
|
||||
│
|
||||
┌────────────────────────▼────────────────────────────────────┐
|
||||
│ Staging Layer (STAGING schema) │
|
||||
│ - Column mapping, type casting, null handling │
|
||||
│ - No business logic, just data preparation │
|
||||
└────────────────────────┬────────────────────────────────────┘
|
||||
│
|
||||
dbt Transforms
|
||||
│
|
||||
┌────────────────────────▼────────────────────────────────────┐
|
||||
│ Transform Layer (TRANSFORM schema) │
|
||||
│ - Join discount conditions, policies, agreements │
|
||||
│ - Calculate: discount tiers, totals, exceptions │
|
||||
│ - Apply: governance flags, audit columns │
|
||||
│ - Models: trf_quotation, trf_chatbot_quotation │
|
||||
└────────────────────────┬────────────────────────────────────┘
|
||||
│
|
||||
Semantic Views (DDL)
|
||||
│
|
||||
┌────────────────────────▼────────────────────────────────────┐
|
||||
│ Semantic Layer (DISTRIBUTE_DDL schema) │
|
||||
│ - sv_quotation: Full BI consumption │
|
||||
│ - sv_chatbot_quotation: LLM-optimized dimensions │
|
||||
│ - Exposes: business logic (discounts, policies) │
|
||||
│ - Hides: implementation details (temp tables, keys) │
|
||||
└────────────────────────┬────────────────────────────────────┘
|
||||
│
|
||||
Cortex Analytics Guardrails
|
||||
│
|
||||
┌───────────────┴───────────────┐
|
||||
│ │
|
||||
┌────────▼──────────────┐ ┌─────────▼─────────────┐
|
||||
│ Snowflake UI │ │ Internal Workspace │
|
||||
│ (BI Dashboard) │ │ (Chatbot + LLM) │
|
||||
│ │ │ │
|
||||
│ - Visual analysis │ │ - Quote Q&A │
|
||||
│ - Sales reporting │ │ - Prompt engineering │
|
||||
│ - Ops monitoring │ │ - Real-time answers │
|
||||
└───────────┬──────────┘ └──────────┬────────────┘
|
||||
│ │
|
||||
│ <1 minute turnaround
|
||||
│
|
||||
End Users (Sales, Ops)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Stack: Why Each Layer?
|
||||
|
||||
### SAP → Snowflake ELT (Daily Batch)
|
||||
- **Why**: Quotation data lives in SAP; Snowflake is the reporting warehouse
|
||||
- **How**: Daily batch extract (05:00 UTC), once-per-day refresh sufficient for quote workflow (real-time inventory not required)
|
||||
- **Trade-off**: Once-daily freshness OK for most quote requests; could escalate to hourly for real-time dashboard if needed
|
||||
|
||||
### Staging → Transform → Semantic Layer (dbt)
|
||||
- **Why dbt?**
|
||||
- **Lineage**: Column-level traceability (SAP field → transform → semantic dimension)
|
||||
- **Testability**: Unit tests on discount calculations, data quality checks
|
||||
- **Team reuse**: SQL-first approach, skill portability across data team
|
||||
- **Governance**: Schema.yml documents business logic and access rules
|
||||
- **Reproducibility**: Version-controlled transformations, audit trail of changes
|
||||
|
||||
- **Why separate layers?**
|
||||
- **Staging**: Decouples source-system brittleness from business logic
|
||||
- **Transform**: Concentrates business rules (discount math, policy flags) in one place
|
||||
- **Semantic**: Exposes curated dimensions to downstream consumers; hides technical debt
|
||||
|
||||
### Snowflake Semantic Views
|
||||
- **What**: Semantic views (DDL) wrap transform models; define dimensions, facts, and synonyms
|
||||
- **Why**:
|
||||
- Define the contract between data and consumers (Sales, Cortex)
|
||||
- Synonyms enable natural-language queries ("desconto cliente" → `discount_customer_pct`)
|
||||
- Consistent schema across BI dashboards and LLM consumption
|
||||
- **Example**: A chatbot user asks "What's the customer discount on this quote?" The semantic layer maps "customer discount" (natural language) to `discount_customer_pct` (column), and Cortex applies guardrails (hides the raw percentage and shows only "within policy" or "requires approval")
|
||||
|
||||
### Cortex Analytics + Guardrails
|
||||
- **What**: Snowflake's native LLM platform + custom governance policies
|
||||
- **Why Cortex, not custom LLM?**
|
||||
- Built-in guardrails: protection of sensitive information, access control, and cost controls
|
||||
- Fast prototyping: Cortex agents consume Snowflake semantic views directly (no ETL to another system)
|
||||
- Governance-as-code: Policies live alongside data, versioned, audited
|
||||
- Security: Data stays in Snowflake; LLM calls don't leak sensitive data
|
||||
- **Trade-off**: Cortex is Snowflake-specific (vs. generic LLM + Anthropic); acceptable for POC, revisit if multi-cloud needed later
|
||||
|
||||
### Dual Consumption (Snowflake UI + Internal Workspace)
|
||||
- **Snowflake UI**: Native BI dashboards for sales/ops teams. Semantic model feeds Power BI or Tableau directly.
|
||||
- **Internal Workspace**: Existing chatbot environment. Cortex queries the semantic layer + prompt engineering layer refines responses.
|
||||
- **Why two interfaces?** Users consume data differently — dashboards for exploration, chatbot for direct answers. Semantic layer unified the source, reducing maintenance.
|
||||
|
||||
---
|
||||
|
||||
## Key Decisions & Trade-offs
|
||||
|
||||
| Decision | Rationale | Trade-off |
|
||||
|----------|-----------|----------|
|
||||
| **Governance-first** | Discount policies, protection of personal and customer information, and exception escalation built into the data layer | Requires upfront effort to define policies; can't skip it for "speed" |
|
||||
| **Semantic views (DDL)** | Single contract between data → BI/LLM | Requires mapping business terms to columns; initial work to get synonyms right |
|
||||
| **Once-daily batch** | Sufficient for quote workflow; reduces ELT complexity | Not real-time; acceptable for POC; upgrade to hourly if pilot shows need |
|
||||
| **Cortex (vs. custom LLM)** | Fast prototyping, built-in guardrails, stays in Snowflake ecosystem | Cortex-only (could integrate with Anthropic later) |
|
||||
| **Separate transform model per interface** | `trf_quotation` for BI, `trf_chatbot_quotation` for LLM — optimize each | Slight duplication; worth it if interfaces have different freshness/cardinality needs |
|
||||
|
||||
---
|
||||
|
||||
## Governance: How We Built Safety In
|
||||
|
||||
Governance is not a checklist; it's embedded in the data layer.
|
||||
|
||||
### 1. **Discount Policy Guardrails**
|
||||
- **Rule**: No discount exceeds the configured policy ceiling
|
||||
- **Implementation**:
|
||||
- Transform model calculates `discount_above_ceiling` flag
|
||||
- Cortex policy hides raw discount percentages; shows only "Within Policy" or "Requires Approval"
|
||||
- Chatbot escalates flagged quotes to ops team (max 10 escalations/hour to prevent bot spam)
|
||||
- **Effect**: Sales can't accidentally propose an unapproved discount; exceptions are visible and auditable
|
||||
|
||||
### 2. **Protection of Personal and Customer Information**
|
||||
- **Rule**: The chatbot must not expose customer names, contact details, or negotiated terms unless the user is authorized
|
||||
- **Implementation**:
|
||||
- `sv_chatbot_quotation` includes only aggregated customer data (customer_id, portfolio_code)
|
||||
- Cortex policy replaces `customer_name` with "[Customer information protected]" in chatbot responses
|
||||
- Exception: Sales team in authorized role can see unmasked values via dashboard
|
||||
- **Effect**: Chatbot can answer quote questions without leaking customer contracts
|
||||
|
||||
### 3. **Access Control & Audit Trail**
|
||||
- **Rule**: Different roles see different data
|
||||
- Sales: Can query `sv_chatbot_quotation` (limited fields, no raw discounts)
|
||||
- Ops: Can query `sv_quotation` (full discount details, for policy review)
|
||||
- Cortex Agent: Can query `sv_chatbot_quotation` only (locked down for LLM)
|
||||
- **Implementation**: Snowflake role-based access control (RBAC) + Cortex `authorized_teams` config
|
||||
- **Audit**: All queries logged; policy violations trigger alerts
|
||||
|
||||
### 4. **Data Freshness Contract**
|
||||
- **Rule**: Chatbot rejects queries on stale data; warns user if data is > 24 hours old
|
||||
- **Implementation**:
|
||||
- dbt model adds `load_timestamp` column
|
||||
- Cortex quality gate checks: `current_timestamp() - load_timestamp < 24 hours`
|
||||
- Failure: Returns "Data unavailable; please retry in 1 hour" (prevents stale-data surprises)
|
||||
|
||||
---
|
||||
|
||||
## Results: From POC to Impact
|
||||
|
||||
### Estimated Impact (Based on User Interviews)
|
||||
- **Quote turnaround**: 5-30 min (email + manual lookup) → <1 min (chatbot query)
|
||||
- Savings per quote: 5-30 minutes
|
||||
- Pilots with 5 sales reps: ~50+ hours/week freed (estimated)
|
||||
- Extrapolation: 30+ sales team × 5 hours/week = 150+ hours recovered org-wide
|
||||
- **Error reduction**: Discount exceptions caught by policy guardrails (vs. manually reviewed)
|
||||
- **Team velocity**: POC shipped in weeks (not months), governance patterns established for scale
|
||||
|
||||
### Why This Matters for Delivery Managers
|
||||
1. **Governance-first approach**: We didn't build a prototype and hope for policy later. Guardrails were baked in from day one.
|
||||
2. **Reusable patterns**: The semantic layer + Cortex + dbt playbook works for any quote/contract/discount domain. Next use cases ship faster.
|
||||
3. **Measurable before scale**: We validated impact with a small pilot; ready to scale to 30+ users with confidence.
|
||||
4. **Ops visibility**: Sales/ops can see what the bot is doing (audit logs, escalation flags); not a black box.
|
||||
|
||||
---
|
||||
|
||||
## How to Use This Repo
|
||||
|
||||
### Deliverables
|
||||
- **`dbt/`**: Anonymized dbt models (staging, transforms, semantic views)
|
||||
- `models/staging/stg_*.sql` — Raw layer transformations
|
||||
- `models/transform/trf_*.sql` — Business logic
|
||||
- `models/distribute_ddl/sv_*.yml` — Semantic views with synonyms
|
||||
- **`cortex/`**: Governance config (information protection, access control, audit rules)
|
||||
- `cortex_governance_config.yaml` — Guardrails, freshness contracts, LLM policies
|
||||
- **`docs/`**: Architecture diagrams, decision logs
|
||||
- `architecture_dag.md` — Data flow Mermaid diagram
|
||||
- `decisions.md` — Architecture decision records (ADRs)
|
||||
- **`README.md`**: This file (the narrative for interviewers/stakeholders)
|
||||
|
||||
### How to Adapt This to Your Domain
|
||||
1. **Replace discount logic** with your domain (e.g., pricing tiers, contract terms, approval workflows)
|
||||
2. **Adjust freshness** (once-daily to hourly/real-time) based on your use case
|
||||
3. **Modify semantic dimensions** to match your business terminology
|
||||
4. **Update Cortex policies** for your personal-information and governance rules
|
||||
5. **Test with a small pilot** (5-10 users) before org-wide rollout
|
||||
|
||||
---
|
||||
|
||||
## Lessons Learned
|
||||
|
||||
### What Worked
|
||||
✅ **Semantic views as the contract**: Defining dimensions + synonyms upfront saved rework later. Sales/ops understood the model; Cortex had clear field mappings.
|
||||
|
||||
✅ **Governance in the data layer**: Embedding discount flags and information-protection rules in dbt/Cortex meant the bot couldn't bypass them. No workaround hacks possible.
|
||||
|
||||
✅ **Once-daily batch sufficient**: For this workflow, 24-hour freshness was OK. We didn't need real-time complexity (yet).
|
||||
|
||||
✅ **dbt for team alignment**: SQL-first + schema docs meant the data team and business stakeholders spoke the same language. Onboarding new people was fast.
|
||||
|
||||
### What We'd Change at Scale
|
||||
⚠️ **Real-time discount updates**: If pilot shows sales need <1hr freshness on policy changes, escalate to hourly ELT.
|
||||
|
||||
⚠️ **Multi-region governance**: Current config is single-region (BR). If expanding to US/CA, add region-specific discount tiers and approval workflows.
|
||||
|
||||
⚠️ **Cortex cost**: Monitor Cortex token usage as chatbot volume scales. Current config has cost controls (monthly budget, per-query max tokens); may need aggressive pruning at scale.
|
||||
|
||||
⚠️ **Semantic view complexity**: If adding more dimensions (15→50), consider splitting into focused semantic views (e.g., `sv_quote_discounts`, `sv_quote_compliance`) to keep queries fast.
|
||||
|
||||
---
|
||||
|
||||
## Tech Stack Summary
|
||||
|
||||
| Component | Tool | Why |
|
||||
|-----------|------|-----|
|
||||
| **Data Warehouse** | Snowflake | Scalable, built-in Cortex, governance features |
|
||||
| **Transformations** | dbt + Jinja | SQL lineage, testability, team skill reuse |
|
||||
| **Semantic Layer** | Snowflake DDL | Data contract, synonyms for NLP, single source of truth |
|
||||
| **AI/Guardrails** | Cortex Analytics | Native LLM, information protection, access control, cost controls |
|
||||
| **Consumption 1** | Snowflake Native App / BI Tool | Dashboard for ops/sales exploration |
|
||||
| **Consumption 2** | Internal Workspace + Prompt Engineering | Chatbot for direct question-answering |
|
||||
| **Infrastructure** | Cloud-native Snowflake | No ops overhead, pay-per-query |
|
||||
|
||||
---
|
||||
|
||||
## FAQ
|
||||
|
||||
**Q: Can I use this with a different warehouse (e.g., BigQuery, Redshift)?**
|
||||
A: Yes. Adapt the dbt dialect (BigQuery: `jinja-sql`, Redshift: `redshift` profile). Cortex is Snowflake-native; substitute with your own LLM guardrails (policy enforcement layer).
|
||||
|
||||
**Q: Is once-daily refresh really enough?**
|
||||
A: For quote research, yes. If you need real-time pricing updates, escalate to hourly ELT. The architecture supports it; just change the schedule in your orchestrator.
|
||||
|
||||
**Q: How do I handle discount exceptions?**
|
||||
A: Cortex policy flags them and escalates to a human review queue (max 10/hour). Ops team approves or rejects in Snowflake; bot learns the decision for future similar quotes.
|
||||
|
||||
**Q: What if I need to scale to 500+ users?**
|
||||
A: Semantic layer architecture scales horizontally. Snowflake handles concurrency. Monitor Cortex token usage and partition semantic views if queries slow down. dbt stays the same.
|
||||
|
||||
**Q: Can I export this as a Snowflake Native App for partners?**
|
||||
A: Yes. Package the semantic views + Cortex policies as an app; partners can install it and use the chatbot without seeing raw data.
|
||||
|
||||
---
|
||||
|
||||
## Next Steps
|
||||
|
||||
1. **Expand pilot**: Rollout to 30+ sales team over 4 weeks
|
||||
2. **Measure impact**: Track quote turnaround times, bot usage, escalation rates
|
||||
3. **Iterate on policies**: Refine discount rules and information-protection controls based on pilot feedback
|
||||
4. **Add new domains**: Reuse the semantic layer pattern for contracts, serviceability, pricing
|
||||
5. **Scale LLM**: Move from proof-of-concept to production volume (monitor cost, latency)
|
||||
|
||||
---
|
||||
|
||||
## Contact & Questions
|
||||
|
||||
This case study demonstrates:
|
||||
- Governance-first architecture for enterprise AI
|
||||
- Semantic layer design (SAP → dbt → Snowflake → LLM)
|
||||
- Operationalizing guardrails and audit trails
|
||||
- Shipping production-ready POCs in weeks
|
||||
|
||||
For questions about the design decisions, data flow, or how to adapt this to your domain, reach out.
|
||||
|
||||
---
|
||||
|
||||
**Co-authored by Copilot**
|
||||
Reference in New Issue
Block a user