Skip to main content

Getting Started

Starlake Skills is an open-source skills bundle that gives your AI coding assistant deep expertise in building, configuring, and operating Starlake data pipelines: a skill for every CLI command and configuration pattern. It works with Claude Code, GitHub Copilot, and Gemini CLI; the installer links the skills into all three by default.

Whether you're setting up a new data project, configuring ingestion pipelines, writing transformations, or deploying orchestration DAGs — Starlake Skills gives your AI assistant deep expertise in every aspect of the Starlake platform.

What You Can Do

CategoryExamples
Ingestion & LoadingAuto-infer schemas, load CSV/JSON/XML, Kafka, Elasticsearch
TransformationSQL/Python transformations with write strategies
ExtractionExtract schemas and data from BigQuery, JDBC sources, REST APIs
Schema ManagementBootstrap projects, Excel-to-YAML, DDL generation
Data QualityExpectations with Jinja2 macros and validation patterns
Semantic LayerBusiness models for BI tools and AI agents in metadata/semantic/
LineageColumn-level, table-level, and ACL dependency tracking
OperationsValidation, metrics, freshness, GizmoSQL, migrations
SecurityIAM policies, RLS, CLS, privacy transformations
OrchestrationAirflow and Dagster DAG generation and deployment
UtilitiesParquet conversion, comparisons, site generation

Supported Platforms

Data Warehouses

  • BigQuery: Native and Spark loaders
  • Snowflake: JDBC connectivity
  • DuckDB: Embedded SQL engine
  • PostgreSQL: JDBC connectivity
  • Redshift: JDBC connectivity
  • Databricks: FS and Spark engines

Processing Engines

  • Spark: Distributed processing
  • Native: Built-in Starlake engine
  • DuckDB: Embedded analytical SQL

Orchestration

  • Apache Airflow: Python DAG generation
  • Dagster: Asset-based orchestration

Data Formats

CSV, JSON, XML, Parquet, Elasticsearch indices, Kafka topics

Skills and Starflow: Two Ways to Work

Starlake Skills provides two complementary approaches depending on the scope of your work:

Starlake Skills: Direct Access to Every Command

Skills integrate directly into your assistant (Claude Code, GitHub Copilot, or Gemini CLI). Each skill gives you deep expertise on a specific Starlake capability — CLI syntax, YAML configuration, write strategies, engine-specific behaviors, and production best practices. Examples in these docs use Claude Code syntax; the same skills trigger from natural language in every supported assistant.

Use skills when you have a targeted task: loading a file, writing a transformation, generating a DAG, or configuring a connection.

You: How do I load CSV files from GCS into BigQuery with deduplication?

Claude: [Uses the `load` skill to provide complete YAML configuration
with UPSERT_BY_KEY_AND_TIMESTAMP write strategy, domain config,
and schema definitions]

Starflow: Guided Methodology for End-to-End Projects

Starflow is an optional guided methodology layer built on top of Starlake Skills. Where individual skills answer "how do I do X?", Starflow answers "what should I do next and why?"

Starflow organizes data pipeline projects into four phases — Discovery, Architecture, Pipeline Design, and Implementation: each with dedicated skills and specialized agent personas that guide you through the full lifecycle.

Use Starflow when you're tackling a broader initiative: starting a new data platform, migrating from legacy ETL, onboarding a team, or reviewing an existing architecture.

You: /starflow-data-architect Design a data platform for our e-commerce analytics

Winston: [Guides you through architecture decisions — layers, engines,
storage, governance — then hands off to implementation skills]

How They Fit Together

Starlake SkillsStarflow
ScopeSingle task or commandMulti-step project lifecycle
ApproachDirect — ask and get an answerGuided — phased workflow with recommendations
Best forLoading, transforming, configuring, deployingDiscovery, architecture, planning, reviews
PersonasNone — you drive5 agent personas (Lea, Winston, Amelia, Quinn, Max)

Starflow skills call on the underlying Starlake Skills during implementation, so the two layers work together seamlessly. You can start with Starflow for planning and architecture, then drop into individual skills for hands-on configuration — or skip Starflow entirely and use skills directly for quick tasks.

Next Steps

  • Quickstart: Install and use your first skill in 5 minutes
  • Setup: Detailed installation and configuration options
  • Skills Catalog: Browse all skills by category
  • Starflow Method: Guided methodology for end-to-end projects