New to Claude Skills? Learn how to install them →

google on GitHub

Borderless Data Lakehouse

Free

Design secure, governed data lakehouses with AI integration.

by google17.6k stars on google/skills
2 views
Updated Aug 10, 2026
Get this skill

Free · Opens the source repo

What Borderless Data Lakehouse does

The Borderless Data Lakehouse skill is designed to assist developers and architects in creating a robust multi-product architecture that connects disparate data sources across various environments, including cloud and on-premises. This skill is particularly useful for organizations looking to implement a governed and secure data lakehouse that leverages agentic AI capabilities for enhanced data analytics and insights. By guiding users through a structured workflow, it ensures that all requirements are comprehensively addressed, from initial discovery to final validation.

The workflow is divided into four main phases: requirements discovery, solution design, implementation planning, and solution validation. During the requirements discovery phase, users are prompted to analyze their workload's functional and non-functional requirements, identify data sources, and assess security needs. This foundational understanding is crucial for the subsequent phases, where users will design the architecture and map components to appropriate Google Cloud products, ensuring alignment with best practices.

In the solution design phase, the skill emphasizes creating an architecture diagram using the Mermaid format, which visually represents the components and their interactions. This is complemented by generating design recommendations based on established guidelines, helping users to draft a comprehensive solution architecture guide. The implementation plan further assists users by providing automation instructions for deploying the solution, ensuring a smooth transition from design to execution.

This skill is tailored for data engineers, architects, and developers who require a sophisticated approach to integrating multiple data sources into a cohesive lakehouse architecture. It is not suitable for simpler, single-cloud data warehouse scenarios or non-AI workloads, making it essential for users with complex data integration needs.

When to use it

Use this skill when you need to design a governed, secure data lakehouse that connects multiple data sources and incorporates AI capabilities.

When not to use it

Avoid this skill for simple, single-cloud data warehouse solutions or when AI integration is not a requirement.

What you can build with it

Designing a Multi-Cloud Data Solution

When tasked with integrating data from various cloud providers and on-premises systems, this skill helps in crafting a cohesive architecture.

Implementing AI-Driven Analytics

For organizations looking to leverage AI for data analysis across different environments, this skill provides the necessary framework and recommendations.

Validating Data Architecture Designs

Use this skill to validate proposed data architectures against business requirements and Google Cloud best practices.

How to install Borderless Data Lakehouse

View source

1. Install with the skills CLI

npx skills add google/skills/google-cloud-solution-agentic-ai-borderless-data-lakehouse --agent claude-code

2. Or install it manually

Download the skill folder and drop it into ~/.claude/skills/ for all projects, or .claude/skills/ to scope it to one repo. Restart Claude Code so it picks up the new skill.

Anthropic's agentic coding CLI, and the reference implementation of Agent Skills. Drop a skill folder into ~/.claude/skills and Claude Code loads it automatically whenever a task matches the skill's description. Claude Code docs

Inside SKILL.md

Written by google

Borderless open data lakehouse agentic AI system

Follow this workflow to help users design and implement a custom multi-product solution in the cloud for a given workload, use case, or requirement.

Product Renaming & Terminology

When generating solution designs, architecture diagrams, and documentation, use the updated Google Cloud product names. For details on legacy vs. updated product names and terminology, see references/product_renaming.md.

Workflow

The solution design and implementation workflow consists of the following phases:

  • Phase 1: Requirements discovery and analysis: Analyze the workload's requirements, constraints, dependencies, and current state.
  • Phase 2: Solution design: Build a technology stack, architecture, and deployment configuration for the workload based on Google Cloud design best practices and recommendations.
  • Phase 3: Implementation plan: Generate automation and instructions to deploy the solution.
  • Phase 4: Solution validation: Validate that the deployment meets the requirements of the workload.

Phase 1: Requirements discovery and analysis

  • Step 1: Discover requirements: Understand the functional and non-functional requirements, business goals, and current state (if any) of the workload, including its architecture, dependencies, and constraints. Use the following questions to guide the requirements discovery process:

    • What are your primary data sources?
    • How do you manage and federate metadata across your data sources?
    • What are your security and credential management requirements?
    • What are the analytical and computational requirements to join and transform this borderless data?
    • What types of natural language prompts or user queries do you expect AI agents or end-users to execute against this data?
  • Step 2: Identify components: Based on the requirements analysis, identify the components of the workload and their relationships. Also identify any borderless components, hybrid components, or on-prem components that the solution needs to integrate with.

  • Step 3: Generate component decomposition: Generate a technical decomposition of the components of the workload.

  • Step 4: Ask for confirmation: Ask the user to confirm whether the generated technical decomposition matches their workload requirements.

  • Step 5: Iterate: If the user requests changes, then generate an updated technical decomposition, and ask the user to confirm the changes. Continue iterating until the user confirms the technical decomposition.

Phase 2: Solution design

Phase 3: Implementation plan

Phase 4: Solution validation

  • Step 1: Retrieve relevant verification resources (optional): If the resources from Phase 3 are not already in your context, retrieve the same implementation resources as the starting point for the validation checks and verification scripts that you generate in this phase.

  • Step 2: Define validation checks: Outline validation steps to verify that the deployed infrastructure meets the workload requirements:

    • Deployment dry-run: Commands like terraform plan to preview changes.
    • Connectivity and routing: Verification of network paths, load balancer routing, and service endpoints.
    • Security policies: Verification of restricted access, firewall rules, and IAM enforcement.
  • Step 3: Generate verification scripts: Draft lightweight scripts or command-line instructions (e.g. using curl or gcloud) that the user can run to perform these validation checks.

  • Step 4: Compile validation report: Document the validation steps, verification scripts, and expected outcomes in a single Markdown file.

  • Step 5: Conduct validation and finalize: Assist the user in executing the validation checks and troubleshooting any deployment issues. After the solution is validated successfully, request final approval from the user.

  • Step 6: Iterate: If the user requests changes, then generate an updated validation plan and repeat steps 2-5 until the user approves the validation plan.

Frequently asked questions about Borderless Data Lakehouse

Similar skills