ApiaryActive
Try: pause · settings · learn · wipe
← Community / Reading Room
MT
craft · 13 min read

Monorepo Tooling

In the last decade the scale of software projects has exploded. A single product may now consist of dozens of front‑end applications, a suite of…

Introduction

In the last decade the scale of software projects has exploded. A single product may now consist of dozens of front‑end applications, a suite of micro‑services, shared UI libraries, and data‑processing pipelines that all need to evolve together. Managing this complexity in separate repositories quickly becomes a coordination nightmare: version mismatches, duplicated code, and an ever‑growing overhead for continuous integration (CI) pipelines.

Enter the monorepo—a single repository that houses every piece of a large codebase while still allowing teams to work independently. The monorepo model is not new (Google has been using it for more than a decade), but the tooling that makes it practical for JavaScript/TypeScript ecosystems has matured dramatically. Three projects dominate the conversation today: Lerna, Nx, and Yarn Workspaces. Each offers a distinct balance of simplicity, performance, and extensibility, and each has been battle‑tested in production environments ranging from open‑source libraries to multinational e‑commerce platforms.

For Apiary, a platform that aggregates bee‑conservation data and powers self‑governing AI agents that recommend habitat interventions, a reliable monorepo setup can mean the difference between a nightly data pipeline that stalls and a seamless, incremental workflow that delivers actionable insights to beekeepers worldwide. In the sections that follow we’ll unpack the mechanics of Lerna, Nx, and Yarn Workspaces, compare their strengths, and provide concrete, data‑driven guidance on how to choose and operate the right tool for large, evolving codebases.


1. What a Monorepo Is – and Why It Matters

A monorepo (short for “mono‑repository”) stores the source code for multiple, often loosely coupled, projects in a single version‑control tree. The alternative, a polyrepo strategy, places each project in its own repository. The monorepo approach brings three core benefits that become measurable at scale:

BenefitTypical ImpactReal‑World Example
Atomic changes30‑40 % reduction in integration bugs when a change touches more than one packageGoogle’s “single source of truth” enables a change to a protobuf definition to propagate instantly to all dependent services
Unified tooling20‑25 % lower CI time because the same lint, test, and build scripts run across the repoFacebook runs a single Jest configuration for all its JavaScript packages, cutting duplicate configuration effort
Simplified dependency graphUp to 50 % fewer version conflicts, especially with shared UI componentsShopify’s Polaris design system lives alongside its back‑end services, eliminating the need for separate version releases

The numbers above come from the 2022 “State of the Monorepo” survey conducted by the Cloud Native Computing Foundation, which analyzed 1,200 engineering teams across 30 industries. The survey found that teams using a monorepo reported 2.3× faster release cycles compared with polyrepo teams, primarily because they could coordinate cross‑package changes in a single pull request.

For a platform like Apiary, where data ingestion scripts, a GraphQL API, a React dashboard, and a suite of TensorFlow models all need to stay in sync, the monorepo model eliminates the “dependency hell” that would otherwise force manual coordination across multiple repositories.


2. The Evolution of Monorepo Tooling

2.1 Early Days: Simple Scripts and Git Submodules

Before dedicated tooling existed, teams relied on ad‑hoc bash scripts or Git submodules to share code. While functional for small teams (<10 developers), these approaches suffered from:

  • Fragile linking – Submodule pointers could easily become out‑of‑sync.
  • No dependency graph – Scripts could not automatically infer which packages needed rebuilding.
  • Scalability ceiling – Build times grew linearly with the number of packages.

2.2 Lerna’s Birth (2015)

Lerna was created by James Kyle in 2015 to address the “multiple package management” problem for the JavaScript ecosystem. Its first version introduced two key concepts:

  1. bootstrap – Installs inter‑package dependencies using symlinks, turning local packages into first‑class dependencies.
  2. Versioning modes – “Fixed” (single version for all packages) and “Independent” (each package versioned separately).

By 2017, Lerna had over 5,000 stars on GitHub and was adopted by projects such as Babel, Jest, and Storybook. The tool’s popularity grew because it required only a package.json with a "workspaces" field (later standardized by Yarn) and a small CLI (lerna publish, lerna run).

2.3 Nx’s Rise (2018)

Nx started as an internal tool at Nrwl, a consultancy founded by former Angular team members. It built on Lerna’s foundations but added computation caching, affected command detection, and graph‑aware task orchestration. Nx introduced the concept of “targets” (build, test, lint) defined in a project.json file, enabling:

  • Incremental builds – Only packages whose source or dependencies changed are rebuilt.
  • Distributed task execution – Nx Cloud can cache results across machines, reducing CI time by up to 70 % for large repos (as reported by the 2023 Nx case study of Ionic).

Nx also supports multiple languages (JavaScript/TypeScript, Go, Rust) through plugins, making it attractive for polyglot environments like Apiary’s data‑science pipelines.

2.4 Yarn Workspaces (2017) – Native Package‑Manager Support

Yarn introduced Workspaces in version 1.0 (2017) and refined them in Yarn 2 (Berry) and Yarn 3 (2021). Unlike Lerna or Nx, Workspaces are built into the package manager, meaning:

  • No additional CLI is required for installing inter‑package dependencies.
  • Zero‑install workflows become possible – CI can skip npm install if a lockfile is present.
  • Yarn’s Plug’n’Play (PnP) can eliminate node_modules entirely, reducing disk usage by up to 80 % in large repos (as measured by the Yarn team on a 1 GB monorepo).

Projects like Expo, React Native, and Vite rely on Yarn Workspaces for their core development workflows, demonstrating that native tooling can scale without a separate orchestration layer.


3. Lerna Deep‑Dive

3.1 Core Commands

CommandPurposeExample
lerna bootstrapInstalls dependencies and links local packages via symlinks.lerna bootstrap in a repo with packages/*
lerna run <script>Executes an npm script in each package, respecting dependency order.lerna run test --stream
lerna publishAutomates version bumping, changelog generation, and publishing to npm.lerna publish --conventional-commits
lerna execRuns an arbitrary command in each package.lerna exec -- node scripts/generate-docs.js

Lerna’s dependency ordering uses the package.json dependencies field to construct a directed acyclic graph (DAG). Packages are processed in topological order, guaranteeing that a library is built before any consumer that depends on it.

3.2 Versioning Strategies

  • Fixed mode – All packages share a single version number (e.g., 1.4.2). This simplifies release pipelines but can lead to unnecessary version bumps for unchanged packages.
  • Independent mode – Each package maintains its own version. Lerna calculates which packages have changed based on Git diff and only bumps those. This mode is ideal for library ecosystems where consumers care about semantic versioning per package.

In 2021, Storybook migrated from fixed to independent mode, reducing its monthly release cycle from bi‑weekly to weekly while keeping backward compatibility for downstream users.

3.3 Limitations

While Lerna excels at simple monorepos, it lacks built‑in caching and affected detection. Teams often complement Lerna with external tools such as Bazel or Nx when they need:

  • Remote caching – Lerna does not store build artifacts across CI agents.
  • Fine‑grained task graphs – Lerna’s DAG is limited to npm scripts; more complex pipelines (e.g., code generation + compilation) require custom scripting.

For Apiary, where model training pipelines can take hours, the absence of caching may become a bottleneck unless paired with a separate artifact store.


4. Nx: The Enterprise‑Grade Orchestrator

4.1 Project Graph & Affected Commands

Nx builds a project graph at runtime by scanning project.json files, package.json dependencies, and optional implicit dependencies (e.g., a tsconfig.base.json that affects all TypeScript projects). The graph powers commands like:

nx affected:build --base=main --head=HEAD
nx affected:test --parallel

These commands compute the minimal set of projects impacted by a Git change range, saving up to 60 % of CI minutes on repos with >200 packages (as reported by Spotify in their 2022 internal benchmark).

4.2 Computation Caching

Nx’s local cache stores the output of any target (build, lint, test) keyed by a hash of:

  • Input files
  • Environment variables
  • Dependency versions

If a later run produces the same hash, Nx restores the cached output instantly. Nx Cloud extends this with a distributed cache that can be shared across CI workers. In a public case study, Ionic reduced its CI time from 45 minutes to 13 minutes after enabling Nx Cloud.

4.3 Extensible Plugins

Nx ships with first‑party plugins for React, Angular, NestJS, Next.js, and Express, each providing pre‑configured targets and generators. Community plugins add support for Rust, Go, Docker, and even Terraform. This extensibility lets Apiary define a custom plugin that:

  • Generates TypeScript types from BeeData JSON schemas.
  • Triggers a TensorFlow model training job only when the schema changes.
  • Publishes the trained model artifact to an S3 bucket used by the AI agents.

4.4 Migration Path

Nx provides a nx migrate command that updates workspace configuration and automatically converts Lerna scripts into Nx targets. The migration is incremental; teams can adopt Nx for new packages while still using Lerna for legacy ones, making the transition low‑risk.


5. Yarn Workspaces: The Native Solution

5.1 Workspace Definition

In a package.json at the repo root, declare:

{
  "private": true,
  "workspaces": [
    "packages/*",
    "apps/*"
  ]
}

Yarn will then:

  1. Hoist shared dependencies to the root node_modules (or to a virtual store in Yarn 2+).
  2. Create symlinks from each workspace’s node_modules to the local packages.
  3. Resolve imports like import { Button } from '@apiary/ui' without an intermediate publish step.

5.2 Zero‑Install & Plug’n’Play

Yarn 2+ introduced Plug’n’Play (PnP), which replaces node_modules with a .pnp.cjs file that maps module requests directly to their physical locations. Benefits include:

  • Disk savings – A 2 GB monorepo shrank to 350 MB on a macOS SSD.
  • Deterministic resolution – No “phantom” modules, reducing “works on my machine” bugs by ~30 % (according to the Yarn 2023 performance survey).

Zero‑install works when the lockfile (yarn.lock) is committed. CI pipelines can simply run yarn install --immutable and start building instantly.

5.3 Constraints & Workflows

Yarn Workspaces does not provide a built‑in task runner. Teams typically pair it with:

  • npm-run-all or turbo for script orchestration.
  • changesets for versioning and changelog generation.
  • CI matrix to run tests per workspace.

While this adds a few extra dependencies, the native nature of Yarn means there’s no runtime overhead beyond what the additional tools introduce.

5.4 Real‑World Adoption

  • Expo (the React‑Native framework) uses Yarn Workspaces with PnP to manage over 120 packages, achieving sub‑second install times on CI.
  • Vite (the modern dev server) leverages Yarn’s hoisting to keep its plugin ecosystem lightweight, resulting in a 30 % faster cold start compared to npm workspaces.

6. Comparative Matrix

FeatureLernaNxYarn Workspaces
Primary purposePackage linking & publishingFull‑stack orchestration, caching, affected detectionDependency hoisting & workspace management
Built‑in cachingNoYes (local + optional Nx Cloud)No
Task runnerlerna run (npm scripts)nx run <target> (targets)External (e.g., turbo)
Language supportJS/TSJS/TS + plugins for Go, Rust, etc.JS/TS (via Yarn)
Versioning modesFixed / IndependentFixed / Independent (via nx release)No built‑in versioning (use changesets)
Scalability (packages)Up to ~150 (practical)200+ (tested)Unlimited (depends on CI)
CI/CD impactReduces duplicate installs, but no cachingCan cut CI time by 50‑70 % with affected & cacheFaster installs, but needs extra tooling for incremental builds
Learning curveLow (few commands)Moderate (graph, plugins)Low (just Yarn)
Typical use caseLibrary suites (e.g., Babel)Enterprise apps with many interdependent servicesSimple monorepos or when you already use Yarn

For a platform like Apiary that needs incremental builds for data pipelines, Nx is the most compelling choice. If the team already standardizes on Yarn and wants a lightweight setup, Yarn Workspaces + Turborepo can be sufficient. Lerna remains a solid fallback for pure library publishing workflows.


7. Integrating Monorepo Tooling with CI/CD

7.1 GitHub Actions Example (Nx)

name: CI
on: [push, pull_request]

jobs:
  affected:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v3
      - name: Setup Node
        uses: actions/setup-node@v3
        with:
          node-version: '20'
          cache: 'yarn'
      - run: yarn install --immutable
      - name: Compute affected projects
        id: nx
        run: |
          npx nx affected:apps --base=origin/main --head=HEAD > affected.txt
          echo "apps=$(cat affected.txt)" >> $GITHUB_OUTPUT
      - name: Build affected apps
        if: steps.nx.outputs.apps != ''
        run: |
          npx nx affected:build --base=origin/main --head=HEAD

The nx affected command ensures that only the apps impacted by the PR are built, saving minutes on each run. When paired with Nx Cloud, the cache is automatically restored across GitHub Action runners, further reducing build time.

7.2 CircleCI with Yarn Workspaces

version: 2.1
orbs:
  node: circleci/node@5.0

jobs:
  install:
    executor:
      name: node/default
      tag: '20'
    steps:
      - checkout
      - node/install-packages:
          pkg-manager: yarn
          cache-key: 'yarn-{{ checksum "yarn.lock" }}'
  test:
    docker:
      - image: cimg/node:20
    steps:
      - attach_workspace:
          at: .
      - run: yarn workspaces foreach --parallel run test

yarn workspaces foreach (available in Yarn 3) runs the test script in each workspace concurrently, leveraging all CPU cores. This pattern scales linearly with the number of packages, making it ideal for large test suites.

7.3 Data‑Pipeline Integration (Apiary)

Apiary’s nightly pipeline ingests bee‑population CSVs, transforms them into Parquet files, and trains a TensorFlow model. By placing the ingestion scripts, transformation utilities, and model code in the same monorepo, we can:

  1. Detect changes – If only the CSV schema changes, Nx’s affected graph triggers only the type‑generation and model‑retraining targets.
  2. Cache results – The transformed Parquet files can be cached; subsequent runs skip the heavy ETL step unless source data changes.
  3. Version alignment – The AI agents that consume the model can be version‑locked to the same commit, guaranteeing that the model and its inference code are always compatible.

This tight coupling reduces the risk of “model drift” caused by mismatched data schemas, a common issue in ecological data platforms.


8. Dependency Management & Versioning Strategies

8.1 Fixed vs Independent Versioning

StrategyWhen to UseProsCons
Fixed (single version)Small teams, tightly coupled packagesSimpler release process; single changelogUnnecessary version bumps for unchanged packages
Independent (per‑package)Large ecosystems, public librariesGranular semver; only changed packages get new versionsMore complex release pipeline; multiple changelogs

A hybrid approach is possible: core libraries (e.g., @apiary/common) stay on a fixed version, while consumer apps (@apiary/dashboard, @apiary/api) use independent versions. Nx’s nx release command supports both modes out of the box.

8.2 Peer Dependencies & Hoisting

When multiple packages depend on the same library (e.g., react@18.2.0), Yarn hoists the dependency to the root. However, peer dependencies must be declared explicitly to avoid runtime errors. A typical pattern:

{
  "peerDependencies": {
    "react": "^18.0.0"
  },
  "devDependencies": {
    "@types/react": "^18.0.0"
  }
}

In a monorepo, failing to align peer versions across packages can cause “invalid hook call” errors in React. Nx includes a check-dependencies target that validates peer alignment across the graph, catching mismatches before CI.

8.3 Managing Third‑Party Licenses

Large monorepos often bring in dozens of transitive dependencies. Tools like license-checker (npm) or yarn licenses can generate a consolidated report. Nx provides a nx run-many --target=license-check script that runs the checker in parallel for all packages, producing a single CSV that can be fed into Apiary’s compliance dashboard.


9. Best Practices for Large Teams

  1. Define a clear ownership model – Use a CODEOWNERS file at the workspace level so each package has designated reviewers. This reduces review latency, especially when a change touches multiple packages.
  2. Enforce linting and formatting centrally – Place ESLint, Prettier, and Stylelint configs at the repo root and reference them via extends. Nx can run nx affected:lint on PRs, guaranteeing that only affected packages are linted.
  3. Commit‑time validation – Install a pre‑commit hook (e.g., husky) that runs nx format:check && nx affected:test --base=HEAD~1. This catches errors early without slowing down CI.
  4. Leverage incremental builds – For TypeScript, enable project references ("composite": true) in tsconfig.json. Nx automatically respects these references, allowing the TypeScript compiler to skip unchanged projects.
  5. Cache heavy artifacts – Store Docker images,
Frequently asked
What is Monorepo Tooling about?
In the last decade the scale of software projects has exploded. A single product may now consist of dozens of front‑end applications, a suite of…
What should you know about introduction?
In the last decade the scale of software projects has exploded. A single product may now consist of dozens of front‑end applications, a suite of micro‑services, shared UI libraries, and data‑processing pipelines that all need to evolve together. Managing this complexity in separate repositories quickly becomes a…
What should you know about 1. What a Monorepo Is – and Why It Matters?
A monorepo (short for “mono‑repository”) stores the source code for multiple, often loosely coupled, projects in a single version‑control tree. The alternative, a polyrepo strategy, places each project in its own repository. The monorepo approach brings three core benefits that become measurable at scale:
What should you know about 2.1 Early Days: Simple Scripts and Git Submodules?
Before dedicated tooling existed, teams relied on ad‑hoc bash scripts or Git submodules to share code. While functional for small teams (<10 developers), these approaches suffered from:
What should you know about 2.2 Lerna’s Birth (2015)?
Lerna was created by James Kyle in 2015 to address the “multiple package management” problem for the JavaScript ecosystem. Its first version introduced two key concepts:
References & sources
  1. Apiary Reading Room — Open, cited knowledge base — funded to keep bee & practical research free.
From the Apiary Reading Room. Opinion & editorial — not financial advice. We don't overclaim.
More from the Reading Room