querydrill

Learn ›Aggregation ›The aggregation pipeline

$match, and why it goes first

$gte$lt$match

Of all the stages, this is the one whose position matters most: every document $match removes is a document no later stage has to look at. The syntax itself you already know.

This is basically:

find()

but inside an aggregation pipeline.

Example

Find completed orders:

db.orders.aggregate([
  {
    $match: {
      status: "completed"
    }
  }
])

Input:

Order 1 → completed
Order 2 → completed
Order 3 → completed
Order 4 → pending

Output:

Order 1
Order 2
Order 3

Multiple conditions

db.orders.aggregate([
  {
    $match: {
      status: "completed",
      userId: 101
    }
  }
])

Exactly the same query operators work:

{
  $match: {
    createdAt: {
      $gte: ISODate("2026-01-01"),
      $lt: ISODate("2026-02-01")
    }
  }
}

Important Performance Rule

Usually:

Filter as early as possible.

Bad conceptual pipeline:

100 million documents
        ↓
$group
        ↓
$match

Better:

100 million
   ↓
$match → 10,000
   ↓
$group

Why?

Because later stages process fewer documents.

It matters even more once indexes are involved: a $match at the front of a pipeline can use one, while the same $match sitting after a $group cannot - by then the documents it is filtering did not come from the collection.

Practise this

1 exercise