The Community Standard Working Group is making steady progress in defining a standard for atmospheric groups that will be the foundation for Habitat's organization data model. We look forward to Habitat organizations being compatible out of the box with the rest of the atmosphere!

Now that we have a clearer picture of how applications will organize spaces within an organization, we've started thinking about how we can maximize the value of storing all your organizational knowledge in one place. In order for both humans and agents to be able to navigate and collaborate, they'll need to be able to search. So, we've added support for search across all of an organization's permissioned records.

Building the search index itself was reasonably simple since we can reuse the spaces sync protocol to keep the index up to date. We spun up a Meilisearch instance and tracked a "cursor" index for each repo's rev and commit hash. However, when a user makes a query, we can't just show them every matching record.

Re-enforcing permissions

With permissioned spaces, each record inherits its permission boundary from the space its in. When querying, we need to filter out records that the user doesn't have permission to. In the opensocial.group proposal, spaces maintain an opensocial.group.access record within them that declares which roles in the group get access. This let's an application create spaces that only "admins" would be able to see. For the search index, this means we just need to include a "roles" field and add the current user's role as a filter.

The more interesting permission enforcement is for Habitat's ReBAC spaces. These spaces can inherit permissions from others and create a graph of permissions. Not only do we need to flatten that graph to the end users when indexing a record, we also need to update all downstream records when a space's permission gets updated. This is the exact use case for Habitat's network.habitat.relationship.* endpoints. We can use network.habitat.relationship.resolveRelations to get the flattened userset for a given space. If we include those in the index, we can just filter by the current user. If a space's permissions are updated, the syncer can call network.habitat.relationship.resolveSpaces to get all downstream spaces that now need to update their records' usersets. 

Now each search query only returns records that a user has permission to. But what does it even mean to search a record? Records can belong to any collection each with its own lexicon of fields, only some of which have searchable text. Let's look at the (simplified) Bluesky post lexicon as an example:

{
  "$type": "com.atproto.lexicon.schema",
  "id": "app.bsky.feed.post",
  "defs": {
    "main": {
      "record": {
        "properties": {
          "createdAt": {
            "format": "datetime",
            "type": "string"
          },
          "reply": {
            "ref": "#replyRef",
            "type": "ref"
          },
          "tags": {
            "items": {
              "type": "string"
            },
            "type": "array"
          },
          "text": {
            "type": "string"
          }
        },
        "type": "object"
      },
      "type": "record"
    }
  },
  ...
}

The "text" and "tags" fields are the only ones we need our search index to query against. However, "createdAt" is still useful in search to sort and filter by. And what do we do with the "reply" ref which is generally just an at:// URI? We need some way to specify how our search index handles each record type.

Search configuration

We're proposing a network.habitat.search.config record type that can be specified per collection type similar to com.atproto.lexicon.schema. The config object will include things like "searchableFields", "filterableFields", "sortableFields" along with other metadata about how a collection type should be handled by search. 

Refs may not be directly searchable but would be extremely useful in ranking algorithms like PageRank. The search config will include metadata about the semantics of refs so that rankers can use that to build a topology of data. 

The search config will also contain information about how a search result is rendered. Since the application can't feasibly build a custom view for every collection type, we'll need a generic framework for how things like the search result "title" should look. 

To bootstrap, we hardcoded a couple search configs into our indexer. Once we finalize the shape of the config record, other lexicon authors can start declaring search configs for their collections just like they do with lexicon definitions. Reusing public atproto broadcast is perfect for search config since it doesn't contain any sensitive info and can be a default that is shared across organizations.

While lexicon authors generally have the best idea of how their collections should be indexed, we also expect organizations to have more specific requirements for their data. Hence the search config should be heirarchical and organizations can override and customize how search results are handled for their org.

This config is a proposed interoperable standard so devs can get their records automatically indexed by Habitat's backend.

Entrypoints

We've finally enabled the search bar at https://home.habitat.network! If you're in an org, you can now see Chalk documents for that org in the search results. We expect Habitat to become a hub for an organization and is a great place for members to start their work via search. 

The network.habitat.search.query endpoint is also available via service auth so any app can reuse it for groups hosted on Habitat. An app can filter results by collection, so they don't need to rebuild search for their app!

Search is a stepping stone for RAG so that Habitat can allow organizations to navigate their data agentically. We added search_records tool to our MCP server so that other agents can also get access to an organization's full context.

We'd love to hear the your thoughts on our search config idea and other thoughts on what it should contain. Join our Discord!