Supporting relations with Umbraco Search

Supporting relations with Umbraco Search

Umbraco Search does not support content relations out of the box. This is a conscious design decision, because relations can be very expensive to maintain from an indexing perspective.

But of course, this can be changed with a bit of code 🤓

You’ll find a full working site in this GitHub repo. As per usual, the admin credentials to the Umbraco backoffice are:

  • Username: admin@localhost
  • Password: SuperSecret123

Indexing the document picker

For document picker values, Umbraco Search indexes only the picked document ID (key) for filtering purposes. This works without relation handling, because the ID is immutable.

But what if the picked document name needs indexing too, so it can be found with textual search?

Well, the actual indexing is pretty straightforward with a property value handler:

public class ContentPickerPropertyValueHandler : IPropertyValueHandler
{
    private readonly IContentService _contentService;

    public ContentPickerPropertyValueHandler(IContentService contentService)
        => _contentService = contentService;

    public bool CanHandle(string propertyEditorAlias)
        => propertyEditorAlias is Constants.PropertyEditors.Aliases.ContentPicker;

    public IEnumerable<IndexField> GetIndexFields(IProperty property, string? culture, string? segment, bool published, IContentBase contentContext)
    {
        // The document picker stores an UDI, so let's parse that.
        var value = property.GetValue(culture, segment, published) as string;
        if (value.IsNullOrWhiteSpace()
            || UdiParser.TryParse(value, out Udi? udi) is false
            || udi is not GuidUdi guidUdi)
        {
            return [];
        }

        // Get the content name for indexing.
        var contentName = GetContentName(guidUdi.Guid, culture, published);
        
        // Index the content name as text for free text search.
        // Retain the default filtering ability by including the content key as a keyword.
        return contentName.NullOrWhiteSpaceAsNull() is not null
            ? [
                new IndexField(
                    property.Alias,
                    new IndexValue
                    {
                        Keywords = [guidUdi.Guid.AsKeyword()],
                        Texts = [contentName!],
                    },
                    culture,
                    segment)
            ]
            : [];
    }

    private string? GetContentName(Guid key, string? culture, bool published)
    {
        // Omitted for brevity.
    }
}

With that in place, a document can be searched by the name of a picked document. At least… the name of the picked document at index time 🫣

Keeping relations up to date

I’ve previously described how you can use a notification handler to trigger reindexing of related content. While this is a perfectly valid approach, Umbraco Search actually has an extension point that’s quite appropriate for this particular use case; the content change strategy.

The content change strategy is responsible for figuring out what to do when any given content changes.

A search index is tied to a specific content change strategy, but Umbraco Search can run any number of content change strategies in parallel. Umbraco Search ships two content change strategies - one for draft content, and one for published content.

I do not recommend implementing your own from content change strategy scratch. They can get insanely complex, particularly when you start factoring in published state tracking in multi-lingual publishing paths 😱

Instead, you should delegate your own custom change handling logic to the default content change strategies:

public class MyContentChangeStrategy : IContentChangeStrategy
{
    private readonly IPublishedContentChangeStrategy _delegateContentChangeStrategy;

    public MyContentChangeStrategy(IPublishedContentChangeStrategy delegateContentChangeStrategy)
        => _delegateContentChangeStrategy = delegateContentChangeStrategy;

    public async Task HandleAsync(IEnumerable<ContentIndexInfo> indexInfos, IEnumerable<ContentChange> changes, CancellationToken cancellationToken)
    {
        // Figure out which effective changes should be handled.
        var effectiveChanges = await CalculateEffectiveChanges(changes, cancellationToken);

        // Delegate the actual handling to the change strategy.
        await _delegateContentChangeStrategy.HandleAsync(indexInfos, effectiveChanges, cancellationToken);
    }

    public async Task RebuildAsync(ContentIndexInfo indexInfo, CancellationToken cancellationToken)
        // Just delegate this entirely to the delegate change strategy.
        => await _delegateContentChangeStrategy.RebuildAsync(indexInfo, cancellationToken);

    private async Task<IEnumerable<ContentChange>> CalculateEffectiveChanges(IEnumerable<ContentChange> changes, CancellationToken cancellationToken)
    {
        // Here you'll figure out what the effective changes are, based on the changes flagged by Umbraco Search.
    }
}

In my scenario, I’m looking to reindex related content when something changes, so I do need a few extra bells and whistles:

  1. The ITrackedReferencesService for resolving content relations.
  2. The IIndexDocumentService for clearing out any database cache for index values, because I’ll be triggering additional reindexing outside the default Umbraco Search handling.

Lastly, I want this implementation to cover both draft and published content, so it’ll be an abstract base implementation (ReindexRelationsContentChangeStrategyBase) for two concrete content change strategies, each supporting their own ContentState (draft or published) 👀

Now, to keep all this at a reasonable scope, I’ve also added a few limitations:

  1. I’ll only support document-to-document references.
  2. Anything past 1000 references per item is explicitly ignored.

With that, the effective change calculation looks like this:

// Added fields:
private readonly ITrackedReferencesService _trackedReferencesService;
private readonly IIndexDocumentService _indexDocumentService;
private readonly ContentState _supportedContentState;

// Concrete implementation:
private async Task<IEnumerable<ContentChange>> CalculateEffectiveChanges(IEnumerable<ContentChange> changes, CancellationToken cancellationToken)
{
    // Get the document IDs of all relevant changes (that is, changes in the supported content state).
    var changesAsArray = changes as ContentChange[] ?? changes.ToArray();
    var documentIds = changesAsArray
        .Where(change => change.ObjectType is UmbracoObjectTypes.Document
                         && change.ContentState == _supportedContentState)
        .Select(change => change.Id)
        .ToArray();

    if (documentIds.Length is 0)
    {
        // No changes to handle here; short-circuit and let the delegate change strategy handle the changes.
        return changesAsArray;
    }
    
    var relatedDocumentIds = new List<Guid>();
    foreach (var documentId in documentIds)
    {
        // Get all relations for the document
        // NOTE: For simplicity we just fetch the first 1000 relations here. If you expect to have content with
        //       more relations than that, consider iterating through the relations with proper pagination.
        var references = await _trackedReferencesService.GetPagedRelationsForItemAsync(
            documentId,
            UmbracoObjectTypes.Document,
            0,
            1000,
            true);
        if (references.Success)
        {
            relatedDocumentIds.AddRange(references.Result.Items.Select(item => item.NodeKey));
        }
    }

    if (relatedDocumentIds.Count is 0)
    {
        // No relations found; short-circuit and let the delegate change strategy handle the changes.
        return changesAsArray;
    }

    // Play nice; check for cancellation before continuing.
    if (cancellationToken.IsCancellationRequested)
    {
        // Cancellation was requested; short-circuit and let the delegate change strategy deal with it.
        return changesAsArray;
    }
    
    // Clear the cached index values to force a rebuild of the document index values.
    await _indexDocumentService.DeleteAsync(relatedDocumentIds.ToArray(), _supportedContentState is ContentState.Published);
    
    // The effective changes to handle are the original input changes plus the related documents.
    return changesAsArray
        .Union(relatedDocumentIds
            .Except(documentIds)
            .Select(documentId =>
                ContentChange.Document(documentId, ChangeImpact.Refresh, _supportedContentState)
            )
        );
}

And the concrete implementation of the abstract base look like this:

// For handling draft content changes:
public class ReindexRelationsDraftContentChangeStrategy : ReindexRelationsContentChangeStrategyBase
{
    public ReindexRelationsDraftContentChangeStrategy(
        ITrackedReferencesService trackedReferencesService,
        IDraftContentChangeStrategy publishedContentChangeStrategy,
        IIndexDocumentService indexDocumentService)
        : base(trackedReferencesService, publishedContentChangeStrategy, indexDocumentService, ContentState.Draft)
    {
    }
}

// For handling published content changes:
public class ReindexRelationsPublishedContentChangeStrategy : ReindexRelationsContentChangeStrategyBase
{
    public ReindexRelationsPublishedContentChangeStrategy(
        ITrackedReferencesService trackedReferencesService,
        IPublishedContentChangeStrategy publishedContentChangeStrategy,
        IIndexDocumentService indexDocumentService)
        : base(trackedReferencesService, publishedContentChangeStrategy, indexDocumentService, ContentState.Published)
    {
    }
}

Two things now remain to complete the puzzle:

  1. Registering the custom content change strategies in the service collection
  2. Re-registering the default document indexes to utilize the custom change strategies.

Both of these fit well in a composer:

public sealed class SiteComposer : IComposer
{
    public void Compose(IUmbracoBuilder builder)
    {
        // Register the custom change strategies.
        builder.Services.AddSingleton<ReindexRelationsPublishedContentChangeStrategy>();
        builder.Services.AddSingleton<ReindexRelationsDraftContentChangeStrategy>();

        // Re-register the default content indexes to use the custom content change strategies.
        builder.Services.Configure<IndexOptions>(options =>
        {
            options.RegisterContentIndex<IIndexer, ISearcher, ReindexRelationsPublishedContentChangeStrategy>
            (
                Constants.IndexAliases.PublishedContent,
                UmbracoObjectTypes.Document
            );
            options.RegisterContentIndex<IIndexer, ISearcher, ReindexRelationsDraftContentChangeStrategy>
            (
                Constants.IndexAliases.DraftContent,
                UmbracoObjectTypes.Document
            );
        });
    }
}

Now when a document changes, any other document with a relation to the changed document will be queued for reindexing alongside the changed one. Therefore, documents remain searchable by the names of their picked documents, even when these change over time 🙌

And… breathe!

Well, that was a mouthful 😯

If you’re feeling a bit confused or exhausted by now, don’t fret. Shy of a full-blown search provider, the content change strategy is probably the most complex extension point of Umbraco Search.

While it’s not part of the official docs (yet), it’s still a powerful concept. It’s also one of those things that can seriously hamper your server performance when misused.

As I mentioned in the beginning, relation maintenance at index time can be quite the workload, so please use it with care 🙏

Also; if you missed it, Umbraco Search becomes part of the CMS in Umbraco 19.

Until next time,

Happy indexing 💜