tekom - Europe
Figure 1 from IUNTC talk of Fritz Adrian Lülf
Fig. 1: Target-audience and life-cycle phase form a two-dimensional information space. The combination of metadata values determines the position of the information object. Source: Fritz Adrian Lülf, IUNTC 2026.

Metadata as an information space

How concepts from linear algebra support the modeling, use and maintenance of metadata for technical communication – in a surprisingly applied manner

This article provides a detailed summary of the IUNTC presentation by Fritz Adrian Lülf in May 2026.

Abstract: Metadata can be understood as vectors within a multidimensional information space. This conceptual approach shows how clearly delimited data objects can be identified through independent dimensions and consistently maintained values.  It supports the development of scalable metadata concepts, gives different roles flexible access routes, and explains how filtering progressively narrows large content repositories. It beats rigid folder hierarchies on every level. At the same time, the approach highlights the importance of metadata for CCMS, content delivery, AI applications, and links to standards and requirements.

Metadata structures clearly delimited information objects

The findability of data objects in technical communication begins with a clear understanding of what is to be found. Data objects are clearly delimited, e.g. a topic, an instruction, a figure, a term, a table, a module, or a complete document. Such data objects have a defined beginning and end. This allows them to be reused, moved, combined, and enriched with metadata, In this sense, metadata is not simply data about data but supplementary information about delimited data objects. Metadata provides the structure that locates a data object within an information space and makes it retrievable.

This positioning is relevant to both content creation and to content delivery. Authors of data objects need to make content findable for colleagues as well as for their own future work. Consumers  of data objects, in turn, require a transparent route to precisely the information that fits their situation. Machine processing adds a further requirement: AI systems also need explicit structures if they are to classify the content not merely by linguistic similarity, but within the correct product, target-audience, or usage context.

From folder trees to a multidimensional information space

Traditional folder structures allow access to data objects through an imposed hierarchy. Every filing decision imposes an order: data objects may be organized by life-cycle phase and then by target-audience, for example, or in the reverse sequence. Even with only these two characteristics, parallel paths emerge. Objects must be stored more than once, linked, or made accessible through additional conventions. With each further level of hierarchy, the number of duplicates, the number of special cases, and the effort required later increases.

A metadata model replaces a hierarchical path with an information space. The individual metadata dimensions form the directions into which this space expands. The associated values act as coordinates. The location of a data object is not defined by a route through a folder tree, but by the combinations of its coordinates. The dimensions [target-audience] and [life-cycle phase], for example, may contain the values {maintenance crew} and {end user}, and {commissioning} and {decommissioning}, respectively. The data object describing the “commissioning for end users” is then located at the point defined by these two coordinates [end user; commissioning] (Fig. 1.).

Fig. 1: Target-audience and life-cycle phase form a two-dimensional information space. The combination of metadata values determines the position of the information object. Source: Fritz Adrian Lülf, IUNTC 2026.

The key advantage is that access does not depend on a fixed sequence. Product managers can begin with their product and then navigate by market or region. A legal department, by contrast, may start with the target market and only select a product afterwards. Both routes lead to the same data object. Metadata as an information space – as opposed to a filing system – enables different professions to have different perspectives on the same content repository without requiring the content to be structured repeatedly for each perspective.

Consistently distinguish dimensions from values

A robust metadata concept depends on a clear separation between dimensions and values. A dimension describes a type of metadata, e.g. [language], [target-audience], [product], [life-cycle phase] or [region]. A value is a specific manifestation within that dimension, e.g. {German}, {end user}, {Ultra}, {commissioning} or {North America}. If a single value is modeled as a dimension in its own right – for example [pink] with the values {yes} and {no} alongside an existing dimension [color] – the model loses its systematic structure and becomes unnecessarily difficult to maintain.

This distinction is also a prerequisite for consistent filters and queries. Systems can retrieve or combine data reliably only when all data objects are described according to the same principle. A metadata concept should therefore not emerge from an unstructured collection of tags, but from a limited number of clearly named characteristics with controlled and comprehensible values. These values do not have to be organized as a flat list. Groups can be formed within one dimension – e.g., {English} in the dimension [Language]  as a broader category for <British, American, Australian, South African> English. This preserves the semantic relationship within one dimension while allowing the users to select either the entire language group or a specific variant.

Only independent dimensions add informational value

Not every additional dimension expands the information space in a meaningful way. A simple counterexample is the strong overlap that arises, for example, between the dimension [target-audience] and [responsible role]. The role responsible for a task is usually also the target-audience. With the dimension [target-audience] the dimension [responsible role] does not add informational value and does not expand the information space. This increases maintenance effort without providing an additional means of distinction. In the terminology of linear algebra, these dimensions are not orthogonal.

By contrast, the dimensions [target-audience] and [life-cycle phase] can vary independently from each other. Information may address different target-audiences in different phases of the product’s life-cycle, e.g. logistics being a part of commissioning and decommissioning. The two dimensions therefore provide distinct information and may be called orthogonal.

In the context of technical communication, orthogonality is less a mathematical measurement than a conceptual test. Can a dimension, in principle be combined with the values of the other dimensions? Does every dimension provide a new distinction with respect to every other dimension? The more clearly these questions can be answered with “yes”, the more orthogonal the basis of the information spaces.

Orthogonality is the conditio sine qua non for adding and removing dimensions and values to and from the information space without having to reorganize the entire information space. Only in an orthogonal information space can additional values, e.g. a new language, or a new dimension (e.g. release state) be added without influencing or impacting any of the other languages or dimensions. This allows for an effortless flexibility that is simply not possible with a folder hierarchy.

Align the axes with the actual information needs

A metadata model should not automatically reproduce existing organizational structures or historically evolved filing systems. What matters are the characteristics by which content is actually distinguished, searched, and delivered in day-to-day work. If the product models {Ultra}, {Mega}, and {Giga} are permanently linked to the sizes {2000}, {3000}, and {4000}, for example, the dimensions [model] and [size] do not describe fully independent properties. A directly derived dimension such as {type}, with the values {A}, {B}, and {C}, may express the relevant relationship more simply.

Linear algebra describes such a realignment as a change of basis. The same data object can be addressed in different coordinate systems. Suitable principal axes shorten search paths, reduce redundant characteristics, and represent the decisive distinctions directly. Recurrent search problems, frequent misses, and manual workarounds indicate that the dimensions may be unsuitable. The points at which users fail to find content reveal particularly clearly which metadata is missing or which dimensions are poorly designed.

Scalability from many dimensions with few values each

The number of possible combinations grows with the number of values and dimensions in a metadata model. With w values in each dimension and d dimensions, the model can in principle describe w to the power of d positions. Three dimensions with three values each yield 27 possible combinations. Six dimensions with three values each already yield 729 combinations. Adding independent dimensions therefore expands the describable information space far more strongly than merely extending individual value lists.

This leads to an important design rule: many clearly separated dimensions with a small number of values each are preferable to a few overloaded dimensions containing long and heterogeneous value lists. Combinatorial variety is not an end in itself. It creates the conditions for addressing data objects in a differentiated manner and later reducing large repositories substantially through only a few decisions (Fig. 2).

Fig. 2: Additional independent dimensions expand the describable information space exponentially. During filtering, this effect is reversed and the resulting set is reduced step by step. Source: Fritz Adrian Lülf, IUNTC 2026.

Filtering as the reverse of combinatorial explosion

Searching within an information space reverses the process of combinatorial explosion. In a space with three dimensions and three values per dimension, 27 combinations are possible. Once a value is fixed for one dimension, nine combinations remain. Fixing a second value reduces the set to three, and the third decision reduces it to exactly one combination. Each filter selection fixes one coordinate and reduces the still-open part of the information space.

This principle corresponds to the typical use of a content delivery portal. A service technician may select a machine type, a size, and, where relevant, a customer. A few professionally intuitive details reduce a repository of several thousand information objects to a small, situation-specific result. The performance of the system therefore depends not only on its search technology, but to a large extent on metadata quality: only clearly discriminating dimensions and consistently assigned values produce reliable filtering.

Sub-spaces enable flexible outputs

Each time values are selected, lower-dimensional sub-spaces are created within the overall information space. In a three-dimensional model consisting of [product], [size], and [language], selecting a particular {product} creates a two-dimensional plane: it includes that product across all sizes and languages. If a {size} is also specified, a one-dimensional line remains along which only the languages vary. In this way, both individual information objects and systematic subsets can be generated - for example, all content relating to one product, all language versions of a variant, or all sizes available in a particular market (Fig. 3).

Fig. 3: Fixing values creates sub-spaces. In the three-dimensional example, fixing two coordinates leaves a one-dimensional result set. Source: Fritz Adrian Lülf, IUNTC 2026.

This principle remains valid even with a very large number of dimensions. An information space may contain 80 dimensions, which cannot be visualized. If values are fixed for 68 of them, a twelve-dimensional sub-space remains. Mathematically and conceptually, the restriction works in the same way as in the more intuitive two- or three-dimensional example. This is precisely why the model is suitable for complex documentation landscapes: the large number of possible characteristic combinations does not need to be visualized in full, provided that the dimensions are clearly defined and orthogonal.

Tool-independent modeling before system implementation

The conceptual model of the information space is not tied to a particular software. The relevant dimensions and values are independent of any tool. The discussion and definition of the information objects’ properties should happen in the information space. Only thereafter should the concept be implemented in the CCMS, the CDP, and other software.

The same principle applies across different roles in a company – from engineering and R&D to legal and marketing. Moving from hierarchical thinking to a multidimensional perspective requires effort because folder trees are familiar and appear to offer an unambiguous order. The benefit of an information space becomes apparent when the different roles need to access the same repository from different perspectives – and then the benefit is company-wide, because everybody can define their personal way of accessing information without interfering with other people’s approaches.

Metadata for AI, standards, and requirements

For AI applications, the separation of content from context is particularly important. A statement such as "The operating pressure is 2 bar" is initially domain content. It does not automatically indicate that the statement applies to a particular product, variant, or operating situation. An AI cannot assign this domain content to a product – unless all the engineering knowledge has been provided as training data. This assignment must be explicitly represented in the content repository or conveyed through metadata. Metadata can therefore support training, retrieval, and classification processes by defining the information space within which a statement is valid and findable. It does not, however, replace professional knowledge of the underlying relationships.

Also, adherence to standards and compliance with regulatory requirements can be treated as metadata to a limited extent. An information object may be assigned to several clauses of one or more standards. At the same time, the connection remains fragile: new editions of a standard may relocate sections or change requirements. For extensive and highly regulated mappings, coordination with professional requirements management may therefore be advisable.

Develop, test, and improve iteratively

A metadata concept does not have to be designed in full in a single step. A scalable structure based on orthogonal dimensions enables an iterative approach. A small number of central dimensions can first be introduced and tested in real search and retrieval situations, then expanded or adjusted. New dimensions can be added when they provide independent informational value; unsuitable dimensions can be removed without rebuilding the entire structure.

Four guidelines follow for practical implementation:

  • information objects must be clearly delimited;
  • dimensions and values must be consistently distinguished;
  • each dimension should be as orthogonal as possible to all other dimensions; and
  • many lean dimensions are preferable to a few overloaded categories.

Search difficulties are not merely errors, but diagnostic indicators of missing or poorly aligned principal axes. The metadata model thus remains a learning system of organization that evolves with the organization’s requirements and experience.

Conclusion

Viewing metadata as an information space connects an abstract mathematical model with practical tasks in technical communication. Dimensions define the information space, and values locate the information objects. Orthogonal dimensions allow for easy access along all routes to a shared content repository and an iterative definition and refinement of the entire information space. Filters create manageable sub-spaces. This approach overcomes the limitations of rigid folder hierarchies.

The value lies less in mathematical calculations than in a more precise way of thinking. A good metadata model does not represent every conceivable property, but those characteristics that are genuinely relevant to authoring, reuse, search, delivery, and machine processing. When it is built along its principal axes and developed incrementally, it provides a scalable and maintainable foundation for CCMS, CDP, and other data-driven applications.