AI & LLMs

AI Assets: Organizing Data, Models, and Evaluation Evidence

Treat an AI workflow as a connected collection of data, configuration, tests, and decisions. Learn what to record and how to maintain it.

AI Assets: Map Your Workflow. Pink, cyan, and lavender glass spheres connected to a central cube, with CyberAssets.xyz branding.
Field notes from CyberAssets.xyz

When someone asks where an AI system lives, pointing to a model is rarely a complete answer. The useful result may also depend on a source collection, a labeling guide, a prompt, a retrieval index, a provider configuration, and the tests used to approve changes. An AI asset inventory makes those connections visible.

The goal is practical continuity: another person should be able to understand what a workflow does, identify the material it relies on, and judge whether a proposed change is acceptable. The AI cyber assets overview provides a starting taxonomy. Here is how to turn that taxonomy into records that help you maintain a real workflow.

1. Define the workflow before listing the files

Write a short purpose statement identifying the user, input, output, and decision supported. “Draft a summary of a maintenance report for a human editor” is more useful than “use AI for documents.” The narrower statement tells you which documents belong in scope and who must review the result.

The NIST AI Risk Management Framework is voluntary guidance intended to help incorporate trustworthiness considerations into the design, development, use, and evaluation of AI systems. A practical inference for an inventory is to keep evidence about the surrounding workflow, rather than documenting a model in isolation.

Record the situations the workflow is intended to handle and the situations requiring another route. Identify who maintains the records, who can approve a change, and how a user reports a problem. These can be simple role names within a small team. What matters is that responsibility remains clear when the original builder is unavailable.

2. Separate source data, prepared data, and labels

Source material and prepared material deserve separate entries. A folder of original documents might be transformed into extracted text, cleaned passages, and retrieval records. Keep the transformation steps identifiable so a person can locate the origin of a questionable passage or rebuild a derived collection.

Record collection boundaries

For each collection, record its source, permitted purpose, access rules, revision, and known gaps. Include dates where freshness matters. Describe the population or subject matter covered instead of writing “complete dataset” without a defined boundary. A collection of equipment manuals, for example, may cover only certain models or languages.

Preserve the labeling guide

Labels need their own explanation. If a reviewer assigns categories such as “installation” and “repair,” preserve the definitions and examples used to make that judgment. Record ambiguous cases and how disagreements were resolved. A label file without its labeling guide can hide assumptions that later become hard to detect.

Use stable references between originals, transformations, and labels. If an original must be removed or corrected, those relationships help identify which derived records need attention.

3. Record the model or provider configuration precisely

For deployable model weights, record the specific revision, file identity, tokenizer or other required components, runtime configuration, and applicable license materials. Include any conversion or adjustment made before use. A familiar model name can refer to several materially different configurations.

For a hosted service, the asset record looks different. Record the provider and model identifier used, relevant account settings, request configuration, and the conditions under which information is sent. Keep credentials in an appropriate secret store; the inventory should describe where authorized access is managed without containing the secret itself.

Record known limits as questions your workflow must answer. How are long inputs handled? What happens when a request fails? Is a response checked before another system consumes it? Give those questions owners and tests. The LLM model selection checklist develops the comparison between hosted access and deployable weights, including the operational work each arrangement leaves with the team.

4. Connect the components that shape an answer

An inventory becomes more useful when it shows relationships. A prompt may depend on an output schema. A search index depends on prepared text and an embedding configuration. An evaluation depends on a test set and a grading guide. Connect these entries explicitly instead of relying on nearby folder names.

ComponentRecord alongside itQuestion it helps answer
Retrieval collectionSource revision and preparation stepsWhich material could the answer use?
Prompt templateVersion and input contractWhich instructions were applied?
Model configurationIdentifier and runtime settingsWhich configuration produced the output?
EvaluationTest revision and grading criteriaWhat evidence supported approval?

Choose the level of detail that supports decisions. A small editorial assistant may need a dozen clear records, not a complex cataloging platform. Start with components whose removal or alteration would change the output. Add detail when a real maintenance question reveals a missing connection.

5. Treat evaluation evidence as an asset

An evaluation record should explain the task, the cases used, the judging criteria, the configuration tested, and the observed results. Save representative failures as carefully as successful examples. They help the next reviewer understand what remains unresolved.

Define quality in terms a person can inspect. For a document summary, a checklist might ask whether the output preserves the required dates, includes the main action, avoids unsupported explanations, and stays within the requested length. Decide which failures block use and which are tolerable with editing. Avoid hiding that distinction inside one average score.

Keep development examples separate from cases reserved for checking a proposed release. Otherwise, repeated revisions can become tailored to a familiar set without showing whether the workflow handles new material. Include short, long, incomplete, and conflicting inputs appropriate to the task.

A record of the evaluation procedure is as important as the result. If a human reviewer changes the rubric, document that change. An apparent improvement means little when two versions were judged by different standards without explanation.

6. Make changes traceable and reversible

Create a release record that identifies the approved combination of data, prompts, model configuration, and evaluation. Give that combination a readable version or date-based identifier. Do not assume that preserving only the code preserves everything needed to understand a result.

When something changes, state the reason and the expected effect. A source update may add a new equipment model. A prompt adjustment may require explicit evidence references. A model change may alter response speed or formatting. Test the affected behavior and keep the previous approved configuration available when practical.

Choose retention deliberately. Debugging records can themselves contain sensitive input or output, so collect only what the maintenance task requires and restrict access appropriately. Record removal decisions and the links affected by them. Retention is a workflow choice to resolve, not an excuse to copy every input forever.

Schedule reviews around meaningful events: a source collection changes, a provider changes a supported configuration, a new use case appears, or users report a recurring failure. Calendar reviews can supplement those triggers.

7. Walk through a hypothetical maintenance assistant

Imagine a team building an assistant that drafts internal summaries of workshop maintenance reports. Its approved task is to produce a short draft for an editor, with no automatic change to maintenance schedules. This boundary keeps the expected output and human responsibility explicit.

The inventory contains the report collection, a text extraction process, a terminology list, a summary prompt, a model configuration, and an evaluation set. Each report retains its source identifier. The prompt asks for the equipment name, observed issue, and stated next action, with missing information identified instead of guessed.

A reviewer notices that handwritten notes sometimes disappear during extraction. The team records the failure against the preparation step and adds affected examples to evaluation. Changing the prompt alone would not repair missing input. The dependency map directs attention to the part of the workflow that lost the information.

After revising extraction, the team checks the affected examples and a separate set of reports, then records the approved combination. Our prompt versioning guide explains how the instruction template fits into this same change process.

Conclusion: inventory the evidence behind useful behavior

An AI asset inventory should help you answer three practical questions: what produced this result, what permission and provenance records accompany the inputs, and what evidence supports the intended use? A model identifier answers only part of that inquiry.

Begin with one bounded workflow. Connect its source data, transformations, instructions, configuration, and evaluation. Give uncertain items a visible status and a responsible person. As those records improve, maintenance becomes easier to explain and changes become easier to assess. The broader cyber assets framework places this work alongside other digital resources that depend on access, rights, and reliable operating records.

Follow the ideas

Updated · Our editorial approach

One useful idea leads to another.

Find another angle on the assets, tools, and choices behind your project.

All field guides