Back to home

Curated capabilities

More than I could do alone.

I curate and develop capabilities by combining human judgment, AI, formal methods, and custom tools into workflows that produce useful, checkable results.

These are human and machine capabilities: ways to think, learn, create, explore, and act with the help of AI and other tools. Prompting, research, software, and visual explanation all play a part. My projects show some of these combinations in use; others are still taking shape.

Intelligence augmentation

Human and machine, working together.

I think of this as a practical form of cyborgization: extending my thinking and my ability to act through machines. I bring purpose, context, and judgment. Models help me explore possibilities, and tools help me calculate, build, remember, and check.

I use these combinations to understand, discover, create, and act. I want to keep expanding those abilities as AI advances and changes how people live.

PartWhat it brings
HumanPurpose, judgment, context, and responsibility.
MachineGeneration, search, memory, calculation, and execution.
TogetherNew ways to think, learn, create, and act, with checks suited to the task.

Capability map

Showing is better than telling.

Here are some of those capabilities in practice, with examples you can explore. The labels distinguish current use, public evidence, and experiments still to be completed.

U: unaided, based on my account unless independently proctored. A: assisted by AI, with the tools and process disclosed. V: finding and correcting errors in results generated by AI. Benchmark scores are reported only after a completed run has been reproduced.

Existing evidenceU / A

Creating and explaining

I use writing, diagrams, interactive examples, and software to make difficult ideas easier to understand. I show how an idea works, what supports it, and where it falls short.

Benchmark neededU / A / V

Data analysis & decision support

I plan to demonstrate how I check messy data, reproduce an analysis, explain a decision using numbers, and identify misleading averages or totals.

Benchmark neededU / V

Search relevance & rubric application

I plan to test how consistently I can rate search results, resolve disagreements, and compare my judgments with a reviewed reference set.

Scaling · Development direction

How far can this go?

As AI becomes more capable, I want to expand what I can understand, create, and act on. Today I guide the process. Next I want to teach more of it to a coordinated swarm of agents through prompting, tools, and eventually training.

Stronger models and more compute may let me explore more ideas at once. I want to find out which parts of my approach can be passed on to agents, and whether the results hold up when I do.

Abstract illustration: one source of human direction branches into coordinated agents, whose proposals converge on shared checks before producing evidence artifacts.
Human direction → coordinated agents → shared evidence standards → checked outputs. A plan for future experiments in scaling. View larger.
  1. Transfer the method

    Describe how I choose questions, build abstractions, look for counterexamples, and assess evidence so agents can follow the process. Specify which decisions require my review.

  2. Coordinate more agents

    Let agents explore and implement in parallel, with Research Kernel preserving the research record. Keep deterministic checks and independent evaluation central as the volume of proposals grows.

  3. Measure the gain

    Compare the process run by agents with the process I guide today. Increase the use of approaches that produce better results under the same budget and checks.

What successful scaling would mean

More useful discoveries, creations, and solutions that stand up to independent checks, for the time and compute spent. I also track mistakes, unsupported claims, repeated effort, and what I learn from failed attempts. More text, or more agents agreeing, doesn't tell me that the process is better.

Current status

I already use the tools and guide the process myself. Passing the whole process to an agent swarm, or training it into a model, is a research goal. I still need to test how much can be automated reliably and what it costs.

I share the principles and the evidence here. I keep the detailed orchestration and training recipes private.

Evidence standard

Check your work.

Each kind of evidence answers a different question. A Git commit records a version; it doesn't establish whether AI was used. Repeating an experiment checks whether a result can be reproduced. Independent review asks whether the result supports the conclusion.

Where possible, the evidence chain is: task → submission → rubric → verifier → receipt → live defense.

Completed

The required artifact exists.

Replayed

A clean run reproduced the committed result.

Ready to publish

The benchmark met its score threshold with no unresolved critical failure.

Externally verified

An identified third party independently evaluated or proctored the work.

Benchmarks and practice

One step at a time.

I also use a capability ledger to practice and test specific abilities. Some benchmarks prepare me for paid projects. It keeps one benchmark active until the result meets its completion criteria and can be reproduced.

For career training, I choose the next benchmark by looking at demand and abilities I still need to demonstrate. I also pursue questions, experiments, and creative projects for learning and discovery.

  1. 1. Observe

    For career benchmarks, look at current roles and the tasks they have in common.

  2. 2. Build

    Use AI and other tools to make something small enough to finish and substantial enough to test the relevant abilities.

  3. 3. Verify

    I review the important judgments, check the scoring, repeat the experiment, and disclose how AI was used. I test the evaluator itself and explain my decisions in a live review.

  4. 4. Graduate

    I publish the result when it meets the evidence requirements. If it falls short, I keep working on the same benchmark.

Explore together

What could we do together?

A question to investigate, an idea to make tangible, a tool to build, or a new way to combine human and machine abilities. Start with something worth exploring and a way to tell whether we're making progress.