Prompting and steering AI
I turn an intention into clear questions, instructions, and constraints. I give a model the context it needs, challenge its assumptions, and refine the approach when it misses the point.
Curated capabilities
I curate and develop capabilities by combining human judgment, AI, formal methods, and custom tools into workflows that produce useful, checkable results.
These are human and machine capabilities: ways to think, learn, create, explore, and act with the help of AI and other tools. Prompting, research, software, and visual explanation all play a part. My projects show some of these combinations in use; others are still taking shape.
Intelligence augmentation
I think of this as a practical form of cyborgization: extending my thinking and my ability to act through machines. I bring purpose, context, and judgment. Models help me explore possibilities, and tools help me calculate, build, remember, and check.
I use these combinations to understand, discover, create, and act. I want to keep expanding those abilities as AI advances and changes how people live.
| Part | What it brings |
|---|---|
| Human | Purpose, judgment, context, and responsibility. |
| Machine | Generation, search, memory, calculation, and execution. |
| Together | New ways to think, learn, create, and act, with checks suited to the task. |
Capability map
Here are some of those capabilities in practice, with examples you can explore. The labels distinguish current use, public evidence, and experiments still to be completed.
U: unaided, based on my account unless independently proctored. A: assisted by AI, with the tools and process disclosed. V: finding and correcting errors in results generated by AI. Benchmark scores are reported only after a completed run has been reproduced.
I turn an intention into clear questions, instructions, and constraints. I give a model the context it needs, challenge its assumptions, and refine the approach when it misses the point.
I explore unfamiliar subjects, compare explanations, and connect ideas across philosophy, computation, governance, and economics. Questions and independent checks help me find out where an explanation holds up.
I use writing, diagrams, interactive examples, and software to make difficult ideas easier to understand. I show how an idea works, what supports it, and where it falls short.
I develop research questions with LLM assistance. Models help me explore and recombine possible hypotheses, witnesses, proofs, and designs. I then test the resulting claims with symbolic tools and reproducible experiments.
I guide the investigation using my custom Research Kernel Protocol MCP, LEAP, and PopperPad. Lean, Z3, Tau, and other deterministic tools check formal claims. My scientific practice emphasizes attempts to falsify competing hypotheses. Skills include hypothesis development, abstraction design, experiment design, and evaluation of results generated by AI.
I trace claims back to their sources, look for evidence against them, and record what they depend on. Research Kernel and PopperPad help me keep that record so I can revisit the conclusions.
I use Lean, TLA+, deterministic verifiers, and replay to check precisely stated properties. I explain the assumptions and what those results establish about the system.
I build systems that let models suggest actions, with explicit rules and checks controlling what can run.
Define invariants, construct failure cases, reproduce defects, test tamper/replay boundaries, and distinguish passing tests from justified system claims.
Turn ambiguous quality requirements into testable criteria; find hallucinations, regressions, instruction failures, brittle scoring, and evaluator mistakes.
EVAL-001: build an LLM Regression Sentinel with reference labels reviewed by a person, deterministic checks, failure taxonomy, adversarial tests, and a memo explaining whether the result is ready to use.
I plan to demonstrate how I check messy data, reproduce an analysis, explain a decision using numbers, and identify misleading averages or totals.
I plan to test how consistently I can rate search results, resolve disagreements, and compare my judgments with a reviewed reference set.
Scaling · Development direction
As AI becomes more capable, I want to expand what I can understand, create, and act on. Today I guide the process. Next I want to teach more of it to a coordinated swarm of agents through prompting, tools, and eventually training.
Stronger models and more compute may let me explore more ideas at once. I want to find out which parts of my approach can be passed on to agents, and whether the results hold up when I do.
Describe how I choose questions, build abstractions, look for counterexamples, and assess evidence so agents can follow the process. Specify which decisions require my review.
Let agents explore and implement in parallel, with Research Kernel preserving the research record. Keep deterministic checks and independent evaluation central as the volume of proposals grows.
Compare the process run by agents with the process I guide today. Increase the use of approaches that produce better results under the same budget and checks.
More useful discoveries, creations, and solutions that stand up to independent checks, for the time and compute spent. I also track mistakes, unsupported claims, repeated effort, and what I learn from failed attempts. More text, or more agents agreeing, doesn't tell me that the process is better.
I already use the tools and guide the process myself. Passing the whole process to an agent swarm, or training it into a model, is a research goal. I still need to test how much can be automated reliably and what it costs.
I share the principles and the evidence here. I keep the detailed orchestration and training recipes private.
Evidence standard
Each kind of evidence answers a different question. A Git commit records a version; it doesn't establish whether AI was used. Repeating an experiment checks whether a result can be reproduced. Independent review asks whether the result supports the conclusion.
Where possible, the evidence chain is: task → submission → rubric → verifier → receipt → live defense.
The required artifact exists.
A clean run reproduced the committed result.
The benchmark met its score threshold with no unresolved critical failure.
An identified third party independently evaluated or proctored the work.
Benchmarks and practice
I also use a capability ledger to practice and test specific abilities. Some benchmarks prepare me for paid projects. It keeps one benchmark active until the result meets its completion criteria and can be reproduced.
For career training, I choose the next benchmark by looking at demand and abilities I still need to demonstrate. I also pursue questions, experiments, and creative projects for learning and discovery.
For career benchmarks, look at current roles and the tasks they have in common.
Use AI and other tools to make something small enough to finish and substantial enough to test the relevant abilities.
I review the important judgments, check the scoring, repeat the experiment, and disclose how AI was used. I test the evaluator itself and explain my decisions in a live review.
I publish the result when it meets the evidence requirements. If it falls short, I keep working on the same benchmark.
Explore together
A question to investigate, an idea to make tangible, a tool to build, or a new way to combine human and machine abilities. Start with something worth exploring and a way to tell whether we're making progress.