A conversation with my manager recently changed how I think about journaling. We were discussing AI, voice notes, second brains, and how people actually improve at their work.
He made a simple observation: improvement requires reflection. Doing more work is not enough. You need to look back at what happened, understand what worked, and carry that learning into the next attempt.
I think of a reflection loop as checking one finished task, keeping the lesson the evidence supports, and using it when a similar decision returns.
Then he pointed out something about the way I had been working with Claude. I was already doing a version of this. I just wasn’t calling it journaling.
My Obsidian wiki was already a working journal
I use an Obsidian wiki alongside my development workflow. Over time, it has accumulated architectural decisions, implementation notes, failed approaches, project conventions, and reasons behind choices I wanted to preserve.
I don’t treat it like a traditional diary. It is closer to a working record of how I think and why something was done a particular way.
If I abandon an optimization because it creates another problem, I want more than the final code. I want the reason the approach was abandoned. If Claude makes a mistake and we correct it, I want that correction to survive so the same discussion doesn’t start from zero next time.
After working this way, Claude began to feel better at working with me. That does not mean the model itself became more intelligent. The environment around it had become better informed.
Hooks could capture the work as it happens
The next question was whether maintaining this journal had to depend on me remembering every useful detail at the end of the day.
Claude Code hooks run at defined points in a session. They can help build a record around the work. The possible loop looks like this:
Load context → Act → Capture outcome → Reflect → Update the wiki → Use it next time
At SessionStart, a hook could bring in relevant project history. The useful part might be an earlier decision, a convention, or a known issue. Loading the whole vault would make it harder to see what matters.
Before a tool runs, PreToolUse can record the proposed action: which command is about to run or which file is about to change. It does not need access to the model’s private reasoning. The practical question is, what are we about to try?
After the action, PostToolUse can capture the result. Did the command fail? Did tests pass? Did the change expose another constraint? Now the record can distinguish the plan from what actually happened.
At the end of the task, a short review can decide what belongs in the wiki. The Stop event can trigger that check. Simply recording every event, however, does not establish that each one is a durable lesson.
For example, an illustrative review of a component change could record:
Decision: Reuse the existing upload component.
Problem: A replacement broke the mobile layout.
Correction: Preserve the responsive wrapper.
Lesson: Check mobile behaviour before replacing a shared component.
The lesson is more useful than fifty lines of raw logs because it gives the next task something to inspect. It also needs its boundary: one failed replacement is not proof that the component should never be replaced.
Capture gives reflection material. It does not perform the reflection.
My manager’s point was about the next attempt
My manager was describing a broader loop: do something, inspect the result, understand it, and change what you do next. Seen that way, my wiki was performing part of the journaling function for me.
I was not writing a chronological diary after every task. A decision could go into the wiki. A voice note could hold an observation. A hook could preserve an outcome. Claude could help turn a messy record into a shorter note worth reviewing.
The cost of keeping a record was going down while the need to reflect remained. That was the part I had not considered before.
Reflection changes what experience teaches us
There is evidence for the value of pausing to review. In a field experiment at a Wipro call centre in India, trainees who spent the last 15 minutes of the day reflecting scored 22.8% higher on the final training test than a group that kept practising. The result concerns that training setting; it does not give us a predicted gain for every job or every AI workflow.
The useful distinction is simpler. More experience does not automatically produce more learning. People need a way to examine what the experience taught them.
AI research has explored a related mechanism. In the Reflexion paper, agents turned feedback from one attempt into written reflections that informed later attempts. The researchers did this without changing the model’s weights. That is not evidence that my Obsidian setup has the same measured effect. It shows why feedback carried forward as text is a meaningful design choice.
The model can stay the same while the next attempt receives better context.
A useful vault has to evolve, not just grow
This is why I now care less about how much an AI system can remember than about what it should carry forward. Experience → Capture → Reflection → Updated understanding → Next action is the loop I want.
Take reflection out and you mostly have an archive. Dumping every Claude session into Obsidian would add commands, duplicated notes, abandoned decisions, and assumptions that may no longer be true. A future session could retrieve an old conclusion and treat it as current guidance.
Sometimes five debugging notes should become one rule. Sometimes an architectural decision needs to be marked obsolete. Sometimes a note should remain an unresolved question because the evidence never settled it.
What deserves to survive into the next decision? That question needs a review. The vault should evolve, not merely expand.
The same pattern applies outside coding
Companies often evaluate AI one task at a time. Can it summarize a meeting, make a presentation, write code, or draft an email? Those are useful questions, but they miss what happens between tasks.
Imagine two employees using the same model for a year. One begins with mostly fresh context each day. The other has a maintained working wiki with decisions, failed experiments, customer objections, meeting outcomes, and short reflections. This is a comparison, not a claim that either employee achieved a measured result.
They have the same model, but the second person can bring more relevant history into the next decision. They have built a better reflection loop around the tool
This idea is appearing in products too. Anthropic’s Reflect feature for Claude, introduced in July 2026, lets people look back over their usage patterns and consider whether their AI use fits their goals. It is a different kind of reflection from reviewing a coding task, but it asks a related question: what keeps happening, and what should I examine before doing it again?
I still want to make the final judgment
I am happy for Claude to help maintain the journal. Hooks can capture events. Obsidian can preserve the record. AI can help combine several messy notes into one observation and bring back an old constraint when it becomes relevant.
But if the system tells me that I keep making the same mistake, learning does not happen because Claude wrote the sentence. It happens when I change what I do next.
AI may let us write fewer chronological diary entries while preserving more of our working history. That makes reflection more important, not less. Capture, organization, and retrieval can support it. They cannot decide for me which lesson still applies.
Before adding more memory to an AI workflow, open one completed conversation. Find a correction you would not want to repeat. What should the next task check before making that decision again?
Reflection is where a record of experience becomes a change in action.






