WEBVTT

1
00:00:00.225 --> 00:00:07.375
You're listening to Peeter Levvuls clearing a two-week Windows XP block with a nineteen-dollar Kimi K3 switch.

2
00:00:07.375 --> 00:00:10.575
Written by Dee-aygo Ferrahro on Promptway.

3
00:00:11.160 --> 00:00:13.547
Peeter Levvuls wanted to install Yahoo!

4
00:00:13.547 --> 00:00:19.172
Messenger from two thousand three inside a Windows XP desktop running in the browser.

5
00:00:19.172 --> 00:00:23.560
Claude Code kept treating the emulator like a cybersecurity problem.

6
00:00:23.860 --> 00:00:24.822
He spent two weeks

7
00:00:24.822 --> 00:00:25.335
bouncing

8
00:00:25.335 --> 00:00:28.197
between model fallbacks and safety blocks.

9
00:00:28.197 --> 00:00:29.397
Then he opened X and

10
00:00:29.397 --> 00:00:29.735
asked

11
00:00:29.735 --> 00:00:29.910
how

12
00:00:29.910 --> 00:00:30.035
to

13
00:00:30.035 --> 00:00:30.272
run

14
00:00:30.272 --> 00:00:32.760
Kimi K3 through a coding agent.

15
00:00:32.760 --> 00:00:33.747
The answer changed

16
00:00:33.747 --> 00:00:33.910
his

17
00:00:33.910 --> 00:00:35.410
afternoon.

18
00:00:35.760 --> 00:00:36.110
Levels

19
00:00:36.110 --> 00:00:36.685
installed

20
00:00:36.685 --> 00:00:37.460
OpenCode,

21
00:00:37.460 --> 00:00:37.935
connected

22
00:00:37.935 --> 00:00:40.485
it directly to Kimi, paid nineteen dollars

23
00:00:40.485 --> 00:00:40.597
for

24
00:00:40.597 --> 00:00:40.660
a

25
00:00:40.660 --> 00:00:41.322
membership,

26
00:00:41.322 --> 00:00:41.735
switched

27
00:00:41.735 --> 00:00:41.897
the

28
00:00:41.897 --> 00:00:42.297
agent

29
00:00:42.297 --> 00:00:42.535
into

30
00:00:42.535 --> 00:00:42.835
Build

31
00:00:42.835 --> 00:00:43.285
mode,

32
00:00:43.285 --> 00:00:43.410
and

33
00:00:43.410 --> 00:00:43.572
let

34
00:00:43.572 --> 00:00:43.660
it

35
00:00:43.660 --> 00:00:43.910
work

36
00:00:43.910 --> 00:00:44.110
through

37
00:00:44.110 --> 00:00:44.247
the

38
00:00:44.247 --> 00:00:44.910
simulator's

39
00:00:44.910 --> 00:00:45.185
to-do

40
00:00:45.185 --> 00:00:45.935
list.

41
00:00:45.935 --> 00:00:46.135
His

42
00:00:46.135 --> 00:00:46.535
report

43
00:00:46.535 --> 00:00:46.797
after

44
00:00:46.797 --> 00:00:46.897
the

45
00:00:46.897 --> 00:00:47.235
switch

46
00:00:47.235 --> 00:00:47.435
was

47
00:00:47.435 --> 00:00:48.022
short:

48
00:00:48.022 --> 00:00:48.635
K3

49
00:00:48.635 --> 00:00:48.860
was

50
00:00:48.860 --> 00:00:49.560
"absolutely

51
00:00:49.560 --> 00:00:50.022
hammering

52
00:00:50.022 --> 00:00:51.710
through" the work.

53
00:00:52.035 --> 00:00:56.272
This is a good founder case study because the task stayed recognizable.

54
00:00:56.272 --> 00:00:58.260
A browser emulator was stuck.

55
00:00:58.260 --> 00:00:59.997
The founder changed the stack.

56
00:00:59.997 --> 00:01:01.710
Work resumed.

57
00:01:01.985 --> 00:01:06.910
It is also a messy model comparison, which is where the useful part begins.

58
00:01:07.160 --> 00:01:13.285
The project K3 walked into The Windows XP simulator sits at pieter.com.

59
00:01:13.285 --> 00:01:23.285
It is the kind of project Levels keeps returning to: old software, browser emulation, a long list of rough edges, and no client waiting for a compliance memo.

60
00:01:23.560 --> 00:01:26.047
He had already been using Claude Code heavily.

61
00:01:26.047 --> 00:01:32.410
In June, he wrote that he had coded almost entirely on a virtual private server with Claude Code for nearly a year.

62
00:01:32.410 --> 00:01:37.235
The switch did not come from a tourist opening two chat tabs and asking for a snake game.

63
00:01:37.235 --> 00:01:42.310
It came from a paying power user who had run into the same refusal pattern for days.

64
00:01:42.585 --> 00:01:51.460
Levels described the immediate problem in public: "Claude Code couldn't do this for 2 weeks." His complaint centered on safety fallbacks.

65
00:01:51.460 --> 00:01:59.810
The model treated requests around the Windows XP environment as risky even though Levels was working on a hobby project he controlled.

66
00:02:00.135 --> 00:02:03.347
The task context changes how I read the refusal.

67
00:02:03.347 --> 00:02:07.497
A guardrail can be reasonable in one environment and maddening in another.

68
00:02:07.497 --> 00:02:10.310
Levels was not asking an agent to probe a bank.

69
00:02:10.310 --> 00:02:11.797
He was trying to make Yahoo!

70
00:02:11.797 --> 00:02:14.435
Messenger work in a browser toy.

71
00:02:14.685 --> 00:02:20.935
He changed four variables, not one The viral version of this story is Kimi K3 beat Claude.

72
00:02:20.935 --> 00:02:24.210
The public record supports a narrower finding.

73
00:02:24.547 --> 00:02:48.747
Layer Before After Model Claude models, with reported safety fallbacks Kimi K3 Harness Claude Code OpenCode Provider path Anthropic through Claude Code Direct Kimi connection after an OpenRouter rate limit Permissions Claude's policy and tool gates OpenCode Build mode with permission bypass enabled Task Windows XP simulator to-do list K3 deserves credit for completing work that had stalled.

74
00:02:48.747 --> 00:02:53.160
Levels also removed several sources of friction around the model.

75
00:02:53.460 --> 00:02:56.247
He published the setup on July 17.

76
00:02:56.247 --> 00:02:57.597
Install OpenCode.

77
00:02:57.597 --> 00:02:59.072
Create a Kimi account.

78
00:02:59.072 --> 00:03:00.385
Pay nineteen dollars.

79
00:03:00.385 --> 00:03:01.847
Get an A P I key.

80
00:03:01.847 --> 00:03:04.172
Connect Kimi Code inside OpenCode.

81
00:03:04.172 --> 00:03:05.410
Switch to Build mode.

82
00:03:05.410 --> 00:03:10.810
His instructions also recommend bypassing permissions, which he already did in Claude Code.

83
00:03:11.135 --> 00:03:14.822
That final setting makes the run faster and less comparable.

84
00:03:14.822 --> 00:03:20.835
A model with broad shell access can finish jobs that a more constrained agent pauses to confirm.

85
00:03:20.835 --> 00:03:24.085
It can also damage more when it guesses wrong.

86
00:03:24.410 --> 00:03:27.435
Levels made a rational trade for a hobby emulator.

87
00:03:27.435 --> 00:03:31.035
I would not copy that trade onto a production database.

88
00:03:31.310 --> 00:03:39.772
Why K3 fit this job Moonshot AI released Kimi K3 on July 16, one day before Levels published his switch.

89
00:03:39.772 --> 00:03:47.285
The model has 2.8 trillion total parameters in a mixture-of-experts architecture, with 104 billion activated for each token.

90
00:03:47.285 --> 00:03:52.360
It accepts a one-million-token context and handles text, images, and video.

91
00:03:52.685 --> 00:03:56.022
The specifications matter less here than the training target.

92
00:03:56.022 --> 00:04:03.347
Moonshot built K3 for long coding sessions, large repositories, terminal tools, and visual feedback loops.

93
00:04:03.347 --> 00:04:07.485
A browser operating-system simulator touches all four.

94
00:04:07.785 --> 00:04:14.035
Moonshot's own technical post says an early K3 build handled most of the team's kernel-optimization work late in development.

95
00:04:14.035 --> 00:04:19.660
The company also reports a 48-hour autonomous chip-design run and a compiler project built from scratch.

96
00:04:19.660 --> 00:04:24.785
Those are vendor case studies, so I treat the measurements as claims until independent teams reproduce them.

97
00:04:24.785 --> 00:04:28.085
They still show what Moonshot tuned the model to attempt.

98
00:04:28.410 --> 00:04:32.960
Levels supplied an outside example with a public project and a named workflow.

99
00:04:32.960 --> 00:04:38.885
On the same day, he also published a macOS 27 browser interface created with K3.

100
00:04:38.885 --> 00:04:41.935
That does not prove the model wins every frontend task.

101
00:04:41.935 --> 00:04:46.035
It does show the visual coding loop was more than a benchmark row.

102
00:04:46.310 --> 00:04:53.410
The cost moved from abstract to nineteen dollars Levels first tried OpenRouter and hit an upstream rate limit.

103
00:04:53.410 --> 00:05:00.785
He then went straight to Kimi, bought the nineteen dollars membership, and used the provider's A P I key with OpenCode.

104
00:05:01.135 --> 00:05:13.260
Moonshot prices the K3 A P I separately at thirty cents per million cache-hit input tokens, three dollars per million cache-miss input tokens, and fifteen dollars per million output tokens.

105
00:05:13.260 --> 00:05:19.360
The company says coding workloads on its official A P I exceed a 90 percent cache-hit rate.

106
00:05:19.685 --> 00:05:22.172
That is vendor-reported cache performance.

107
00:05:22.172 --> 00:05:28.797
Your bill depends on the harness, prompt reuse, context size, and how often the agent rewrites its own plan.

108
00:05:28.797 --> 00:05:34.185
K3 always thinks, and long autonomous runs can burn output tokens quickly.

109
00:05:34.485 --> 00:05:37.847
The relevant founder number remains nineteen dollars.

110
00:05:37.847 --> 00:05:43.185
That was cheap enough for Levels to stop arguing with his old setup and try a new route.

111
00:05:43.435 --> 00:05:48.247
Where the case study gets uncomfortable K3 has its own failure modes.

112
00:05:48.247 --> 00:05:50.960
Moonshot lists three in the release notes.

113
00:05:51.310 --> 00:05:54.422
First, the model expects preserved thinking history.

114
00:05:54.422 --> 00:06:00.447
A harness that drops earlier reasoning, or a mid-session model swap, can make quality unstable.

115
00:06:00.447 --> 00:06:08.210
Moonshot recommends starting K3 in a compatible harness instead of dropping it into a conversation another model began.

116
00:06:08.585 --> 00:06:11.197
Second, the model can act too aggressively.

117
00:06:11.197 --> 00:06:16.635
Moonshot says K3 may make unexpected decisions when intent is ambiguous.

118
00:06:16.635 --> 00:06:22.235
The company recommends explicit constraints in the system prompt or AGENTS.md.

119
00:06:22.585 --> 00:06:29.222
Third, Moonshot concedes that the overall user experience still trails the strongest proprietary models.

120
00:06:29.222 --> 00:06:36.960
A public Kimi Code issue filed after launch also reports the terminal interface hanging during a trivial prompt at maximum effort.

121
00:06:36.960 --> 00:06:40.060
The failure surface changed with the model.

122
00:06:40.435 --> 00:06:47.322
Levels solved one kind of agent friction by choosing a model trained to keep going and a harness configured to let it.

123
00:06:47.322 --> 00:06:53.010
The same combination can turn a vague instruction into a long, expensive mistake.

124
00:06:53.285 --> 00:06:59.285
What I would copy from the switch I would copy the routing decision and leave the permission bypass behind.

125
00:06:59.560 --> 00:07:07.347
When a coding agent refuses a legitimate task twice, write down the task, the exact refusal, the branch state, and the acceptance test.

126
00:07:07.347 --> 00:07:10.310
Start a fresh session in a second harness with a second model.

127
00:07:10.310 --> 00:07:12.985
Give it the same repository and the same test.

128
00:07:12.985 --> 00:07:17.610
Compare the completed diff, elapsed time, tool calls, and regressions.

129
00:07:17.910 --> 00:07:21.272
Do not continue the old conversation after the switch.

130
00:07:21.272 --> 00:07:25.260
K3's own documentation warns against that path.

131
00:07:25.585 --> 00:07:30.297
Keep the second agent inside a disposable branch and a scoped environment.

132
00:07:30.297 --> 00:07:34.285
Levels can rebuild a browser simulator when an agent gets inventive.

133
00:07:34.285 --> 00:07:37.085
Your billing system is less funny.

134
00:07:37.335 --> 00:07:41.635
The founder lesson Levels did not wait for a benchmark committee.

135
00:07:41.635 --> 00:07:45.835
He had a blocked task, spent nineteen dollars, and changed the route.

136
00:07:46.160 --> 00:07:53.735
The result is strong evidence that Kimi K3 belongs in the coding-agent rotation for long, visual, tool-heavy work.

137
00:07:53.735 --> 00:08:01.985
It is weak evidence that K3 is categorically better than Claude because the harness, provider, and permissions changed with the model.

138
00:08:02.259 --> 00:08:09.934
The evidence is enough for a useful operator story: a stuck job, a visible change, and a result someone else can test.

139
00:08:10.259 --> 00:08:12.471
Run K3 on a fresh branch.

140
00:08:12.471 --> 00:08:14.459
Keep the acceptance test fixed.

141
00:08:14.459 --> 00:08:17.959
Keep the permissions narrower than Peeter Levvuls did.

142
00:08:18.520 --> 00:08:19.370
That's the story.

143
00:08:19.370 --> 00:08:22.282
You'll find every source linked at the end of the article.

144
00:08:22.282 --> 00:08:24.945
Thanks for listening, on Promptway.
