Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
|
gpjt's submissions
login
1.
Extending Raschka's GPT-2: an MoE trained from scratch on an RTX 3090
(
gilesthomas.com
)
1 point
by
gpjt
6 days ago
|
past
|
discuss
2.
Putting my Jax-trained models on the Hugging Face Hub
(
gilesthomas.com
)
2 points
by
gpjt
13 days ago
|
past
|
discuss
3.
Why do OpenAI's GPT-2 weights beat mine? Part four: digging into dropout
(
gilesthomas.com
)
1 point
by
gpjt
20 days ago
|
past
4.
First patient to undergo live AI-assisted brain surgery has tumour removed
(
bbc.com
)
6 points
by
gpjt
21 days ago
|
past
5.
Adding diagrams to my static site generator with D2
(
gilesthomas.com
)
3 points
by
gpjt
22 days ago
|
past
6.
Use the built-in GELU, don't roll your own
(
gilesthomas.com
)
1 point
by
gpjt
28 days ago
|
past
7.
A Quick(ish) Chinchilla Check
(
gilesthomas.com
)
1 point
by
gpjt
40 days ago
|
past
8.
I use AI on this blog
(
gilesthomas.com
)
1 point
by
gpjt
47 days ago
|
past
|
1 comment
9.
Why do OpenAI's GPT-2 weights beat mine? Part three: testing overtraining
(
gilesthomas.com
)
2 points
by
gpjt
48 days ago
|
past
10.
Why do OpenAI's GPT-2 weights beat mine? Part two: the bugfix
(
gilesthomas.com
)
8 points
by
gpjt
48 days ago
|
past
11.
Why do OpenAI's GPT-2 weights beat mine?
(
gilesthomas.com
)
4 points
by
gpjt
50 days ago
|
past
12.
Benchmarking Qwen 3.6 35B MoE (3B active) on an RTX 3090
(
gilesthomas.com
)
16 points
by
gpjt
54 days ago
|
past
13.
Building intuition about LLM parameter counts
(
gilesthomas.com
)
2 points
by
gpjt
68 days ago
|
past
14.
Poppy the training box, part 1: the beginnings
(
gilesthomas.com
)
3 points
by
gpjt
70 days ago
|
past
15.
From bigrams to GPT-2, one component at a time (in Jax)
(
gilesthomas.com
)
1 point
by
gpjt
70 days ago
|
past
16.
Building a Jax training loop for an LLM training run
(
gilesthomas.com
)
2 points
by
gpjt
78 days ago
|
past
17.
Thoughts on Role Confusion
(
gilesthomas.com
)
3 points
by
gpjt
84 days ago
|
past
18.
Flax debugging: making a hash of things
(
gilesthomas.com
)
2 points
by
gpjt
3 months ago
|
past
19.
10Gb/s Ethernet: switching to a Broadcom SFP+ module
(
gilesthomas.com
)
195 points
by
gpjt
3 months ago
|
past
|
170 comments
20.
Jax: Commitment Issues
(
gilesthomas.com
)
4 points
by
gpjt
3 months ago
|
past
21.
Jax Back Ends and Devices
(
gilesthomas.com
)
2 points
by
gpjt
3 months ago
|
past
22.
Using Safetensors with Flax
(
gilesthomas.com
)
2 points
by
gpjt
3 months ago
|
past
23.
First Looking into Jax
(
gilesthomas.com
)
3 points
by
gpjt
3 months ago
|
past
24.
10Gb/s Ethernet: using mini-heatsinks with a 10GBASE-T SFP+ module
(
gilesthomas.com
)
3 points
by
gpjt
4 months ago
|
past
25.
10Gb/s Ethernet: what I did to get it working in my home
(
gilesthomas.com
)
232 points
by
gpjt
4 months ago
|
past
|
177 comments
26.
10Gb Ethernet: what I had to (re)learn
(
gilesthomas.com
)
1 point
by
gpjt
4 months ago
|
past
|
1 comment
27.
LLM from scratch, part 33 – what I learned from the appendices
(
gilesthomas.com
)
5 points
by
gpjt
4 months ago
|
past
28.
LLM from scratch (32l) – Interventions: updated instruction fine-tuning results
(
gilesthomas.com
)
1 point
by
gpjt
4 months ago
|
past
29.
How an LLM becomes more coherent as we train it
(
gilesthomas.com
)
3 points
by
gpjt
5 months ago
|
past
30.
LLM from scratch, part 32k – Interventions: gradient accumulation
(
gilesthomas.com
)
2 points
by
gpjt
5 months ago
|
past
More
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: