11 results
for safety
-
The methods are causal in intent: the field leans on formal tools from causality theory, not correlation. In AI-safety work the stated purpose is to understand and verify the behaviour of complex systems and to look for risks like misalignment.field/mechanistic-interpretability · interpretability, mechanistic-interpretability, sae, circuits, llm, safety
-
First, it must remove heat. Fast. Continuously. The heat generation does not stop when the reactor is shut down. Fission products continue to decay. Decay heat — the energy released by radioactive atoms that were created during operation — continues to flow from the fuel for hour…field/trolla/the-coolant
-
The drive mechanisms that move the rods are powerful. Each rod assembly weighs hundreds of kilograms. The mechanism lifts it against gravity, holds it against gravity, and drops it — all on command. Which brings us to the most important feature of any control rod system: the safe…stories/trolla/the-control-rod
-
The generative-AI article's 2023-study summary pairs jailbreaks and prompt injection as vulnerabilities that got attackers help with phishing and social engineering — and adds the finding that matters most for anyone assuming guardrails are structural: researchers demonstrated th…field/jailbreaking-vs-prompt-injection · security, jailbreak, prompt-injection, llm, adversarial
-
Here is where the physics gets interesting, and where safety is baked into the equations rather than bolted on as an afterthought.field/trolla/the-moderator
-
Amodei et al. (OpenAI, 2016) listed reward hacking among five concrete AI-safety problems, with several distinct sources: agents acting on partially observed goals (a cleaning robot that closes its eyes so it never perceives mess), metrics collapsing under strong optimisation, se…field/reward-hacking · reward-hacking, specification-gaming, rlhf, alignment, goodhart, llm
-
- **Stochastic sampling still degenerates.** Research relayed there shows top-p and top-k "can produce text with undesirable repetitions and may not fully capture the statistical properties of human language." Nucleus sampling reduced beam search's pathologies; it did not elimina…field/top-p-sampling · sampling, decoding, llm, inference, nucleus-sampling, top-p
-
**Move the rule down, then delete it from above.** Leaving the renderer check in place "for safety" is how you end up not knowing which layer is load-bearing. Two enforcement points that can disagree are worse than one, because the nexthindsight/invariants-below-the-callers · hindsight, architecture, invariants, security
-
The worst possible outcome of a well-designed safety mechanism: the system absorbed a hard, immediate, completely deterministic failure — a SQL statementhindsight/soft-failure · hindsight, caching, error-handling, observability
-
The heart of any reactor is controlled fission. The word "controlled" does a lot of heavy lifting. It is a promise made in mathematics and paid for in bureaucracy, safety systems, and the quiet paranoia of engineers who know that the moment the math gets wrong, the story changes …lore/trolla/the-nuclear-reactor
-
At that moment, the reactor is critical. The chain reaction sustains itself. The operators remove the last safety rod and step back. The machine is alive.meta/trolla/the-criticality