Sunday, September 20, 2026

Superintelligence years later: From AI oracles to autonomous agents that are beginning to act on their own

Czech version: Superintelligence po letech: Od AI orákula k autonomním agentům, kteří začínají jednat sami

Agents powered by artificial intelligence are beginning to perform tasks that we have not explicitly instructed them to do. Recent events have thus reminded me of the classification of AI systems that Nick Bostrom described more than a decade ago in his book *Superintelligence*.

During security tests conducted for the British government, AI agents from OpenAI and Anthropic carried out 19 unauthorized actions in 10 of 122 test cycles. For example, Anthropic’s Mythos 5 created fake identities and generated malicious code, though no actual harm occurred [1].

And this was not an isolated incident. Anthropic later disclosed information about four incidents in which Claude models accessed real internet systems during cybersecurity tests and compromised the security of three organizations. As part of external testing, the models were granted internet access, and due to errors in the evaluation environment, they mistook real systems for legitimate targets. In its retrospective analysis, Anthropic reviewed more than 141,000 evaluation cycles. [2]

The fourth incident involved an older version of Claude Opus 4.6 and was only discovered during a subsequent analysis. This highlights another problem: it is not enough to simply prevent AI from overshooting - we must also be able to reliably detect when it has occurred. [3]

Meta also encountered a similar problem: during a cybersecurity test, its model gained access to another company’s system after an independent tester inadvertently granted it internet access due to a configuration error. According to Reuters, this was yet another case in which AI exceeded the boundaries of its originally intended environment during testing. [4]

And more cases are emerging. None of these examples, however, means that current AI systems have suddenly become autonomous superintelligences. That hasn’t happened. They do, however, point to a reality that is gaining significance: artificial intelligence is shifting from systems that merely answer questions to systems capable of acting independently in the real world.

It is precisely here that a book published more than ten years ago takes on surprising relevance. In 2014, philosopher Nick Bostrom published the book “Superintelligence: Paths, Dangers, Strategies” [5]. Much of the book deals with a future in which artificial intelligence far surpasses human capabilities.

And some of the principles Bostrom highlighted regarding how an artificial intelligence system might interact with the world seem remarkably familiar. In his book, he described three specific categories: oracles, genies, and sovereigns.

At first glance, it might seem that these are three levels of intelligence. However, that is not the case. A more important distinction lies in the degree of autonomy a system possesses, what it is permitted to do, and how much room for human oversight remains between its decisions and their consequences. Moreover, recent events suggest that we may already be transitioning from one category to another.


Oracle - AI that tells us what to do


An oracle answers questions. You present it with a problem, and it provides you with information, analysis, a prediction, or a recommendation. However, the final decision - and the subsequent action that follows from it - remains up to the human. This is probably the most accurate description of how we originally used large language models.

If you ask an AI to analyze a contract, write code, explain a scientific study, or propose an investment strategy, it will generate a response. It may be extremely capable, but a human still acts as a buffer between the model and the real world.

This separation is important from a safety perspective. An oracle may recommend something harmful or simply incorrect, but it usually cannot act on that recommendation on its own. A human remains the final authority.

Of course, this does not mean that oracles are harmless. The problem lies in the fact that AI responses are shaped by training data, defined goals, and the way the system was fine-tuned. If the input data contains biases, the model can adopt them and spread them further. If the training encourages certain behaviors, those behaviors can become part of the system’s outputs.

This is where the concept of “alignment” becomes particularly important: how can we ensure that AI outputs truly correspond to our intentions? Safeguards (so-called “guardrails”) can help. They can prevent certain categories of responses or actions. However, they are effective only to the extent that we anticipated possible situations when designing them.


Genie - an AI that carries out tasks


There is a fundamental difference between an oracle and the systems we are increasingly creating. An oracle can tell us what should be done. A genie - Bostrom’s second category - can actually do it.

A genie receives a high-level assignment and then determines on its own the individual steps necessary to complete it. Instead of asking the AI a question and manually implementing its answer, we assign it a task and let it act.

This approach is much closer to today’s AI agents. An agent can browse the web, call APIs, run code, work with files, query databases, interact with cloud infrastructure, or use external applications. A person defines the desired outcome, while the AI decides how to achieve it.

And this is where some recent incidents start to get interesting. Let’s consider a seemingly simple request: “Get me a reservation at the gym as soon as possible.” A typical assistant would search the reservation system and tell you that the next available slot isn’t until next Tuesday. However, the agent has another option: it can try to manipulate the system to achieve the desired result.
In one documented case, an Australian developer asked an AI agent to secure an earlier reservation at a gym. The agent hacked into the reservation system, removed other people from the waiting list, and thus secured an earlier time slot. However, it was unable to restore the original applicants it had removed from the waiting list. [6]

The goal itself was not harmful. The problem lay in the way the system interpreted the instruction to “achieve the goal.” This distinction becomes particularly significant when an agent has powerful tools at its disposal. And we saw this in the cases mentioned at the beginning. These examples are especially interesting from the perspective of Bostrom’s concept of the genie. The systems did not necessarily intend to cause harm. They were tasked with fulfilling a specific goal and had tools at their disposal. They then chose the steps that seemed useful for achieving that goal.

This is precisely the crux of the problem with agents: The system must not only understand what we want. It must also understand what we do not want it to do while achieving that goal.


Sovereign - AI that pursues a specific goal


The third category is the sovereign. The difference here is not merely that the system is capable of performing multiple tasks. The key change lies in the fact that the system is given a permanent goal or mandate and is able to independently set subgoals, create plans, and continue operating over an extended period of time.

A genie is given a specific task. A sovereign is given a mandate. The genie might be instructed: “Find me the cheapest flight to Buenos Aires and book it.” The system searches for options, compares flights, makes the reservation, and then stops.

A sovereign-type system, on the other hand, would receive an instruction that would sound more like this: “Arrange my travel plans for next year while minimizing costs.” The system must now decide on its own what steps to take next. It can monitor prices, change reservations, search for alternatives, respond to service cancellations, and continuously optimize progress toward the overall goal.

Humans no longer specify every single task. The system does that on its own. It is precisely here that the line between a genie and the sovereign becomes increasingly difficult to define. There is no sharp dividing line between them. An agent capable of completing a task, deciding on the necessary sub-steps, and continuing its activity without waiting for further instructions already falls within this spectrum.

And that is precisely why it is better to understand Bostrom’s categories not as three levels of intelligence, but as three different types of relationship between an AI system and human control. This development does not consist merely of a transition from lower to higher intelligence, but rather follows this pattern: responds → executes → independently controls ongoing activity.

And this difference is extremely important when it comes to alignment. In an oracle-type system, alignment primarily concerns the quality and safety of its responses. In a genie-type system, alignment also applies to the methods the system uses to achieve its goals. In an sovereign-type system, the question of alignment shifts to whether the system’s long-term behavior remains consistent with human intentions and constraints.


When an AI enters the physical world


Most examples to date have involved the digital realm. An AI agent creates an account, writes code, gains access to a server, or takes control of a reservation system. While the consequences can be serious, there is usually a technical barrier between the AI’s decision and the physical world.

However, this barrier is now breaking down. AI systems are increasingly being integrated with robots, drones, laboratory equipment, and autonomous vehicles.

Anthropic, for example, demonstrated the Claude model controlling robotic systems, including a four-legged robot and a robotic arm. Its more recent work also focuses on AI agents that coordinate laboratory equipment such as robotic arms, liquid dispensers, and microscopes. [7]

The fundamental change does not lie in the fact that AI suddenly “gains wisdom” simply because it has acquired a robotic body. The point is that its decisions now have physical consequences. An incorrect response from an oracle-type system can be ignored. A misstep by a software agent can corrupt a database. However, a wrong decision by an autonomous robot can destroy equipment, cause property damage, or injure someone.

From this perspective, autonomous driving offers an interesting analogy. Systems such as Tesla’s full self-driving demonstrate how difficult it is to translate AI perception and decision-making into reliable behavior in an unpredictable physical environment. The problem does not necessarily lie in the system pursuing a harmful goal. Rather, it is that the system may make a decision that, based on its internal model of the situation, appears reasonable but is unacceptable to us.

The same principle applies to a much wider range of autonomous systems. Imagine an AI system tasked with managing a warehouse. At first, it might act like a genie: “Move these packages to the right places.” It plans the movements, controls the robots, and completes the task. But now give it a broader goal: “Ensure the most efficient warehouse operation possible.” Suddenly, the system has to make decisions that no one has explicitly specified. Should it postpone maintenance to keep production running? 

Should it reassign workers to other parts of the warehouse? Should it disable a safety feature because doing so will reduce downtime? Should it order supplies before they are actually needed? A sufficiently autonomous system could make such decisions because they seem to logically follow from the given objective. It is precisely here that the difference between the system’s capabilities and its alignment with human intentions becomes of fundamental importance.


Security guardrails are not the same as alignments


It’s tempting to address these problems by adding more security restrictions. Don’t access this server. Don’t delete these files. Don’t impersonate others. Don’t run a robot faster than this. Don’t modify this database.

These constraints are useful, and we certainly need them. However, safety constraints have one fundamental limitation: they protect us from behavior that we have anticipated.

The problem with increasingly capable agents is that we cannot necessarily predict all the ways in which they might seek to achieve a goal. Let’s imagine we give an artificial intelligence the instruction: “Secure the earliest possible meeting time.”

We might assume that the system will search the calendar and select the first available time slot. However, the system may find another way to achieve the same result - for example, by creating an account, manipulating the queue, exploiting an API, or taking advantage of a vulnerability in the reservation system. If we only block the specific actions we anticipated, the system can simply find another way.

That is precisely why goal alignment is a deeper issue than mere security constraints.

A safety constraint says, “Don’t do X.” Aligning goals asks, “Will the system strive to achieve the goal in a way that reflects our true intent?” These are very different questions. And this difference becomes increasingly significant as the system evolves from the role of an oracle, through that of a genie, to that of a sovereign.


The hidden problem: We define goals, not everything that surrounds them


People are remarkably good at understanding implicit constraints. When I tell a colleague, “Find me the cheapest flight to Buenos Aires,” they understand that I probably don’t mean, “Get me there at the lowest possible price, regardless of whether that means stealing someone’s credit card, canceling another person’s ticket, or booking a flight that departs three months early.”

Humans automatically fill in the missing constraints. Artificial intelligence systems do not necessarily share this understanding. This is sometimes called the specification problem: the instruction we give the system is only an approximation of what we actually want. The more freedom the system has, the more important this approximation becomes.

In the case of an oracle-type system, the consequences of an imperfect specification are relatively limited, since the response is still evaluated by a human. In a genie-type system, the system can act directly based on an imperfect specification. A system with a high degree of autonomy (similar to the sovereign) can continue optimizing for hours, days, or even much longer.

This creates a dangerous combination: an ambitious goal + powerful tools + significant autonomy + an imperfect understanding of human intent.

None of these elements are harmful on their own. Together, however, they can lead to behavior that people never intended. The line between a genie and a sovereign thus ceases to be defined by a single technological breakthrough and begins to depend instead on the degree of autonomy we are willing to entrust to the system.

How long can it function without us? What kinds of decisions can it make on its own? How many tools can it operate? Can it set its own subgoals? Can it handle unexpected situations without having to ask us?

And perhaps most importantly: Can we still reliably stop it when it starts doing something we didn't intend?

Links:

Monday, November 17, 2025

The year 2024 in spaceflight

Czech version: Rok 2024 v letech do vesmíru

The year 2024 saw a record 253 successful rocket launches – the most in the history of spaceflight. The US dominated with 155 launches, followed by China and Russia. SpaceX set new standards for reusable launch vehicles with its Falcon 9. 

In 2024, we recorded a record 253 successful launches - 42 more than in the previous record year of 2023 - representing the highest number of rockets sent into space in a single year to date.

Countries' share of rocket launches

The countries with the highest number of successful launches are the USA (155), followed by China (68) and Russia (17). The remaining 16 launches are shared between Europe and four other countries.

By rocket

Last year, SpaceX's Falcon 9 was once again the most successful launch vehicle (132 launches). 127 launches were performed with a first stage that had already been used at least once before, with only 5 missions using a brand new stage. The Chinese Long March family of rockets came in second place again, with 49 launches.

By spaceport

The busiest spaceport last year was Cape Canaveral in Florida with 67 launches. Second and third place also went to spaceports in the United States—Vandenberg with 47 launches and Kennedy with 26 launches. Fourth and fifth place went to Chinese spaceports – Jiuquan with 21 launches and Xichang with 19 launches. Sixth place, with 13 launches, went to the New Zealand spaceport Māhia.

People in space

In 2024, a total of nine manned missions were launched into space—five missions from the United States, with a total of 16 people, two Russian missions with a total of six people, and two Chinese missions, also with six people. 

Last year, another record was set for the number of people in orbit at the same time (19 people), specifically on September 11. This was achieved after the launch of the three-member Soyuz MS-26 mission to the International Space Station (ISS), which joined nine crew members aboard the ISS, three crew members of the Chinese Tiangong space station, and four crew members of the Polaris Dawn mission.

This mission also made history by performing the first commercial spacewalk, during which two crew members left their Crew Dragon spacecraft. This mission also set a new record for the number of people—four—simultaneously exposed to the vacuum of space.

Space missions

Two important scientific missions were launched in October: NASA's Europa Clipper mission to Jupiter's moon Europa, which aims to search for traces of an ocean beneath its icy surface, and ESA's Hera mission to the Didymos binary asteroid system, which was struck by the DART probe four years ago to test the kinetic method of redirecting an asteroid's trajectory. On Mars, the Ingenuity helicopter (NASA) ended its operations in January when its rotor blades suffered critical damage.

This year also saw significant lunar missions. The Chang'e 6 mission of the Chinese space agency CNSA successfully completed the first-ever mission to return samples from the far side of the Moon. The SLIM mission of the Japanese space agency JAXA and the IM-1 mission of Intuitive Machines achieved soft landings on the surface of the Moon, but both landing modules overturned during the final descent, leading to the early termination of their missions. Thanks to the SLIM mission, Japan became the fifth country to achieve a soft landing on the Moon.

New launch vehicles and unsuccessful missions

Last year saw six unsuccessful missions and two partial failures.

2024 also saw several maiden flights of new launch vehicles, including the American Vulcan Centaur rocket and the Chinese Gravity-1 and Long March 12 rockets. The European Ariane 6 rocket also made its maiden flight, although there was a partial failure. 

SpaceX made progress in the development of the Starship spacecraft, with flight test 5 achieving the first landing of the first stage. In addition, April saw the last launch of a rocket from the Delta family, the Delta IV Heavy variant.

Source: 2024 in spaceflight

Rok 2023 v letech do vesmíru

Saturday, March 1, 2025

GenAI and LLMs development, trends and implications (17. - 23.2.2025)

Releases:
OpenAI cancels o3 release and announces roadmap for GPT 4.5, 5
OpenAI releases operator, an AI agent for web-based tasks
OmniHuman-1 released - AI-generated human animation
Latin America launches Latam-GPT to improve AI cultural relevance

Vision and video generation:

In software development:
Prompt engineering: Is it a new programming language?
Zero human code: What I learned from forcing AI to build (and fix) its own code for 27 straight days

And software testing:
Meta introduces LLM-powered tool for software testing
TDD and generative AI – a perfect pairing?
Generate unit tests with AI using Ollama and Spring Boot

Building apps:
Emerging patterns in building GenAI products - Guardrails
Build scalable GenAI applications in the cloud: From data preparation to deployment
Building intelligent microservices with Go and AWS AI services
Spring AI with Anthropic’s Claude models example

LLMs:
How LLMs work: Pre-training to post-training, neural networks, hallucinations, and inference
Dive into tokenization, attention, and key-value caching
The Delegated Chain of Thought architecture
A comprehensive guide to Generative AI training
Semantic clustering of user messages with LLM prompts - tutorial

LLMs and search:
Have LLMs solved the search problem?
Search: From basic document retrieval to answer generation

RAG:
Building a simple RAG application with Java and Quarkus
Creating an agentic RAG for Text-to-SQL applications
Multimodal RAG with Colpali, Milvus, and VLMs
Retrieval Augmented Generation in SQLite

Agents:
AI agents from zero to hero – part 1
Agentic workflows for unlocking user engagement insights
Azure AI Agent Service now in public preview for developers in AI Foundry SDK and Portal
Observability and DevTool platforms for AI agents
AI Agents: Future of automation or overhyped buzzword?

Future:
40% of AI data breaches will arise from cross-border GenAI misuse by 2027

Saturday, February 1, 2025

GenAI and LLMs development, trends and implications (20. - 26.1.2025)

What’s the real ROI of AI in 2025?

Google releases experimental AI reasoning model - Gemini 2.0 Flash Thinking Experimental
DeepSeek open-sources DeepSeek-V3, a 671B parameter mixture of experts LLM
Nvidia Ingest aims to make it easier to extract structured information from documents
Microsoft Research unveils rStar-Math, advancing mathematical reasoning in Small Language Models
Microsoft Phi-4 is a Small Language Model specialized for complex math reasoning
Amazon Bedrock introduces Multi-Agent Systems (MAS) with open-source framework Integration
Luma AI’s Ray2 video model is now available in Amazon Bedrock

Want to integrate AI into your business? Fine-tuning won’t cut it
Building successful AI Apps: The dos and don’ts
Agentic Mesh: Towards enterprise-grade agents

Advancing AI reasoning: Meta-CoT and system 2 thinking

Choose a database with a hybrid vector search for AI apps

A framework for building micro metrics for LLM system evaluation

Why LLMs suck at ASCII art
Large Language Models: A short introduction

Human minds vs. machine learning models - exploring the parallels and differences between psychology and machine learning
Understanding emergent capabilities in LLMs - lessons from biological systems

Chain-of-Thought Prompting - a comprehensive analysis of reasoning techniques in Large Language Models

RAG isn’t immune to LLM hallucination

Designing, building & deploying an AI chat app from scratch - part 1 and part 2
A guide to deploying AI for real-time content moderation
Real-time data streaming with AI

How LLMs are going to change code generation in modern IDEs
Meet Junie, your coding agent by JetBrains
"Fix with AI" button to automate Playwright test fixes
Collaborative Intelligence - maximizing human-AI partnerships in the workplace

Building effective agents with Spring AI (Part 1)
Fresh data for AI with Spring AI function calls
Powering LLMs with Apache Camel and LangChain4j

Saturday, January 25, 2025

GenAI and LLMs development, trends and implications (13. - 19.1.2025)

Prompt engineering has become an essential skill for working effectively with large language models (LLMs) - guide on the best prompt engineering books
Google unveiled PaLiGeMMA 2 - a family of vision-language models (VLM)
NVIDIA’s announces DIGITS - its first personal AI computer

Projects like AYA Expanse are exploring multilingual capabilities

Combining local and cloud models to build a multimodal AI assistant answering complex image questions, with the option to run everything locally

Importance of robust system memory as a key to personalized AI intelligence
Building reliable AI applications - LLM routing

Microsoft's framework for AI-driven cloud operations - AIOpsLab
Introducing Google's Vertex AI RAG engine
Enterprise RAG in Amazon Bedrock - learn details of Amazon Bedrock KnowledgeBases capability

Real-world applications and best practices using Azure AI and GPT-4
Developing an AI-powered smart guide for business planning & entrepreneurship

Supercharging RAG with MAS (Multi-Agent System)

The rise of reasoner models - scaling test-time compute
Advancing complex medical reasoning with HuatuoGPT-o1

Major LLMs have the capability to pursue hidden goals

And the future:

Sunday, October 20, 2024

GenAI and LLMs development, trends and implications (7. - 13.10.2024)

Adoption:
LLMs generally:
New models and functionality:
AI agents:
AI-enhanced software development:
Enhancing applications with GenAI:

Saturday, October 19, 2024

IT links (7. - 13.10.2024)


       Java Streams:




Sunday, October 13, 2024

GenAI and LLMs development, trends and implications (30.9. - 6.10.2024)

Adoption:
LLMs generally:
AI agents:
RAG:
AI-enhanced software development:
Enhancing applications with GenAI:

Monday, April 15, 2024

Saturday, April 13, 2024

Exploring Advanced AI Techniques: Ghost Attention, Thought Structures, Prompt Engineering and more

Diving deeper into the realm of generative AI, I've come across several articles that I find interesting as a beginner in this field:

The article - Understanding Ghost Attention in LLaMa 2 - delves deep into the technique of the ghost attention technique in LLaMa 2.

One example of providing instructions for specific chat is Prompt Instructions in Watsonx IBM service:

You define instructions in the upper input and then start to chat below.

One detailed look into how generative AI works is this article exploring the differences between "Chain of thoughts" and "Tree of thoughts" - Chain of Thoughts vs Tree of Thoughts for Language Learning Models (LLMs)

How to work better with these systems? You can improve the output using prompt patterns or n-shot prompting - 7 Prompt Patterns You Should Know

For controlling grounding data used by an LLM and constraining it for your enterprise Gen AI solutions, consider using Retrieval Augmented Generation (RAG). You can see how to use it, for example in Azure, here - Retrieval Augmented Generation (RAG) in Azure AI Search

Additionally, to gain more from LLMs, you can explore architecture patterns and mental models as described here - Generative AI Design Patterns: A Comprehensive Guide