Showing posts with label Distributed System. Show all posts
Showing posts with label Distributed System. Show all posts

Thursday, 15 January 2026

Audacious Plan to Train AI in Space

What happens when humanity's appetite for artificial intelligence outpaces our planet's ability to feed it? Someone decided the answer was "leave the planet."

In December 2025, something delightfully absurd happened 325 kilometers above Earth. A 60-kilogram satellite named Starcloud-1, carrying an NVIDIA H100 GPU, trained an AI model on the complete works of Shakespeare. Model learned to speak in shakespeare English while orbiting our planet at 7.8 kilometers per second.

"To compute, or not to compute"—apparently, in space, the answer is always "compute."

This wasn't a publicity stunt. It was a proof of concept for what might be the most bold infrastructure play in the history of computing: moving AI training off Earth entirely.

First reaction: "This is insane. I should buy more chip stocks !"

Second reaction, after reading the physics: "Wait, this might actually work."





Dirty Secret Nobody Wants to Talk About at AI Conferences

Here's something the AI industry prefers to whisper about over drinks rather than announce on keynote stages: we're running out of power. Not in some distant climate-apocalypse scenario. Now. Today. While you're reading this.

The numbers read like a horror story for grid operators:

Data centers consumed approximately 415 terawatt-hours of electricity globally in 2024—roughly 1.5% of all electricity generated on Earth. By 2030, that figure is projected to more than double to 945 TWh. That's Japan's entire annual electricity consumption. For computers. Training models to argue about whether a hot dog is a sandwich.

Virginia's data centers alone consume 26% of the state's electricity. Imagine explaining to your neighbors that their brownouts are because someone needed to train a chatbot to write better cover letters.
But here's where it gets truly uncomfortable: to train the next generation of frontier models—think GPT-6 or whatever Claude's grandchildren will be called—we'll need multi-gigawatt clusters. 

5 GW data center would exceed the capacity of the largest power plant in the United States. These clusters don't exist because they can't exist with current terrestrial infrastructure.

Breakthrough is not coming from earth solar panel and fusion reactors is 10 to 20+ Year.

It might come from the one place where solar works really, really well.

325 kilometers straight up.

The Physics of Space: Nature's Cheat Codes

Starcloud's white paper makes a case that initially sounds like venture capital science fiction. But then you check the physics, and... huh. It actually works. Let me break down as per paper why space is basically running a different game engine than Earth.

Cheat Code #1: Infinite Solar Energy (Seriously)

Solar panels in Earth orbit receive unfiltered sunlight 24/7. No atmosphere absorbing photons. No weather. No pesky night cycle if you pick the right orbit. A dawn-dusk sun-synchronous orbit keeps a spacecraft perpetually riding the terminator line between day and night—eternal golden hour, but for electricity.

The capacity factor of space-based solar exceeds 95%, compared to a median of 24% for terrestrial installations in the US. The same solar array generates more than 5X the energy in orbit than it would on your roof.  This seems like coding gains we get from claude code :-) 


Cheat Code #2: The Universe's Free Air Conditioning

Deep space is cold. Like, really cold. Cosmic microwave background sits at approximately -270°C. A simple black radiator plate held at room temperature will shed heat into that infinite cold at approximately 633 watts per square meter.

Cooling algorithm is very different on earth vs space.

Earth Cooling:
Evaporative cooling towers consuming billions of gallons of water. Chillers running 24/7. 
Microsoft literally sinking servers in the ocean like some kind of tech burial at sea.

Space Cooling:

Point a black plate at the void. Wait. Physics does the rest. No water. No chillers.

Just thermodynamics being thermodynamic.


Cheat Code #3: No Land Law in Orbit

Perhaps the most underrated advantage. On Earth, large-scale energy and infrastructure projects routinely take a decade or more to complete due to environmental reviews, utility negotiations, zoning battles, and that one guy at every town hall meeting who's convinced 5G causes migraines.

In space? You dock another module and keep building. When xAI had to resort to natural gas generators for their Memphis cluster because the grid wasn't ready, they weren't just solving a technical problem—they were demonstrating the bureaucratic fragility of terrestrial infrastructure.

What does Cost Math look like

Starcloud's white paper presents this comparison for a 40 MW data center operated over 10 years:

Terrestrial 10-Year Cost:
Energy: ~$140M (@$0.04/kWh)
Land, permits, cooling infrastructure
Water, maintenance, grid upgrades
Total: ~$167 million+
Space 10-Year Cost:
Solar array: ~$2M
Launch: ~$5M (next-gen vehicles)
Radiation shielding: ~$1.2M
Total: ~$8 million

That's a 20x difference, driven almost entirely by energy costs.

Now, before you start a space data center SPAC, let's be honest about what this analysis conveniently ignores: the actual compute hardware. 40 MW of GPU capacity costs somewhere in the neighborhood of $12-13 billion. That's... a lot of billions.

But here's the thing: you pay for that hardware whether it's sitting in a concrete bunker in Iowa or floating above the atmosphere. The operational cost delta remains. And as models get larger and training runs stretch from weeks to months, that delta compounds like the most patient venture capitalist in history.

What This Means for Everyone Betting Billions on the Ground Datacenter

If orbital data centers become economically viable at scale.

Hyperscaler Dilemma

 Microsoft, Google, Amazon, and Meta have collectively committed over $200 billion in capital expenditure on terrestrial data center infrastructure. These are sunk costs with multi-decade payback periods. Do they pivot to space and write off billions? Do they wait and risk being leapfrogged? The prisoner's dilemma dynamics here are brutal. Someone will defect first.

Sovereign AI Gets Complicated

Countries racing to build domestic AI capabilities have assumed the limiting factor is talent and chips. If it turns out the limiting factor is energy, and the solution is orbital infrastructure, the competitive landscape shifts dramatically. Quick: who controls orbital launch capacity? Who can deploy and maintain space-based infrastructure? These aren't questions most national AI strategies have seriously considered.

Environmental Narrative Flips

Right now, AI's carbon footprint is a vulnerability—a PR problem and increasingly a regulatory target. Orbital data centers, powered entirely by solar energy and requiring no water for cooling, transform AI infrastructure from environmental liability to potential climate solution. That's a narrative shift worth billions in avoided regulatory friction alone.


Design Principles That Actually Matter

What makes Starcloud's approach interesting isn't just "put computers in space"—it's how they're thinking about building something that can survive and scale in a hostile environment. 

There's some genuine distributed systems wisdom here:

Modularity: Everything is designed to be added, replaced, or abandoned independently. No single-point-of-failure architecture. This is microservices thinking applied to hardware, which is either brilliant or terrifying depending on your ops experience.

Incremental Scalability: You don't build a 5 GW space station and pray it works. You launch 40 MW modules, validate they function, scale up. It's the same philosophy that made AWS successful: don't bet everything on one deployment.

Failure Resiliency: In space, you can't send a technician. Components will fail. The system has to route around damage like the internet was originally designed to route around nuclear attacks. Graceful degradation isn't optional—it's existential.

Ease of Maintenance: Or rather, the complete absence of it. Everything has to be either radiation-hardened enough to outlast its usefulness, or cheap enough to abandon. There's no middle ground.


Wednesday, 23 July 2025

Impossible Dream: How facebook $85 Server Conquered Half the Planet

Story of engineering excellence in the age before clouds, and the distributed computing lessons that changed everything



You can read this POST with AI also 
ChatGpt  Perplexity Claude

Chapter 1: The Dorm Room That Shook the World

It was February 4th, 2004, and Mark Zuckerberg had a problem. His $85-per-month server was melting.

Twenty-four hours earlier, he'd launched "The Facebook" from his Harvard dorm room—a simple PHP application running on a basic LAMP stack that any college student could understand. Now, 1,200 Harvard students were frantically refreshing pages, checking profiles, and poking each other with an enthusiasm that was literally breaking the internet.

"We need more servers," someone said, watching the CPU usage spike to 100% and stay there.

But Zuckerberg's roommate, a computer science student, had a different idea. "What if we don't need bigger servers? What if we need smarter architecture?"

And with that question, one of the greatest distributed computing adventures in history began.

The First Distributed Computing Lesson: When you can't scale up, scale out—but do it intelligently.


Chapter 2: University Student aha moment 

By March 2004, Facebook was expanding to other Ivy League schools. The obvious solution was to throw all users into one massive database and hope for the best. But the engineering team noticed something fascinating: Harvard students mostly talked to other Harvard students. Yale students connected with Yale students. Princeton was its own social universe.

"Why are we fighting human nature?" asked one engineer, staring at server logs. "Let's work with it instead."

They made a decision that would echo through distributed computing history: separate database instances for each university. Not because it was technically elegant, but because it matched how humans actually behaved.

The results were magical. Database queries that once crawled across massive datasets now zipped through smaller, focused collections. Server load distributed naturally. The architecture scaled not through brute force, but through understanding.

Months later, when Facebook was serving dozens of universities with blazing speed, that engineer would realize they'd stumbled upon something profound: the best distributed systems aren't the ones that fight reality—they're the ones that embrace it.

The Second Distributed Computing Lesson: Design your data architecture around user behavior patterns, not technical elegance.


Chapter 3: Sharding Revolution

By 2005, Facebook faced its first existential crisis. Students were no longer staying within their university bubbles. They wanted to connect with high school friends at different colleges, summer camp buddies scattered across the country, and family members everywhere.

The beautiful university-based system was crumbling.

"We need to completely rethink databases," announced the head of engineering in a meeting that would last many hours. "If we can't keep users separated by school, we'll separate them by... randomness."

The room fell silent. Random database sharding? It sounded insane.

But they were desperate, and desperate times call for revolutionary thinking. They embarked on the most audacious database experiment of the early internet: random sharding across thousands of MySQL instances, with shard IDs embedded in every piece of content.

The catch? They had to eliminate cross-database JOINs entirely. Every query that once elegantly connected related data across tables now had to be redesigned from scratch.

"It's like rebuilding a skyscraper while people are living in it," muttered one engineer, refactoring the friends system for the hundredth time.

But it worked. By 2009, Facebook was processing 200 billion page views monthly using this revolutionary approach. They'd proven that with enough engineering creativity, traditional databases could support planet-scale applications.

The Third Distributed Computing Lesson: When existing technologies don't fit your scale, don't accept their limitations—reimagine them entirely.

Chapter 4: Great Cache Stampede

2007 brought a new nightmare: the cache stampede.

Picture this: A popular piece of content expires from cache simultaneously across thousands of servers. Suddenly, thousands of database queries slam the backend all at once, creating a cascading failure that brings down the entire system.

"It's like everyone in a theater trying to exit through the same door,"  They needed traffic control.

Enter the "leases" system—one of the most elegant solutions in caching history. When cache data expired, only one server got permission (a "lease") to fetch fresh data from the database. Everyone else waited patiently for the result.

But that was just the beginning. Facebook's caching infrastructure became a marvel of distributed computing:

  • 1 billion cache requests per second flowing through custom-optimized memcached
  • 521 cache lookups for an average page load, orchestrated with surgical precision
  • UDP optimization that squeezed every microsecond of performance from the network
  • Regional cache pools that shared data across continents

One engineer summed it up perfectly: "We didn't just build a cache. We built the world's largest memory bank, and every byte in it had to be exactly where it needed to be, exactly when it needed to be there."

The Fourth Distributed Computing Lesson: At planet scale, caching isn't optimization—it's fundamental architecture.

Chapter 5: PHP Impossibility

2008 brought the PHP crisis.

Facebook was serving hundreds of millions of users with a programming language that parsed and executed code from scratch on every single page request. It was like rebuilding a car engine every time you wanted to drive to the grocery store.

"PHP is killing us," said someone, staring at server utilization charts.

Traditional solutions like opcode caching helped, but Facebook needed something revolutionary. What they came up with sounded like science fiction: compile PHP to C++ and then to native machine code.

The HipHop project started as a weekend hackathon experiment. "Let's see if we can make PHP as fast as C++," said one engineer, probably not realizing they were about to rewrite the rules of web development.

The results defied belief:

  • 50% CPU reduction immediately upon deployment
  • 90% of Facebook's traffic running on their custom PHP compiler by 2010
  • 70% more traffic served on the same hardware

They had essentially created a new programming language that looked like PHP but performed like compiled code. It was engineering audacity at its finest.

The Fifth Distributed Computing Lesson: Don't let programming language limitations define your performance ceiling—rewrite the language if you have to.


Chapter 6: Impossible Geography

As Facebook exploded globally, they faced a puzzle that kept engineers awake at night: How do you serve users in Japan, Brazil, and Germany with the same millisecond responsiveness as users in California?

The solution was a masterpiece of distributed systems thinking: geographic replication with intelligent routing.

West Coast servers became the "source of truth"—all writes happened there. But reads could happen anywhere, served by mirror databases that synchronized with the masters through carefully orchestrated replication.

Here's where it got really clever: When you updated your status, Facebook set a special cookie that ensured you'd see your own changes immediately (served from the West Coast), while your friends around the world would see the update within seconds as it propagated through the global infrastructure.

"It's like having a conversation that happens simultaneously in multiple time zones," explained one engineer. "Everyone hears you speak in real-time, even though the sound waves take different amounts of time to reach each person."

By 2010, this system was handling users across six continents with response times that felt local everywhere.

The Sixth Distributed Computing Lesson: Global consistency is less important than local performance—design for eventual consistency with smart routing.

Chapter 7: Security Paradox

2009 brought Facebook's most dangerous challenge yet: securing 300 million users with no blueprint to follow.

Modern authentication systems, OAuth, and sophisticated security frameworks simply didn't exist. Facebook had to build everything from scratch while hackers around the world tried to break in.

"We're writing the security playbook for the planet-scale internet," said the security chief, "and we're doing it while under attack."

Their solutions became legendary:

  • Distributed session management using their own cache infrastructure
  • Custom API authentication for the Facebook Platform launch
  • Geographic session routing that kept users secure across continents
  • Privacy controls that learned from early mistakes and influenced industry standards

The Facebook Platform launch in 2007 added another layer of complexity: How do you let thousands of third-party developers access user data without compromising security?

Their answer: revolutionary API design with granular permissions, rate limiting, and authentication systems that later influenced how the entire internet handles third-party integrations.

The Seventh Distributed Computing Lesson: Security at planet scale requires custom solutions that evolve with your architecture—you can't retrofit security onto distributed systems.


Chapter 8: Hardware Symphony

By 2009, Facebook was orchestrating a symphony of silicon across multiple continents.

60,000 servers. Think about that number. Before cloud computing, before Infrastructure as a Service, a college website had assembled more computing power than most governments.

But the real magic wasn't in the quantity—it was in the orchestration:

  • Multi-tier load balancing with Layer 4 and Layer 7 routing intelligence
  • Custom flow control to handle their unique "incast" networking problems
  • Geographic distribution that automatically shifted traffic based on capacity and performance
  • Predictive scaling that bought and configured hardware months before it was needed

The efficiency was staggering: 1 million users per engineer—a ratio that remains impressive even by today's standards.

The Eighth Distributed Computing Lesson: Planet-scale infrastructure requires predictive thinking and symphonic coordination—you can't just add servers and hope for the best.


Chapter 9: Culture Revolution

Perhaps the most important innovation wasn't technical—it was cultural.

"Move fast and break things" wasn't just a slogan; it was survival strategy. In a world where Facebook had to build everything from scratch, traditional software development practices would have been corporate suicide.

Facebook's engineers developed a culture of fearless innovation:

  • Experiment with radical solutions like compiling PHP to C++
  • Contribute innovations back to the open-source community
  • Measure everything and optimize based on data, not opinions
  • Question fundamental assumptions about how internet infrastructure should work

It was an obligation to think differently for every enginner. Normal thinking wouldn't get to 500 million users.

This culture enabled a small team to repeatedly achieve the impossible, building custom solutions that often became industry standards.

The Ninth Distributed Computing Lesson: Technical innovation requires cultural innovation—create an environment where impossible solutions are just engineering challenges waiting to be solved.

Ultimate Lesson

Facebook's journey from $85 server to half a billion users proves that the most important ingredient in any distributed system isn't the infrastructure—it's the engineering mindset that refuses to accept "impossible" as a final answer.

They didn't wait for someone else to solve planet-scale computing. They invented planet-scale computing.

Great engineering doesn't adapt to limitations. Great engineering eliminates limitations by building the impossible solutions that become tomorrow's standard infrastructure.

Sometimes the best way to solve an impossible problem is to prove it's not impossible.

Friday, 20 June 2025

Bitcoin Distributed Systems Masterpiece

On October 31, 2008, amid a global financial crisis, an anonymous figure named Satoshi Nakamoto quietly published a nine-page paper that would revolutionize not just finance, but the entire field of distributed systems. Titled "Bitcoin: A Peer-to-Peer Electronic Cash System," this document didn't just propose a new currency—it solved fundamental problems in computer science that had stumped researchers for decades.

While most people see Bitcoin as digital money, computer scientists recognize it as something far more profound: a masterclass in distributed systems engineering that introduced groundbreaking solutions to some of the field's most challenging problems.

The Problem Bitcoin Solved: Trust in a Trustless World

The core challenge Bitcoin addressed was creating "an electronic payment system based on cryptographic proof instead of trust, allowing any two willing parties to transact directly with each other without the need for a trusted third party." This might sound simple, but it represents one of the hardest problems in distributed systems: achieving consensus among untrusted parties across an unreliable network.

Before Bitcoin, this seemed impossible without a central authority. How do you prevent double-spending in a digital system where there's no bank to verify transactions? How do you ensure all participants agree on the same version of truth when they can't trust each other?

Revolutionary Distributed Systems Concepts

1. Decentralized Consensus Through Proof-of-Work



Bitcoin's most ingenious innovation was solving the consensus problem through proof-of-work, where "the majority decision is represented by the longest chain, which has the greatest proof-of-work effort invested in it." This elegant mechanism transforms computational work into votes, creating a democratic system where "proof-of-work is essentially one-CPU-one-vote."

Unlike traditional Byzantine fault tolerance algorithms that require knowing network participants, Bitcoin's consensus works with anonymous, dynamically changing participants. This breakthrough opened the door to truly open, permissionless distributed systems.

Software Systems Lesson: Modern blockchain platforms, distributed databases, and even content delivery networks now use variations of stake-based consensus mechanisms inspired by Bitcoin's proof-of-work innovation.

2. Distributed Timestamp Server Architecture




Bitcoin implements a distributed timestamp server that "works by taking a hash of a block of items to be timestamped and widely publishing the hash." Each timestamp includes the previous one, "forming a chain, with each additional timestamp reinforcing the ones before it."

This creates an immutable, ordered record of events across a distributed network—solving the fundamental problem of establishing temporal ordering in distributed systems without synchronized clocks.

Software Systems Lesson: This concept now powers distributed logging systems, audit trails, and supply chain tracking applications where establishing tamper-proof chronological ordering is crucial.

3. Peer-to-Peer Network Resilience



Bitcoin's network design embraces the chaotic nature of distributed systems. The network requires "minimal structure" and operates on a "best effort basis," where "nodes can leave and rejoin the network at will."

The system gracefully handles network partitions, node failures, and message losses. "New transaction broadcasts do not necessarily need to reach all nodes. As long as they reach many nodes, they will get into a block before long."

Software Systems Lesson: Modern microservices architectures and content distribution networks adopt similar principles of eventual consistency and graceful degradation rather than requiring perfect synchronization.

4. Merkle Trees for Efficient Verification



To address scalability concerns, Bitcoin introduced the use of Merkle trees for "reclaiming disk space" while maintaining cryptographic integrity. This allows "simplified payment verification" where users can verify transactions "without running a full network node."

Software Systems Lesson: Merkle trees are now fundamental to distributed storage systems, content verification protocols, and git version control systems, enabling efficient integrity checking across large datasets.

5. Mempool & Distributed Transaction Coordination



To handle transaction processing across an untrusted network, Bitcoin created the mempool - a distributed buffer where "new transactions are broadcast to all nodes" and each node independently maintains its own queue of unconfirmed transactions. This eliminates the need for a central transaction coordinator while creating a decentralized fee market where miners select transactions based on economic incentives, solving the fundamental distributed systems challenge of resource allocation without central authority.

Software Systems Lesson: The mempool's design principles now power modern distributed queueing systems like Apache Kafka (distributed message buffers), Kubernetes job schedulers (priority-based resource allocation), CDN edge caches (eventually consistent distributed state), and microservices architectures that use economic throttling and circuit breakers to handle load spikes gracefully without central coordination.


6. Economic Incentives for Distributed Cooperation

Perhaps Bitcoin's most overlooked innovation is how it solves the free-rider problem in distributed systems. By making "the first transaction in a block a special transaction that starts a new coin owned by the creator of the block," Bitcoin creates economic incentives for network participation.

This alignment of individual incentives with network health ensures the system remains secure and operational without central coordination.

Software Systems Lesson: Modern distributed systems increasingly incorporate token-based incentives, from IPFS's Filecoin for distributed storage to various proof-of-stake networks.

Fault Tolerance and Security Properties

Bitcoin demonstrates remarkable resilience properties that distributed systems engineers dream of:

Byzantine Fault Tolerance

The system remains secure "as long as honest nodes collectively control more CPU power than any cooperating group of attacker nodes." This provides practical Byzantine fault tolerance with up to 49% malicious participants—far exceeding traditional BFT systems.

Self-Healing Network

When network partitions heal or new nodes join, they can "accept the longest proof-of-work chain as proof of what happened while they were gone." The system automatically converges to a consistent state without manual intervention.

Probabilistic Finality

Bitcoin introduces the concept of probabilistic finality, where "the probability drops exponentially as the number of blocks the attacker has to catch up with increases." This provides practical certainty without absolute guarantees—a crucial insight for large-scale distributed systems.

Lessons for Modern Software Systems

1. Embrace Eventual Consistency

Bitcoin proves that strong consistency isn't always necessary. Systems can be incredibly robust with eventual consistency and conflict resolution mechanisms.

2. Design for Adversarial Conditions

Bitcoin assumes participants may be malicious and designs accordingly. Modern systems should consider adversarial scenarios from the start, not as an afterthought.

3. Use Cryptographic Primitives for Trust

Rather than relying on network security or trusted parties, Bitcoin uses cryptographic proofs. Modern APIs, microservices, and distributed databases increasingly adopt zero-trust architectures with cryptographic verification.

4. Align Incentives with System Goals

Bitcoin's economic model ensures participants want the system to succeed. Modern distributed systems benefit from carefully designed incentive structures, whether through economic rewards, reputation systems, or service credits.

5. Build for Horizontal Scalability

Bitcoin's design allows unlimited participants without coordination overhead. Modern cloud-native applications adopt similar patterns with stateless services and distributed data structures.

Real-World Applications

The distributed systems principles pioneered by Bitcoin now power:

  • Supply Chain Transparency: Companies like Walmart can use blockchain-inspired systems to track food from farm to store
  • Digital Identity: Self-sovereign identity systems that don't rely on central authorities
  • Distributed Storage: IPFS and other systems that create resilient, decentralized file storage
  • IoT Networks: Device networks that can operate without central coordination
  • Audit and Compliance: Immutable logging systems for regulatory compliance

The Bigger Picture

Bitcoin didn't just create digital money—it proved that large-scale coordination is possible without central authority. As Nakamoto concluded, "We have proposed a system for electronic transactions without relying on trust." This breakthrough has implications far beyond cryptocurrency.

The principles demonstrated by Bitcoin—decentralized consensus, cryptographic proof over trust, economic incentives for cooperation, and resilient peer-to-peer architectures—are becoming foundational to how we build distributed systems in an increasingly connected world.

Whether you're building microservices, designing IoT networks, creating distributed databases, or developing decentralized applications, the lessons from Bitcoin's revolutionary approach to distributed systems remain as relevant today as they were in 2008.

The nine-page paper that started as a solution to digital payments ended up teaching us new ways to build trust, achieve consensus, and create resilient systems in our distributed digital world. That's the true genius of Bitcoin— what it showed us was possible.

Monday, 1 May 2023

Distributed Context Propgation

Why distributed context ?

Nowadays, modern applications are designed to be distributed across multiple data centers and regions, leveraging diverse technology stacks. As a result, a typical application architecture may span multiple geographical locations and incorporate various technologies, such as microservices, containers, and server less computing, to meet the complex demands of modern businesses. Application stack might look something like this.
 

Application Stack



In distributed systems, several challenges arise when it comes to maintaining certain aspects such as global state, telemetry recording, long-duration workflow pipeline checkpointing, and feature flags. These challenges are critical to ensuring the reliability, scalability, and fault tolerance of the system.

One of the primary challenges is managing global state, which refers to the data that needs to be shared and synchronized across multiple instances of the application. This can be a complex task, especially in highly distributed environments, where consistency and concurrency control are crucial.

Another key challenge is recording telemetry, which involves collecting, analyzing, and visualizing data about the system's performance, health, and usage. This helps teams monitor and troubleshoot the system, identify bottlenecks, and optimize its overall efficiency.

In addition, keeping track of long-duration workflow pipelines is essential for ensuring that complex tasks are completed reliably and efficiently. This requires checkpointing and resuming workflows from a particular point in the event of failures or other disruptions.

Finally, managing feature flags, which enable the selective enabling or disabling of certain features in the system, is critical for rapid iteration and experimentation. However, managing feature flags across multiple instances of the application can be challenging, as teams need to ensure that the correct flag configuration is propagated to all instances.


What is solution ?


Several solutions exist to tackle these challenges, such as distributed caching, message-oriented middleware, or specialized tools like Zookeeper for managing state. While these solutions can be effective, they often come with the complexity of managing a sophisticated infrastructure. Additionally, implementing a complex infrastructure for a relatively small use case can lead to more issues than solutions.

To address these challenges, OpenTelemetry offers a novel approach called "context and propagation" that operates at the transport layer. The idea behind this approach is to provide a simple and unified way to propagate information across different components of the distributed system.

Inspired by OpenTelemetry's context and propagation, I would like to introduce an idea that could help address these challenges more effectively. This approach involves utilizing a lightweight middleware layer that sits between the different components of the distributed system. The middleware layer would be responsible for managing the state, telemetry, and feature flags, and would provide a simple and standardized API for the application components to interact with.

By using this approach, teams can simplify the management of complex infrastructure while ensuring that their distributed system remains reliable, scalable, and fault-tolerant. Additionally, this approach could help teams experiment with new features and iterate quickly without compromising on the overall performance and stability of the system.



How does context API look ?

The proposed Context API would have several key features to help manage the state, telemetry, and feature flags of a distributed system more effectively. Some of these features include:

  1. Context Name: This feature would enable the association of a user-friendly name with all the state, telemetry, and feature flags in the system. This name can be used to index and query the information easily.

  2. Time to Live: Some context information is short-lived and should be automatically managed by the underlying implementation. This feature would enable the setting of a time to live for the context information, allowing for automatic expiration and cleanup when no longer needed.

  3. Abstract Data Types: The Context API would support abstract data types such as counters, maps, and sets, allowing for flexible and efficient management of the state, telemetry, and feature flags.

  4. Durability: The context information would be durable, meaning it would survive application restarts and remain available for a longer duration, enabling teams to understand the behavior of the distributed system over time.

By incorporating these features, the proposed Context API would provide a lightweight and standardized way to manage the complex state, telemetry, and feature flags of a distributed system, enabling teams to focus on building reliable and scalable applications without the burden of managing a complex infrastructure.


API -> Implementation

One of the key benefits of designing a system with a clear separation between the API and its implementation is the flexibility it offers. This approach allows teams to easily swap out and upgrade different components of the system without impacting the overall functionality of the API.

Moreover, a clear separation between the API and implementation enables the support of heterogeneous ecosystems, where different components of the system may use different technology stacks or operate in different environments. This allows teams to leverage the strengths of each technology stack and environment, without compromising the overall functionality and interoperability of the system.

In summary, a clear separation between the API and implementation provides teams with flexibility, interoperability, modularity, and maintainability, allowing them to build complex and reliable systems that can support the needs of diverse and dynamic ecosystems.

What operations are supported on Abstract data type

Let's take a closer look at the sample API and the operations it supports. One of the benefits of the proposed Context API is its ability to support a wide range of use cases, and new abstract data types (ADTs) can be added to address more advanced scenarios.

At its core, the Context API provides a simple and standardized way to manage state, telemetry, and feature flags in a distributed system.

The Context API can be extended with new ADTs to support more advanced use cases. For example, a new ADT could be added to support distributed locks, allowing different components of the system to coordinate access to a shared resource. Another ADT could be added to support distributed queues, enabling components to communicate asynchronously in a reliable and scalable manner.




A sample implementation that leverages the concepts discussed above is available at https://github.com/ashkrit/corejava/tree/master/playground/src/main/java/context. This implementation showcases how the proposed Context API can be used to manage state, telemetry, and feature flags in a distributed system.

The sample implementation uses an H2 database for the persistence of state, and this database is abstracted through the ContextProviderClient.java interface. This abstraction enables the use of any backend system to store states, providing flexibility and adaptability.

One of the key challenges of the Context API is state replication in a distributed system. To address this, various approaches can be taken, depending on the requirements of the system. One option could be to use a REST API that has access to a replicated database system, where all state management is done through the API. Alternatively, a multi-data center cache, such as Redis, could be used to store and replicate state across different regions. In some cases, a distributed and replicated file system could also be used to store and manage state.

Overall, the choice of backend system and replication strategy will depend on the specific needs of the distributed system and the desired level of fault tolerance, consistency, and scalability. The flexibility and abstraction provided by the Context API make it possible to use a variety of backend systems and replication strategies, enabling teams to choose the approach that best meets their requirements.

Overall, the sample implementation provides a practical example of how the Context API can be used to simplify the management of state, telemetry, and feature flags in a distributed system, while promoting flexibility, scalability, and reliability.

Sunday, 28 June 2020

Ship your function

Now a days function as service(FaaS) is trending in serverless area and it is enabling new opportunity that allows to send function on the fly to server and it will start executing immediately.   

Code as data as code.

This is helps in building application that adapts to changing users needs very quickly.
Function_as_a_service is popular offering from cloud provider like Amazon , Microsoft, Google etc.

FaaS has lot of similarity with Actor model that talks about sending message to Actors and they perform local action, if code can be also treated like data then code can also be sent to remote process and it can execute function locally. 

I remember Joe Armstrong talking about how during time when he was building Erlang he used to send function to server to become HTTP server or smtp server etc. He was doing this in 1986!

Lets look at how we can save executable function and execute it later.
I will use java as a example but it can be done in any language that allows dynamic linking. Javascript will be definitely winner in dynamic linking. 

Quick revision
  Lets have quick look at functions/behavior in java


Nothing much to explain above code, it is very basic transformation.

Save function
Lets try to save one of these function and see what happens. 


Above code looks perfect but it fails at runtime with below error

java.io.NotSerializableException: faas.FunctionTest$$Lambda$266/1859039536 at java.io.ObjectOutputStream.writeObject0(ObjectOutputStream.java:1184) at java.io.ObjectOutputStream.writeObject(ObjectOutputStream.java:348) at faas.FunctionTest.save_function(FunctionTest.java:39) at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)

Lambda functions are not serializable by default.
Java has nice trick about using cast expression to add additional bound, more details are available at Cast Expressions.

In nutshell it will look something like below


This technique allows to convert any functional interface to bytes and reuse it later. It is used in JDK at various places like TreeMap/TreeSet as these data structure has comparator as function and also supports serialization.
   
With basic thing working lets try to build something more useful.

We have to hide & Serialized magic to make code more readable and this can be achieved by functional interface that extends from base interface and just adds Serializable, it will look something like below


Once we take care of boilerplate then it becomes very easy to write the functions that are Serialization ready.



With above building block we can save full transformation(map/filter/reduce/collect etc) and ship to sever for processing. This also allows to build computation that can recomputed if required.

Spark is distributed processing engine that use such type of pattern where  it persists transformation function and use that for doing computation on multiple nodes. 

So next time you want to build some distributed processing framework then look into this pattern or want to take it to extreme then send patched function to live server in production to fix the issue. 

Code used in in post is available @ faas