Showing posts with label J2EE. Show all posts
Showing posts with label J2EE. Show all posts

Wednesday, February 17, 2010

Java Transactions – 2 phase commit protocol

Most of us know about ACID properties. To understand Atomicity in detail, lets look at Two Phase Commit Protocol (2PCP). When a client starts a transaction using Transaction Manger, it associates a coordinator with the transaction known as Transaction Coordinator (TC) :) Now the question is, how does 2PCP helps maintaining Atomicity of a transaction?

In the phase 1 of 2PCP, controller asks all the participant if they can commit. Depending upon the response TC can enter phase 2, if all the participants agreed to commit, TC will confirm this decision with all the participants in phase 2 otherwise roll back decision is sent to all the participants. Also, after phase 1 each participant as well as TC saves enough information in their transaction logs so that they can roll back in case of failure. After phase 1, all the participants which agreed to commit remain blocked till they get phase 2 message.

Failure conditions: What if transaction controllers fails after phase 1? The participants will remain in locked state. Have you ever noticed when you start an application server, it tries to recover the incomplete transactions? Look at the startup logs when you see it next time. What is does is, it looks as the transaction logs and tries to recover the previous state. This is main reason that after phase -1 each participant save enough information in the disk, so that it can recover from the failed transaction. There are other failure scenarios like “What if resource A fails after phase -1?” All these kind of scenarios result in transaction timeout and it rolls back. Resource A when come back, it should read the transaction logs it saved before it crashed and bring itself to consistent state.

How about transactions running in a cluster? Each node in a cluster can start and manage its own transactions. If one of the nodes crashes (say Node1), then HA Manager (some kind of P2P implementation which is highly available and controls all the nodes) with the help of Node2's transaction manager reads the shared transactional logs and recover the locked resources (which is result of failure of transaction controller) Locked resource could be a poisoned message in a JMS queue or a locked row in a database (sometimes many rows get locked due to page level locking – DB2 does it.)

Saturday, June 20, 2009

Distributed Cache

Every application requires some level of caching, either its finite value (values which remain same for the application) cache or cache which holds stateful objects like HTTP Session objects. If the application is running in the same JVM, there is no problem*(don’t miss the *) but then we have a single point of failure.

Problems in caching: I would rather say important decisions which are mostly taken wrong.
1. What data to cache - should business entities / transactional data be cached?
2. How to clear the cache? Should soft references be used? How to remove the cached items which are not used?

So, running the same application in a clustered environment makes the above decision making process even harder.

There are different solutions to solve the above problem like JCache which takes care of distributing and synchronizing the objects in cache. Now a new project Infinispan [JBoss] expose the same interface as JCache [JSR-107] but takes a one step further to provide highly available data grid platform. So, what’s the difference? The difference is how data is being cached. In a typical caching systems the data is just replicated to provide high availability & therefore total memory available remains the same even after running the application in cluster [If we replicate the same object @ N servers, we need N * sizeOf(Object) memory in heap.] On the other hand data grid solutions increase the total effective heap size by replicating the objects in any given K servers [ K < N & K > 0]. So, what does it mean? It means that I can choose which objects I want to replicate in the cloud and how many times.

According to me data grid solutions make more sence. If you read the project introduction of Infinispan carefully, JBoss is going to concentrate on it and stop supporting JCache soon.

The purpose of this post is to lay down the background of what I will be blogging about in the coming few post. In the next post, we'll talk about J2EE Clustering, associated issues and its solutions.