Stop the Hollyweb! No DRM in HTML5.   

Sunday, January 8, 2012

Correcting Hadoop's HDFS java.io.IOException Errors

This past Christmas and New Year, like the last three years now, I accompanied my wife to Bogota, Colombia for the holidays. The difference about this trip was; our pet Mickey was going with us, we had no side trips planned, and I was taking my new Lenovo T420 with me. We arrived in Bogota late at night on December, 22nd, and after being greeted by family, we took a cab to my mother-in-laws house. There’s something both exciting and terrifying about cab rides in Bogota, but after time, for me at least, it’s just fun.

I woke up the next morning well rested and we began our Christmas vacation. I began setting up a new wireless router that I had brought with me so that I could work remotely from wherever was most comfortable. After having installed the router, I connected with my new laptop and set my priorities for the task that I needed to complete. Was the typical task, update my time, respond to some emails, finalize peer reviews, and check up on database backups. After these tasks were completed, I had time to relax.

What was planned for us was day to day, but mostly we would go out to have lunch or dinner with friends and family, then return home. On a few nights, we engaged in consuming heavy amounts of adult beverages and dancing which is the custom in Colombia. Most of our time however, was spent at home. I would wake up early, before anyone else, and sit at the dining room table, next to the window with my laptop and watch the sun rise over the mountains enjoying some fresh Colombian coffee.

With nothing to do so early in the morning, I decided to log onto my desktop back in DC. When I connected to my desktop, I saw that I had left open a ssh connection to a small Hadoop cluster I had setup to do some testing, except HDFS was not working properly. This was the perfect opportunity to find out what went wrong in my install and configuration of the cluster. The only catch was, is that I would have do everything from the command shell, no gui. I had followed Michael Noll’s “Running Hadoop on Ubuntu”,but now, I was getting errors in the namenode logs.

I was getting java.io.IOException errors. In Michael Noll’s how-to, he describes how he addressed this error by reformatting the cluster. He described how he stopped all running daemons and deleted the /app/hadoop/tmp/hdf/name/data directory and then ran bin/hadoop namenode –format . Somewhere in my troubleshooting my errors and researching online, I found that it was also a good idea to add the following properties to the hdfs-site.xml configuration file.

< !-- Adding dfs.data.dir dfs.name.dir 1/1/2012. -- >
< property >
< name > dfs.data.dir < /name >
< value > /app/hadoop/tmp/dfs/name/data < /value >
< final > true < /final >
< /property >
< property >
< name > dfs.name.dir < /name >
< value > /app/hadoop/tmp/dfs/name < /value >
< final > true < /final >
< /property >

Also, if you get permission denied (publickey,password), you may want to check that the paths for the properties you added to the hdfs-site.xml file are correct. If this problem persist, you might try running the following on all nodes;

sudo chown –R hduser:hadoop /app/hadoop

Some of the other errors that I ran into were as follows;

Cannot lock storage /app/hadoop/tmp/dfs/name. The directory is already locked.
org.apache.hadoop.hdfs.server.namenode.FSNamesystem: Fatal Error : All storage directories are inaccessible.
ERROR org.apache.hadoop.hdfs.server.namenode.NameNode: java.net.BindException: Problem binding to Address already in use

I did not do a good job at keeping notes during my troubleshooting these issues, but it seemed that whenever I would try to fix one thing, a different error would pop up. I did find that I had made an typo in the hdfs-site.xml file. At the end of each path, I had added a /. Therefore, instead of /app/hadoop/tmp/dfs/name, I had /app/hadoop/tmp/dfs/name/. But once I corrected that and delete all data in the HDFS directory and then ran format, everything worked! So here is how that went.

After stopping all daemons and correcting the paths in the hdfs-site.xml file, I then deleted all data in the HDFS directory on all nodes.

hduser@bigdata1:/app/hadoop/tmp/dfs$">hduser@bigdata1:/app/hadoop/tmp/dfs$ sudo rm -rf *

Then, I ran the format.

hduser@bigdata1:/usr/local/hadoop/hadoop$ bin/hadoop namenode –format

The output looks like;

12/01/04 07:59:15 INFO namenode.NameNode: STARTUP_MSG:
/************************************************************
STARTUP_MSG: Starting NameNode
STARTUP_MSG: host = bigdata1/172.20.10.92
STARTUP_MSG: args = [-format]
STARTUP_MSG: version = 0.20.2
STARTUP_MSG: build = https://svn.apache.org/repos/asf/hadoop/common/branches/branch-0.20 -r 911707; compiled by 'chrisdo' on Fri Feb 19 08:07:34 UTC 2010
************************************************************/
12/01/04 07:59:15 INFO namenode.FSNamesystem: fsOwner=hduser,hadoop,adm,dialout,fax,cdrom,floppy,tape,audio,dip,video,plugdev,fuse,lpadmin,netdev,admin,sambashare
12/01/04 07:59:15 INFO namenode.FSNamesystem: supergroup=supergroup
12/01/04 07:59:15 INFO namenode.FSNamesystem: isPermissionEnabled=true
12/01/04 07:59:15 INFO common.Storage: Image file of size 96 saved in 0 seconds.
12/01/04 07:59:15 INFO common.Storage: Storage directory /app/hadoop/tmp/dfs/name has been successfully formatted.
12/01/04 07:59:15 INFO namenode.NameNode: SHUTDOWN_MSG:
/************************************************************
SHUTDOWN_MSG: Shutting down NameNode at bigdata1/172.20.10.92
************************************************************/

Next, I started HDFS.

hduser@bigdata1:/usr/local/hadoop/hadoop$ bin/start-dfs.sh

It’s output was;

starting namenode, logging to /usr/local/hadoop/hadoop/bin/../logs/hadoop-hduser-namenode-bigdata1.out
bigdata2: starting datanode, logging to /usr/local/hadoop/hadoop/bin/../logs/hadoop-hduser-datanode-bigdata2.out
bigdata3: starting datanode, logging to /usr/local/hadoop/hadoop/bin/../logs/hadoop-hduser-datanode-bigdata3.out
bigdata4: starting datanode, logging to /usr/local/hadoop/hadoop/bin/../logs/hadoop-hduser-datanode-bigdata4.out
bigdata1: starting secondarynamenode, logging to /usr/local/hadoop/hadoop/bin/../logs/hadoop-hduser-secondarynamenode-bigdata1.out

I ran the HDFS Admin Report to see the status of my cluster.

hduser@bigdata1:/usr/local/hadoop/hadoop$ bin/hadoop dfsadmin –report

The report displays the following;

Configured Capacity: 206701436928 (192.51 GB)
Present Capacity: 186873368576 (174.04 GB)
DFS Remaining: 186873294848 (174.04 GB)
DFS Used: 73728 (72 KB)
DFS Used%: 0%
Under replicated blocks: 0
Blocks with corrupt replicas: 0
Missing blocks: 0

-------------------------------------------------
Datanodes available: 3 (3 total, 0 dead)

Name: 172.20.10.127:50010
Decommission Status : Normal
Configured Capacity: 68900478976 (64.17 GB)
DFS Used: 24576 (24 KB)
Non DFS Used: 6721228800 (6.26 GB)
DFS Remaining: 62179225600(57.91 GB)
DFS Used%: 0%
DFS Remaining%: 90.24%
Last contact: Wed Jan 04 08:00:20 EST 2012


Name: 172.20.10.128:50010
Decommission Status : Normal
Configured Capacity: 68900478976 (64.17 GB)
DFS Used: 24576 (24 KB)
Non DFS Used: 6732939264 (6.27 GB)
DFS Remaining: 62167515136(57.9 GB)
DFS Used%: 0%
DFS Remaining%: 90.23%
Last contact: Wed Jan 04 08:00:20 EST 2012


Name: 172.20.10.48:50010
Decommission Status : Normal
Configured Capacity: 68900478976 (64.17 GB)
DFS Used: 24576 (24 KB)
Non DFS Used: 6373900288 (5.94 GB)
DFS Remaining: 62526554112(58.23 GB)
DFS Used%: 0%
DFS Remaining%: 90.75%
Last contact: Wed Jan 04 08:00:17 EST 2012

After this I started MapReduce.

hduser@bigdata1:/usr/local/hadoop/hadoop$ bin/start-mapred.sh

It’s output was this;

starting jobtracker, logging to /usr/local/hadoop/hadoop/bin/../logs/hadoop-hduser-jobtracker-bigdata1.out
bigdata3: starting tasktracker, logging to /usr/local/hadoop/hadoop/bin/../logs/hadoop-hduser-tasktracker-bigdata3.out
bigdata2: starting tasktracker, logging to /usr/local/hadoop/hadoop/bin/../logs/hadoop-hduser-tasktracker-bigdata2.out
bigdata4: starting tasktracker, logging to /usr/local/hadoop/hadoop/bin/../logs/hadoop-hduser-tasktracker-bigdata4.out

Success!




Now there are three web interface URLs that you can use to check up on your clusters health, there are;




http://localhost:50030/ – web UI for MapReduce job tracker(s)
http://localhost:50060/ – web UI for task tracker(s)
http://localhost:50070/ – web UI for HDFS name node(s)










Wednesday, June 29, 2011


It's time again for SQL Saturday Washington DC! With last year's event being such a huge success, this year's event will be even better! We are expecting over 250 attendees to show up for a full day of training from some of the best speakers in the SQL Server community.

This is one event you don't want to miss! Everyone who attends will get a SQL Saturday T-Shirt and lots of swag! Free breakfast and lunch will be provided. We will have several SQL Server vendors that will be presenting the latest upgrades and solutions. And at the end of the day, we will be raffling off some awesome door prizes provided by our many sponsors!

SQL Saturday is a FREE one day training events for SQL Server professionals. SQL Saturday was initially the idea of three DBAs that wanted a Code Camp style event just for SQL Server professionals. It began with the first SQL Saturday in Tampa, Florida. The event was such a huge success, more events followed. After over 40 successful events, The Professional Association for SQL Server took over the administration of SQL Saturday events.

SQL Saturday events are usually divided into three tracks consisting of BI, Database Development, and Database Administration. Typically there will be five sessions per track. SQL Saturday tries to recruit local SQL Server professionals to present at the event, but occasionally, more nationally known speakers may also present.

SQL Saturday is all about sharing issues and solutions, and gaining knowledge that will make you a better SQL Server professional.

We’ll see you there!!!

Register at: http://www.sqlsaturday.com/96/eventhome.aspx

FOLLOW US!



Friday, June 10, 2011

SQL Server: An error occurred while executing batch. Error message is: The directory name is invalid.

I logged onto the server this morning to delete some files because the disk had run out of free space. After deleting some files, I went to run some maintenance scripts in SSMS and received the following error:

An error occurred while executing batch. Error message is: The directory name is invalid.


After doing some online research, I found out that the reason for this error is because SQL Server cannot find the temp folder in which to store the query results. To correct this, try logging off and back on to the machine that you are running SSMS on. If the error persists, reboot the machine.

Thursday, June 2, 2011

Cannot edit job steps in SSMS 2008 R2

When attempting edit the job step, or view the job step details, I received the following error:

Creating an instance of the COM component with CLSID {AA40D1D6-CAEF-4A56-B9BB-D0D3DC976BA2} from the IClassFactory failed due to the following error: c001f011. (Microsoft.SqlServer.ManagedDTS)


You may also get this error when attempting to create a new job step.

I did a search and found several old post requesting help with the same problem. Microsoft corrected the problem with the last Hotfix; Cumulative Update package 7 for SQL Server 2008 R2.

For more information on Cumulative Update package 7 for SQL Server 2008 R2 and how to download it, go here; http://support.microsoft.com/kb/2507770

Wednesday, May 18, 2011

MADExop: The Mid Atlantic Developer Expo in Hampton, VA June 30 - July 1, 2011


Ok, so I get an email from Andrew Duthie at Microsoft asking me to get the word out about the upcoming Mid Atlantic Developer Expo (MADExop) in Hampton, VA. Now this was the first time I had heard about it and really wanted go after I visited the website and read all the cool sessions they had lined up. But then, I saw that for only $20 you could bring your child for an all day kid geekout session! Now, what could be better? You go for two days of awesome and, did I mention cheep?, developer training from the best developers around and inspire your child to follow in your footsteps at the same time!

Take a look at the lineups for Thursday and Friday sessions:
Click here for a PDF of the Day 1 Agenda
Click here for a PDF of the Day 2 Agenda

Register here!
MADExpo 2011 Registration!

Thursday, May 12, 2011

Idera's SQLsafe Restore Error "too recent to apply to the database"

Last week, my manager came to me and asked if I could perform our first restore since installing the new SQLsafe Backup and Recovery software. Because I was the one that pushed for SQLsafe to take over as our Backup and Recovery solution for all production SQL Servers, I was more than happy to show off SQLsafe's ease of use in restoring databases. I select the point in time that I wanted to restore to and clicked NEXT.

I was expecting for the database to be restored in no time, but instead, and with horror written all over my face, I received the following error message in the "Result Text" with a BIG RED "Error" Progress indicator next to my Restore status.

-------------------------------------------- snip -------------------------------------------

<!--[if gte mso 9]> Normal 0 false false false EN-US X-NONE X-NONE

" Server instance: INSTANCE/NAME, Database: mas500_pl

The log in this backup set begins at LSN 94000000227900001, which is too recent to apply to the database. An earlier log backup that includes LSN 94000000221200001 can be restored.

RESTORE LOG is terminating abnormally."

Normal 0 false false false EN-US X-NONE X-NONE

--------------------------------------------------------------------------------------------

With embarrassment, I turned to my manager and told him that I would call Idera's support staff and get help with restoring the database. I called the support number and got routed to a voice mailbox. However, within just a few minutes, I got a call back from Carl at Idera's Customer Support. I explained to him the issue I was having and after some brief research on his part, he said based on the message I was receiving that it looked like another backup had been done and that the SQLsafe backup that I was attempting to use was not the most current.

I then opened SSMS and started the restore database wizard. I selected the database that I wanted to restore and then clicked file and add to browse to the folder. Then SURPRISE! The default backup folder appeared with recent backups of the database that I was attempting to restore.

I then clicked on the SQL Server Agent and found a Full Backup job scheduled to run every Sunday at 2:00 am which was after the SQLsafe Full backups.

To restore the database, I first had to restore the native Full backup from Sunday morning with the following sql script;

RESTORE DATABASE [mas500_pl]

FROM DISK = N'C:\Program Files\Microsoft SQL Server\MSSQL.1\MSSQL\BACKUP\mas500_pl_db_201105010200.BAK'

WITH FILE = 1, NORECOVERY, NOUNLOAD, STATS = 10

GO

Next I opened SQLsafe and restored each Differential, one at a time, with "Force restore", "Ingnore Checksum Errors", and with non-recovery, "Not accessible", until the last Differential. For the last Differential, I restored it with recovery or "Fully accessible".

So a BIG Thank You goes out to Carl at Idera's Customer Support!

Friday, November 12, 2010

PASS Summit 2010 Keynote David DeWitt

This was the best presentation during all of PASS Summit 2010. Everyone enjoyed Dr. DeWitt's keynote session on SQL Query Optimization.