Thursday, November 13, 2014

Decommissioning DataNodes on HDP 2.0

Issues & Adventures with Decommissioning Data Nodes on HDP 2.0

Use the Ambari UI to decommission...

  1. Decommission the HBase RegionServer
  2. Decommission the NodeManager
  3. Finally Decommission the DataNode


The HBase RegionServer and NodeManager statuses will to to "Decommissioning" and eventually to "Decommissioned".

Use have to use the Ambari Rest API to see that the DataNode is on its way to being Decommissioned...

> curl -s --user admin:40rt0n http://prdslsldsafht25.myfamilysouth.com:8080/api/v1/clusters/prdslsldsafht/hosts/prdslsldsafht12.myfamilysouth.com/host_components/DATANODE | less

Look for "desired_admin_state" : "DECOMMISSIONED".

What seems to be the correct steps...
One would think that Decommissioning should simply work, but it seems that there are bugs that require the procedure to be specific...

  1. First make the Active NameNode NN1
  2. Perform the decomm of the DN (Via the Ambari UI).
    1. You will notice that NN2 does not recognize the decomm of the DN.
    2. There is a bug that indicates that the Decomm action might be taken by a random NN,
      Which may invalidate my hypothesis...
      1.   https://issues.apache.org/jira/browse/AMBARI-4927

What goes wrong if Decomm occurs with NN2 and the Active NN...

  • If the Active is NN2 and Standby is NN1...
    • The decommissioning of a node, is only recognized by NN1, NN2 will continue to write to the datanode as NN1 trys to ensure that all blocks are replicated off of it. NN2 is at fault as no more new blocks should be sent to a decommissioning node.
    • Thus it is possible for Standby-NN1 to be decommissioning a DN while Active-NN2 is writing to the DN.
    • It may be that HA decommissioning tests are always running with Active-NN1 and Standby-NN2 and never the other way around.


Decommissioning nodes would not fully decommission because a very few blocks continue to be indicated as under-replicated, but I could not find those blocks and files via fsck to handle them.

NN UI...
| Decommissioning Datanodes : 2
| Node Transferring
| Address Last
| Contact Under Replicated Blocks Blocks With No
| Live Replicas Under Replicated Blocks
| In Files Under Construction Time Since Decommissioning Started
| prdslsldsafht12 10.211.25.122:50010 2 232143 0 4 0 hrs 8 mins
| prdslsldsafht13 10.211.25.123:50010 2 261441 0 3 0 hrs 8 mins

hdfs fsck...
|  Total dirs:    140574
|  Total files:   5367953
|  Total symlinks:                0 (Files currently being written: 5353)
|  Total blocks (validated):      5683249 (avg. block size 40621873 B) (Total open file blocks (not validated): 1796)
|  Minimally replicated blocks:   5683249 (100.00001 %)
|  Over-replicated blocks:        385172 (6.7773204 %)
|  Under-replicated blocks:       0 (0.0 %)
|  Mis-replicated blocks:         0 (0.0 %)
|  Default replication factor:    3
|  Average block replication:     3.0742598
|  Corrupt blocks:                0
|  Missing replicas:              0 (0.0 %)
|  Number of data-nodes:          75
|  Number of racks:               5
| FSCK ended at Thu Nov 13 21:26:31 MST 2014 in 1680353 milliseconds


I simply pull the nodes, making them dead-nodes by turning off the DN process.
Stangely, Ambari shows that the nodes are in a decommissioned state.
But, I still have the option to stop the DataNode service, so I do that.
The NN UI continues to indicate that the 2 nodes are still in the process of decommissioning.
So, we need to make the NN think that the node is dead by removing the hostnames from dfs.exclude and running -refreshNodes.
After running -refreshNodes, the nodes no-longer display in the Decommissioning nodes list.
And, the main NN UI page shows that 2 nodes are Decommissioned.

Thus, we can by-pass the Decommissioning process if it can't resolve the last few under replicated blocks.

The following link was helpful in that it indicated that it was possible for the decommission process to completely stall on a final few under-replicated blocks. One should be able to find the files associated those under replicated blocks and perform an appropriate handling on then. But, in our case, fsck was not finding any under-replicated blocks at all.
http://stackoverflow.com/questions/17789196/hadoop-node-taking-a-long-time-to-decommission


Friday, May 21, 2010

Extend a laptops Keyboard and Mouse to other Computers/Displays




I've recently set myself up with a trading system that consists of 2 large displays and 2 laptops and their displays.
I am using both laptops. The first laptop is used mainly to display trading charts, and to develop and upgrade my indicator to NinjaTrader 7. I don't do real trades with NinjaTrader 7, as it is still Beta software.
The second laptop, is used to display trading charts, and to do real trading, using the stable version of NinjaTrader 6.5.
The problem is that it can be very confusing to move between 2 mice and 2 laptop keyboards. So , I Googled for a way of extending a computers laptop and keyboard over to another computer. Essentially, VNC without the display.

The solution that I found was a project called Synergy-Plus. This code allows you to set up your main interface-computer (the one that you sit in front of most of the time), as the server, and another computer as a client. Once, the software is set up, when your mouse hits the left of right edge of you monitor, the mouse and keyboard control jumps over to the other computer and allows you to control its mouse, keyboard and windows.
Synergy-Plus is has really reduced the confusion in the area by allowing me to control both laptops through a single laptop's mouse and keyboard.

Synergy-Plus supports various Operating Systems including Windows, Linux and MAC.
I am currently using it with windows 7 64-bit with no problems so far.

Friday, February 5, 2010

Solaris: Cron Jobs Don't Run (!bad user) locked account

It took a bit of time to find out that my cron jobs where not running for a given user on a Solaris box.

Looking at /var/cron/log I was seeing "!bad user" (userName).

But, the user was fine, I was using it, I'd already su'ed to become the user several times.

Turns out that the problem is that, cron does not like users that have been locked out due to not changing the password. I never saw a request to change the password because I always su to become the user rather than logging in as that user.

So, to fix the problem, run "passwd -u userName" as root or via sudo. After that, the cronjobs run fine.

Apparently, the fact that on Solaris, cron does not run the jobs of a locked user, is not documented in any visible manner.

Sunday, January 3, 2010

The Acer Aspire Revo R3610 as a Linux Server.


For many years I have been building Linux servers for my own web and mail services.
I've always limited by hardware costs to $500 or less by reusing old equipment.
Over the last several months, I have had hard drive failures on 3 of my servers, two of which I could live without. Last week, when the third system (my mail server) started to experience failures, I started looking for a replacement server.

I started to look at the PogoPlug as a possible solution. My new PogoPlug had just recently arrived in the mail, and it seemed like a good candidate as a mail server replacement. With is ultra small footprint and its low power consumption, I started hacking it with grand designs in mind. Unfortunately, after getting to the point of being able to install OpenPogo packages to a USB drive, I became cautious about making too many changes to the PogoPlugs sofware. Partly, due to the fact that I have come to enjoy what PogoPlug does best, making data on a hard drive that you have at home, available on the web for yourself and others (if you choose), in a safe a secure manner. Thus, I dropped any further major tweeking for now, until I can see a cleaner safer way of adding additional Linux services to the PogoPlug via OpenPogo, without compromising it's security or bricking it.

I took a trip to Fry's hoping to find a solution via a Shuttle X2700 mini-pc server. I only wanted to spend about $400 at most on the new server, but found that I would be at about $600 after buying the bare bones system, then adding the CPU memory and hard drive. So I lost interest in that route and looked at what they offered in terms of complete systems. Here, I came across the Acer Revo. It's an affordable mini-pc system that uses the ATOM processor, the R3610 is a 64bit processor with 2 cores. Linux actually reports 4 cores because each core can run 2 threads. I bought the R3610 for $329 plus tax. This gets me all I need for a server plus more.

The Acer Aspire Revo R3610 came with Windows 7. But I want Linux to be the main OS.
When I initially tried to boot the Ubuntu installer from a USB drive. I was disappointed to see the message "No Operating system found". I reformatted to USB drive to NTFS and used some procedures that I found on the net to make the USB drive bootable via an original Windows Vista bootable Install CD, but that did not work either as I got messages like "Bootmngr was not found." or "OS was not found.". The simple solution was just a few steps.
  1. Download the iso for Ubuntu Server or Ubuntu Desktop.
  2. Cleanly format a USB drive as FAT32.
  3. Use the "Universal Netboot Installer" to place a bootable install of 64 bit Ubuntu Linux on the USB Drive. Just point the unetbootin utility to the Ubuntu iso and the USB drive letter of the newly formatted FAT32 USB drive.
  4. Plug the USB drive into the Acer Revo.
  5. Set the Acer Revo to boot from the USB drive.
  6. Partition about half of the 160Gig drive to run Linux, and leave the other half for Windows 7.
  7. Install Ubuntu...
Ultimately, I'd like to be able to run Windows 7 VM os to Linux via Xen.
But, I'll have to leave that experiment for later...

Acer Aspire Revo Specs:
AR3610-U9012Genuine Windows® 7 Home Premium , Intel® Atom™ Processor N330 (1MB L2 cache, 1.60GHz, 533MHz FSB), 2GB (1/1) DDR2 800 SDRAM, 160GB SATA hard drive, multi-in-one card reader, NVIDIA® ION™ graphics, gigabit LAN, 802.11b/g/Draft-N WLAN

Wednesday, May 6, 2009

Rails (2.3.2) Install, MySql issues

I recently set up Rails 2.3.2 on my Windows Vista Laptop.
I just installed Ruby and Rails (2.3.2) on my Windows Vista Laptop.
I came across a problem in getting MySql working with Rails...

I used the following tutorial to perform the install:
At Step 3, if you click on "About your application’s environment", you may get the message "We're sorry, but something went wrong". If you do, continue with the steps below. If not, you are probably okay.

When you click on "About your application’s environment" here is what you should see:

Ruby version 1.8.6 (i386-mswin32)
RubyGems version 1.3.1
Rack version 1.0 bundled
Rails version 2.3.2
Active Record version 2.3.2
Action Pack version 2.3.2
Active Resource version 2.3.2
Action Mailer version 2.3.2
Active Support version 2.3.2
Application root C:/Ruby/firstproject
Environment development
Database adapter mysql
Database schema version 0


If you get the message "We're sorry, but something went wrong", here's what you need to do...

Dowload and install the following file into \Ruby\bin:

Create the "firstproject" database and the "rails" user using MySql tools. Then...

Edit the file in /Ruby/firstproject/config/database.yml:

development:
adapter: mysql
database: firstproject
username: rails
password: yourpassword
host: localhost
port: 3306
pool: 5
timeout: 5000

test:
adapter: mysql
database: firstproject
username: rails
password: yourpassword
host: localhost
port: 3306
pool: 5
timeout: 5000

production:
adapter: mysql
database: firstproject
username: rails
password: yourpassword
host: localhost
port: 3306
pool: 5
timeout: 5000

Restart rails, and you should get the appropriate response when you click on "About your application’s environment".

-Best...

References:


Keywords:
  • Problems with mysql
  • gem mysql install fails
  • webrick "can't convert fixnum into string"

Friday, March 6, 2009

QuickenBooks Upgrade to 2009 Pro...

Just upgraded from QuickBooks 2006 Pro to 2009 Pro.
After the upgrade, QuickBooks converts the data file so that it is compatable with the latest version (2009 Pro), and this process seemed to run fine.

Problem:
The main problem that I ran into after the upgrade was that QuickBooks 2009 Pro would crash when I would try to run the last bank reconciliation report (that was last done on the prior version).
I ran "Help->Update QuickBooks" to update the program, but it continued to crash when I would run the prior back reconciliation report.

Solution:
Then I found "File->Utilities->Rebuild Data". It required that you execute another backup. But, after the rebuild, QuickBooks 2009 Pro no-longer crashed when running the last bank reconciliation report.

So, if you are up against this problem of QuickBooks crashing after an upgrade.
  1. Install QuickBooks 2009.
  2. Use the "Help->Update QuickBooks" to update the program to its latest.
  3. Let QuickBooks 2009 Pro update your data file to it's latest.
  4. Run "File->Utilities->Rebuild Data" on the file to ensure that avoid the "crash on reporting bug".

-Best...

Saturday, January 31, 2009

Cygwin "unable to remap..."

After thinking that Cygwin was finally "stably" installed on Vista-64bit, I started to see "unable to remap..." errors while running a perl CPAN module install.

What was that command that I had to run?.... Oh yeah, "rebaseall".

  • Start up an ash session from a cmd.exe window.
  • cd c:\cygwin
  • \bin\ash.exe
  • Then from the ash prompt: /bin/rebaseall
  • Then start up a Cygwin window... and everything should be working again.
-endCommunication!