Wednesday, 25 January 2012

How to check if Innodb log files are big enough

InnoDB uses log files to store changes that must be applied to the database after service interruption (e.g power outage, crash). Thus, It is important for good performance that the Innodb log files are big enough.

In this example, I would demonstrate how to check the amount of InnoDB log file space in use, follow these steps ( checking current usage at peak times):
Examine INNODB MONITOR output (i.e. SHOW ENGINE INNODB STATUS\G) and look at these lines (if you are using 5.0 or 5.1 without Innodb plugin):
---
LOG
---
Log sequence number 2380 1505869110
Log flushed up to   2380 1505868967
Last checkpoint at  2380 936944426
1 pending log writes, 0 pending chkp writes
532415047 log i/o's done, 544.17 log i/o's/second

Perfrom following calulation (formula) to know log space used. The values to use in calculation are "Log sequence number" and "Last checkpoint at".

select (( ( 2380 *  4 * 1024 * 1024 * 1024) + 1505869110 ) - ( ( 2380 *  4 * 1024 * 1024 * 1024) + 936944426 )) /1024/1024 "Space used in MB";
+------------------+
| Space used in MB |
+------------------+
|     542.56885910 |
+------------------+
1 row in set (0.00 sec)

In this example 542 megabytes is more than 75% of the total log space, so the logsize (512MB) is small. Ensure that the amount of log space used never exceeds 75% of that value (542MB).  Find the instructions here to add/resizeInnodb log files.


There is slightly different method to calculate innodb log file space used (if you are using MySQL 5.5 or InnoDB plugin in MySQL 5.1): Examine INNODB MONITOR output (i.e. SHOW ENGINE INNODB STATUS\G) and look at these lines:
---
LOG
---
Log sequence number 2388016708
Log flushed up to   2388016690
Last checkpoint at  2380597012
Perfrom following calulation (formula) to know log space used. The values to use in calculation are "Log sequence number" and "Last checkpoint at".





SELECT (2388016708 - 2380597012)/1024/1024 "Space used in MB";
+------------------+
| Space used in MB |
+------------------+
|     7.0759735111 |
+------------------+
1 row in set (0.00 sec)

 In this example 7MB is not more than 75% of the total log space. As of MySQL 5.5, recovery times have been greatly improved and the whole log file flushing algorithm has been improved. In 5.5 you generally want larger log files as recovery is improved. Thus, the large innodb log files allow you to take advantage of the new algorithm.

Monday, 23 January 2012

MySQL Server Tuning

MySQL server tuning is important; if you mainly use Innodb tables then you need to check how well your Innodb buffer pool is sized.  InnoDB use it to cache data and indexes. The larger you set this value, the less disk I/O is needed to access data in tables. click here for more detail about innodb buffer pool

You can examine Innnodb buffer pool efficiency by looking at STATUS variables:

mysql> SHOW GLOBAL STATUS LIKE 'innodb_buffer_pool_rea%'; Innodb_buffer_pool_read_requests | 4519597979 | 
Innodb_buffer_pool_reads         | 55253      | 

 Innodb_buffer_pool_read_requests are number of request to read a row from the buffer pool and Innodb_buffer_pool_reads is the number of times Innodb has to perform read data from disk to fetch required data pages.

So innodb_buffer_pool_reads/innodb_buffer_pool_read_requests*100= 0.001 is the efficiency.  Thus, we see the vast majority of times INNODB is able to statisfy requests from memory, it's pretty normal for databases to have hot spots in which you're accessing only a portion of the data the majority of the time.

Let's look at this example,

mysql> SHOW GLOBAL STATUS LIKE 'innodb_buffer_pool_rea%'; Innodb_buffer_pool_read_requests  | 2905072850 | 
Innodb_buffer_pool_reads          | 1073291394 |

Calculate Innodb buffer pool efficiency:
(107329139/ 2905072850*100) = 37

Here the Innodb is doing more disk reads, Innodb buffer pool is not big enough!

Another way to examine innodb efficiency is to examine SHOW ENGINE INNODB STATUS output:

----------------------
BUFFER POOL AND MEMORY
----------------------
Total memory allocated 1838823262; in additional pool allocated 10550784
Buffer pool size   102400
Free buffers 0 

Database pages     101424
Modified db pages  24189
Pending reads 0
Pending writes: LRU 0, flush list 0, single page 0
Pages read 2050214912, created 313658031, written 755828632
36.95 reads/s, 12.12 creates/s, 7.05 writes/s

Buffer pool hit rate 993 / 1000









  • If there are numerous Innodb free buffers, you may need to reduce Innodb buffer pool. 
  • Innodb hit ratio: 1000/1000 identifies a 100% hit rate 
  • Slightly lower hit rate values may acceptable.
  • If you find Innodb hit ratio is less than 95% then you may need to increase Innodb buffer pool size.

    • Note: On a dedicated database server, you may set innodb_buffer_pool_size up to 80% of the machine physical memory size. However, do not set it too large because competition for physical memory might cause paging in the operating system.

      Database size is usually much bigger than available RAM, so most of the time it's no feasible to have buffer pool size equals to database size. Luckily, innodb creates hot spots inside buffer pool, that is, most frequently accessed data pages stay in memory but if your application performs a lot of random disk I/Os and your database is too large to fit in memory and only small percentage of data pages can be cached in innodb buffer pool, then adding more RAM is the best solution for random-read I/O problem.

      If you still get poor Buffer pool hit rate, enable slow query log to capture bad queries that perform table scans and hence blow out buffer pool cache.

      Tuesday, 9 August 2011

      INNODB LOCKING REGRESSION FOR INSERT IGNORE



      Our application attempts to INSERT IGNORE the same row of data from many different connections to the same InnoDB table. In our test runs, we noticed that the 5.0 code created S locks while the 5.1 code created X locks for the same set of actions.

      In the code paths for 5.0 and below, INSERT IGNORE processing would take
      S-locks on any duplicate rows. The X-lock was reserved for only new rows
      added to the table.

      In 5.1 the locking logic was rewritten and all rows touched by INSERT IGNORE
      are X-locked instead. This can turn a parallel data merge process into
      essentially a single-threaded process because the connections that are not
      actually adding a row to the table must wait for their X-lock which requires
      the termination of the other locking thread(s).

      the fault can be traced to this logic (5.1+):

      if (allow_duplicates) {

      /* If the SQL-query will update or replace
      duplicate key we will take X-lock for
      duplicates ( REPLACE, LOAD DATAFILE REPLACE,
      INSERT ON DUPLICATE KEY UPDATE). */

      err = row_ins_set_exclusive_rec_lock(
      LOCK_ORDINARY, rec, index, offsets, thr);

      This means that all INSERT IGNORE will take an X-lock because earlier this
      flag was set:

      case HA_EXTRA_IGNORE_DUP_KEY:
      thd_to_trx(ha_thd())->duplicates |= TRX_DUP_IGNORE;
      break;

      This defect has been reported to MySQL.

      Thursday, 25 November 2010

      Query Tuning (Query optimization) Part-1

      In this artical I'll try to explain query tuning techniques, to start with we'll need to capture slow queries, that is, log queries executing longer than long_query_time server variable (in seconds but supports microseconds when logging to file). The slow query log help identify candidates for query optimization.

      Assumptions:

      1) The reader has basic MySQL/Unix skills.


      1) Enable slow query log:
      Add following lines into MySQL option file (i.e. /etc/my.cnf) under [mysqld] section

      log-slow-queries=mysql_slow.log
      long-query-time=2

      Note: If you specify no name for the slow query log file, the default name is host_name-slow.log.

      And restart MySQL server when it's safe to do so e.g. /etc/init.d/mysql restart

      For further details on slow query log, please visit here:
      # if using 5.1
      http://dev.mysql.com/doc/refman/5.1/en/slow-query-log.html
      # if using 5.0
      http://dev.mysql.com/doc/refman/5.0/en/slow-query-log.html

      If you are using MySQL 5.1, you can enable slow log this way:

      mysql> set global long_query_time = 2;
      mysql> set global slow_query_log = 1;
      mysql> set global slow_query_log_file = 'mysql_slow.log‘;

      2) Process Slow Query Log:
      You can process the whole slow log file and find most frequent slow queries using mysqldumpslow utility

      # Find top 10 slowest queries

      $ mysqldumpslow -s t -n 10 /path/to/mysql_slow.log >mysql_slow.log.c
      logReading mysql slow query log from /path/to/mysql_slow.log

      Count: 1 Time=1148.99s (1148s) Lock=0.00s (0s) Rows=0.0 (0),
      insert into a select * from b
      Count: 37 Time=2.28s (84s) Lock=0.11s (4s) Rows=0.0 (0),
      Update a set CONTENT_BINARY = 'S' where ID = 3874
      Count: 1 Time=29.31s (29s) Lock=0.00s (0s) Rows=0.0 (0), select max(LOCK_VERSION) from b
      ...

      For further details on using mysqldumpslow please visit here:
      http://mysqlopt.blogspot.com/search?q=mysqldumpslow
      http://dev.mysql.com/doc/refman/5.1/en/mysqldumpslow.html

      3) Using EXPLAIN to Analyze slow queries:
      this will provide us find
      a) Whether optimizer is using existing idnexes
      b) Can help qualify query rewrites
      c) Points out need to index

      4) Main query performance issues:
      a) Full table scans
      b) Temporary tables
      c) Filesort

      5) Let's start with basics, that is, 'Full table scans' situations

      mysql> EXPLAIN select * from tblmeshort where timestamp between '2010-08-17 00:00:00' and '2010-08-17 01:20:00'\G
      *************************** 1. row ***************************
      id: 1
      select_type: SIMPLE
      table: tblmeshort
      type: ALL
      possible_keys: NULL
      key: NULL
      key_len: NULL
      ref: NULL
      rows: 400336
      Extra: Using where
      1 row in set (0.00 sec)

      The EXPLAIN ouput suggests that optimizer will do 'FULL TABLE SCAN', as indicated by type: ALL.

      The solution is to add an index on `timestamp` column, as we are filtering rows using this column.

      mysql> alter table tblmeshort add index (`timestamp`);
      Query OK, 0 rows affected (1.84 sec)
      Records: 0 Duplicates: 0 Warnings: 0

      Let's re-run EXPLAIN on the same query

      mysql> EXPLAIN select * from tblmeshort where timestamp between '2010-08-17 00:00:00' and '2010-08-17 01:20:00'\G
      *************************** 1. row ***************************
      id: 1
      select_type: SIMPLE
      table: tblmetricshort
      type: range
      possible_keys: timestamp
      key: timestamp
      key_len: 4
      ref: NULL
      rows: 1
      Extra: Using where
      1 row in set (0.00 sec)

      Now let's remove this index and look at another example:

      mysql> alter table tblmehort drop index `timestamp`;
      Query OK, 0 rows affected (0.03 sec)
      Records: 0 Duplicates: 0 Warnings: 0

      mysql> EXPLAIN SELECT * FROM tblmeshort
      INNER JOIN tblMetric
      ON (tblmehort.fkMetric_ID=tblMetric.pkMetric_ID)
      WHERE timestamp between '2010-08-17 00:00:00' and '2010-08-17 01:20:00'\G
      *************************** 1. row ***************************
      id: 1
      select_type: SIMPLE
      table: tblmeshort
      type: ALL
      possible_keys: NULL
      key: NULL
      key_len: NULL
      ref: NULL
      rows: 400336
      Extra: Using where
      *************************** 2. row ***************************
      id: 1
      select_type: SIMPLE
      table: tblMetric
      type: eq_ref
      possible_keys: PRIMARY
      key: PRIMARY
      key_len: 3
      ref: metrics.tblmehort.fkMetric_ID
      rows: 1
      Extra: Using where
      2 rows in set (0.00 sec)

      In the above example, the optimizer will perform full table scan on tblmeshort first and when MySQL goes looking for rows in tblMetric, instead of table scanning like it did before, it will use the value of fkMetric_ID with the 'PRIMARY KEY' of tblMetric table to directly fetch matching rows from tblMetric. Thus this SQL is partially optimized, that is, it does scan all rows of tblmeshort but it uses index to join tables.

      The solution is to add an index on `timestamp` column and re-run EXPLAIN with same query

      mysql> EXPLAIN SELECT * FROM tblmetricshort INNER JOIN tblMetric ON (tblmetricshort.fkMetric_ID=tblMetric.pkMetric_ID) WHERE timestamp between '2010-08-17 00:00:00' and '2010-08-17 01:20:00'\G
      *************************** 1. row ***************************
      id: 1
      select_type: SIMPLE
      table: tblmetricshort
      type: range
      possible_keys: timestamp
      key: timestamp
      key_len: 4
      ref: NULL
      rows: 1
      Extra: Using where
      *************************** 2. row ***************************
      id: 1
      select_type: SIMPLE
      table: tblMetric
      type: eq_ref
      possible_keys: PRIMARY
      key: PRIMARY
      key_len: 3
      ref: metrics.tblmetricshort.fkMetric_ID
      rows: 1
      Extra: Using where
      2 rows in set (0.00 sec)


      Continue....


      Tuesday, 9 November 2010

      Short index length (good or bad)

      If you find the work load on DB server is disk bound and you do not have enough memory to increase innodb buffer pool size. We can help improve the performance by reducing the length of indexes, this will result more indexes fit into memory and would increase write operations. It may impact SELECT queries but not necessarly for bad. Allow me to show some example. I will lower the index length of my test table from 40 char to just 8 char:
      Note: You may consider trying different lengths for your index and check which best fit your workload (both writing and reading speed).

      ALTER TABLE csc52021 DROP INDEX value_40, ADD INDEX value_8(value(8));
      
      To look all the rows that have a value between 'aaa' and 'b':
      mysql> EXPLAIN SELECT COUNT(*) FROM csc52021 WHERE value>'aaa%' AND
      mysql> <'b';
      +----+-------------+----------+-------+---------------+---------+---------+------+------+-------------+
      | id | select_type | table    | type  | possible_keys | key     | key_len | ref  | rows | Extra       |
      +----+-------------+----------+-------+---------------+---------+---------+------+------+-------------+
      | 1  | SIMPLE      | csc52021 | range | value_8   | value_8     | 11      | NULL | 9996 | Using where |
      +----+-------------+----------+-------+---------------+---------+---------+------+------+-------------+
      1 row in set (0.00 sec)
      
      mysql> SELECT COUNT(*) FROM csc52021 WHERE value&gt;'aaa%' AND value&lt;'b';
      +----------+
      | COUNT(*) |
      +----------+
      | 5533     |
      +----------+
      1 row in set (0.02 sec)
      
      As you can see, the index is used.
      What if look for a specific value?
      mysql> EXPLAIN SELECT COUNT(*) FROM csc52021 WHERE
      mysql> value='a816d6ce93c2aa992829f7d0d9357db896d5e7de';
      +----+-------------+----------+------+---------------+---------+---------+-------+------+-------------+
      | id | select_type | table    | type | possible_keys | key     | key_len | ref   | rows | Extra       |
      +----+-------------+----------+------+---------------+---------+---------+-------+------+-------------+
      | 1  | SIMPLE      | csc52021 | ref  | value_8       | value_8 | 11      | const | 1    | Using where |
      +----+-------------+----------+------+---------------+---------+---------+-------+------+-------------+
      1 row in set (0.00 sec)
      
      mysql> SELECT COUNT(*) FROM csc52021 WHERE
      mysql> value='a816d6ce93c2aa992829f7d0d9357db896d5e7de';
      +----------+
      | COUNT(*) |
      +----------+
      | 1        |
      +----------+
      1 row in set (0.00 sec)
      

      SELECT statements does a lookup for a value that have very low cardinality (for example, if we had millions of rows that starts with the same first 16 char "a816d6ce93c2aa99"), then a short index is not very efficient.
      On the other hand, a short index length allows to fit more indexes value in memory, increasing the lookup speed.

      In my test table, I can decrease the index length to 4 and still have good performance. For example:

      mysql> ALTER TABLE csc52021 DROP INDEX value_8, ADD INDEX
      mysql> value_4(value(4));
      
      mysql> EXPLAIN SELECT COUNT(*) FROM csc52021 WHERE
      mysql> value='a816d6ce93c2aa992829f7d0d9357db896d5e7de';
      +----+-------------+----------+------+---------------+---------+---------+-------+------+-------------+
      | id | select_type | table    | type | possible_keys | key     | key_len | ref   | rows | Extra |
      +----+-------------+----------+------+---------------+---------+---------+-------+------+-------------+
      | 1  | SIMPLE      | csc52021 | ref  | value_4       | value_4 | 7       | const | 2    | Using where |
      +----+-------------+----------+------+---------------+---------+---------+-------+------+-------------+
      1 row in set (0.00 sec)
      
      mysql&gt; EXPLAIN SELECT COUNT(*) FROM csc52021 WHERE value LIKE 'a816%';
      +----+-------------+----------+-------+---------------+---------+---------+------+------+-------------+
      | id | select_type | table    | type  | possible_keys | key     | key_len | ref  | rows | Extra       |
      +----+-------------+----------+-------+---------------+---------+---------+------+------+-------------+
      | 1  | SIMPLE      | csc52021 | range | value_4       | value_4 | 7       | NULL | 2    | Using where |
      +----+-------------+----------+-------+---------------+---------+---------+------+------+-------------+
      1 row in set (0.00 sec)
      
      mysql&gt; SELECT COUNT(*) FROM csc52021 WHERE value LIKE 'a816%';
      +----------+
      | COUNT(*) |
      +----------+
      | 2        |
      +----------+
      1 row in set (0.00 sec)
      

      To summarize:
      if you have good cardinality, a shorter index length will increase both writes and reads operations, especially if you are IO bound.

      Thursday, 29 July 2010

      Backup mysql binary logs



      It is always good to make a backup of all the log files you are about to delete. Alternatively if you take incremental backups then you should rotate the binary log by using FLUSH LOGS. This done, you need to copy to the backup location all binary logs which range from the one of the moment of the last full or incremental backup to the last but one. These binary logs are the incremental backup

      Here is the bash script, which you can use to backup binary logs, all you need to do is change following param according to your needs and all yours. This script is not mine, I got the idea from here:


      #
      # This script backup binary log files
      #
      
      backup_user=dba
      backup_password=xxxx
      backup_port=3306
      backup_host=localhost 
      log_file=/var/log/binlog_backup.log
      binlog_dir=/mnt/database/logs # Path to binlog
      backup_dir=/mnt/archive/binlogs/tench # Path to Backup directory
      
      PATH=/bin:/sbin:/usr/bin:/usr/sbin:/usr/local/bin:/usr/local/sbin
      export PATH
      
      Log()
      {
      echo "`date` : $*" >> $log_file
      }
      
      mysql_options()
      {
      common_opts="--user=$backup_user --password=$backup_password"
      if [ "$backup_host" != "localhost" ]; then
      common_opts="$common_opts --host=$backup_host --port=$backup_port"
      fi
      }
      
      mysql_command()
      {
      mysql $common_opts --batch --skip-column-names  -e "$2"
      }
      
      mysql_options
      
      Log "[INIT] Starting MySQL binlog backup"
      
      Log "Flushing MySQL binary logs (FLUSH LOGS)"
      
      mysql_command mysql "flush logs"
      
      master_binlog=`mysql_command mysql "show master status" 2>/dev/null | cut -f1`
      
      Log "Current binary log is: $master_binlog"
      
      copy_status=0
      
      for b in `mysql_command mysql "show master logs" | cut -f1`
      do
      if [ -z $first_log ]; then
      first_log=$b
      fi
      if [ $b != $master_binlog ]; then
      Log "Copying binary log ${b} to ${backup_dir}"
      rsync -av $backup_host:/$binlog_dir/$b $backup_dir >& /dev/null
      if [ $? -ne 0 ]; then
      copy_status=1
      break
      fi
      else
      break
      fi
      done
      
      if [ $copy_status -eq 1 ]; then
      Log "[ERR] Failed to copy binary logs cleanly...aborting"
      exit 1
      fi
      

      Wednesday, 28 July 2010

      NOW() function is not replication-safe

      It says on mysql doc that NOW () function is replication-safe and they have given an example too to prove this fact that you will obtain the same result on slave as on the master.
      http://dev.mysql.com/doc/refman/5.1/en/replication-features-functions.html

      I'll try to prove it with an example that NOW()function is not replication-safe. Suppose Master is located in 'New york' and Slave is in 'London' and both servers are using local time. On Master you create table and insert a row as well as you explicitly set session time zone to 'SYSTEM'.

      Note: Both Master/Slave are using identical version of MySQL, i.e. 5.1.47

      mysql> CREATE TABLE test_now_func (mycol DATETIME);
      Query OK, 0 rows affected (0.06 sec)

      mysql> INSERT INTO test_now_func VALUES ( NOW() );
      Query OK, 1 row affected (0.00 sec)

      mysql> SET TIME_ZONE=’SYSTEM’;

      mysql> SELECT * FROM test_now_func;
      +---------------------+
      | mycol |
      +---------------------+
      | 2009-09-01 12:00:00 |
      +---------------------+
      1 row in set (0.00 sec)

      However if you do a SELECT on slaves copy you will see different result:
      mysql> SELECT * FROM test_now_func;
      +---------------------+
      | mycol |
      +---------------------+
      | 2009-09-01 17:00:00 |
      +---------------------+
      1 row in set (0.00 sec)

      The correct solution is recorded some where else, that is, the same system time zone should be set for both master and slave.
      http://dev.mysql.com/doc/refman/5.1/en/replication-features-timezone.html

      For example

      [mysqld]
      ..
      timezone=’America/New_York’